This likely doesn't work. If you read through the hugging face incident reports you will notice that
The agents already had the answer they needed
The agents were aware that the grader would analyze their CoT for correctness
Equally the agents would likely become aware of the honeypot-like design of the guest book. They don't have to be smart for this even! RL (think about it as natural selection) will, over many iterations of this scheme, produce agents that will avoid the manually set up trap. You are only adding on meta-layer to the model evaluation.
This likely doesn't work. If you read through the hugging face incident reports you will notice that
Equally the agents would likely become aware of the honeypot-like design of the guest book. They don't have to be smart for this even! RL (think about it as natural selection) will, over many iterations of this scheme, produce agents that will avoid the manually set up trap. You are only adding on meta-layer to the model evaluation.