Is Eval Gaming Downstream of Verbalized Eval Awareness? Not when it's reflexive.
Code and data available at github.com/KieronKretschmar/latent-awareness TL;DR * We take two eval-gaming model organisms (Hua et al.'s (2025) organism and RogueQwen) and apply direct preference optimization (DPO) to their chain-of-thought (CoT) to reduce how often they verbalize situational awareness (vSA): reasoning aloud that they might be evaluated or deployed. *...
Aug 1011