Evaluation Awareness in Small(ish) Models
by John Robertson and Zach Perlman
TL;DR We seek to identify open source reasoning models which are both small enough for white-box interpretability and display evaluation-gaming behavior. We find that how often models verbalize their awareness varies from very rarely to a third of the time, and appears largely unrelated to model size. Within the same...
Oct 17