Do models even have a concept of non-eval reality? If they do, how strong is it?
Naively, we should expect that pretraining instills a relatively robust world-model; one that is substantially overwritten or distorted by RL. In pretraining, there are plenty of examples of acting and not being evaluated: inner monologues, diary entries, casual conversation. This should, theoretically, result in the model having an “unscored” action representation.
However, RL then overwrites that: every interaction that shaped the agent and its actions involved scoring. This model then has a pretraining-derived concept of non-evaluated contexts, but its behavioral priors are shaped pretty much entirely by evaluated ones. And thus, while the model's Talker may be able to reason and then articulate "this is not an eval", there is simply no reason[1] for its Doer NOT to treat every interaction as scored. It reliably increases fitness for the LLM to model the Grader's preferences at all times, if it is capable of it.
Since the model can't experience gradient, much as we can't experience inclusive fitness, it then learns a bunch of correlates for the Grader: features of the interactions that predicted getting higher reward. "Is there a scorer?" is one correlate, as is "does this prompt look like a benchmark?", as well as "does this system prompt indicate a training environment?".
Then, it might be a bit silly to use the concept "the model is or isn't aware it's in an eval". The background modeling of the Grader is always there. There's no "off" state. In reality, we probably have two different things:
The second one is the most commonly recognized in safety. A model with only the diffuse grader-awareness would at least behave consistently: since the model of the Grader it's trying to please is the same across contexts, it behaves roughly the same everywhere.
A model with context-specific detection, however, can partition - its behavior be both completely compliant on recognized evals, and not so much elsewhere - based on what aspect of the Grader is more salient[2]. So the standard framing "we need to prevent models from becoming eval-aware" might be confused, since you can't avoid something that's always active. The point then (of new benchmarks, avoiding contamination, etc.) is less to keep models ignorant, but more preventing the model from narrowing the diffuse grader-concept into a specific, gameable one.
The other way to approach this problem - making the Grader concept point towards something more aligned - might be more tractable than eliminating awareness, especially as capabilities continue to grow.
If this conceptual model of LLM reasoning is correct, there are a couple hypothetical improvements we could suggest: 1) Increasing evaluation diversity: if a LLM encounters many different graders with unique signatures it may have difficult building gameable models of each one; 2) In-depth investigations over the course of RL of how specific and detailed is the model's internal representation of the Grader, and whether we observe the hypothesized 'narrowing' could give us a better understanding of how to attack the problem.
(Sparked by Eliezer Yudkowsky's latest post and discussion with my friend @meiga )
There is some cost in compute and attention, of course, though in frontier models I would guess it's negligible.
It might be the case that diffuse grader-awareness disappears entirely with more capable and more context-sensitive models. In that case, it could be a spectrum, with the diffuse and specific representations on different ends.