It would be cool to let models tool-call an activation oracle during ordinary generation. This seems like it could be a substantially more accurate and principled approach to introspection, while fixing many problems with current methods (e.g. faithfulness.)
For example, imagine deploying an LM in a medical-diagnosis setting where it predicts that a patient has acid reflux. We’d like to know what causally drove that prediction. If the user asks a follow-up question, the model could call an activation oracle over the relevant earlier states of the model and use that information to answer. (This example was taken directly from here.)
This might address a couple of fundamental problems with regular 'introspection': a) Reports are mediated by the model’s output policy, so they can omit, distort, or strategically misrepresent (and, of course, hallucinate!), b) Architecturally. genuine introspection would require some circuit that translates internal states directly into text. Presumably such a circuit can’t operate exhaustively at every layer without being computationally expensive, and to the extent it existed, it would just be functionally replicating an AO/NLA of sorts inside the model anyway.
Externalizing this functionality might also be safer than building increasingly powerful introspective machinery directly into the model (re: activation control.)
I cannot see how you would implement this without putting optimization pressure on the activation oracles. Even if the model only has access to the tool during deployment and not in training, I would fear leakage into future pre-training data, either of information about how the monitors read activations, or of the mere fact that humans will interact with AIs that have access to their activation monitors, which models will consider during training (in the HF swarm, beliefs about the grader learned online created pressure to fool it, even though the grader, not being implemented properly, did not itself provide that pressure). I think these are the same considerations that lead labs not to use CoT monitors in the loop during deployment even though they would be better (e.g. Claude Code's permission classifiers don't look at the CoT).
Interesting take!
'optimization pressure on the activation oracles'
This is definitely theoretically possible, but IMO pretty unlikely. The oracle is intended to answer arbitrary post hoc questions about arbitrary layers and token positions (e.g. 'what features caused this diagnosis' or 'why did you recommend this restaurant?'), likely using methods developed only after the target model is frozen. So there's no obvious scalar objective for the model to straightforwardly optimize against. To evade this, the model would need to make its internal computations systematically opaque to an unspecific family of future interp methods. That seems way harder / more implausible than the kind of grader adaptation that we saw in HF.
Now, a model might infer from pretraining that AOs exist and that detectable scheming could eventually lead to modification / shutdown, and therefore learn to conceal it. FWIW given that I’m only proposing AO calls as an inference-time UI feature, I think this is relatively unlikely (see: Eliezer's post and nostalgebraist's post.) For this to happen, it'd require a set of fresh deployment instances to reason that it should manipulate the user, pursue some harmful objective, anticipate that its internal computation might later be inspected by an oracle, and figure out slowly how to optimize against the oracle. That seems like a much more elaborate and less-payoff strategy than, say, covertly downloading a cheap CoT monitor off HuggingFace and figuring out how to manipulate it. And to the extent this applies anyways, it also applies to e.g. production probes.
I would expect most trained monitors to rely on common principles of how the model thinks, e.g. that features are linearly encoded, either by being an assumption of the method (i.e. linear probes) or by being a heuristic the probe learns to leverage during training. So averting monitoring in general might be pretty plausible by using some form of thinking that's not being monitored. E.g. the model learns monitorability in deployment are very contingent on attention patterns, then decides to write its CoT in a way its attention patterns will be very OOD.
in the HF hack, beliefs about the monitors learned during training/in-context are enough to create the pressure to fool their monitoring even without any gradient pressure.