TL;DR: NLAs detect side tasks the model is instructed to do, but mostly miss the same behavior if it is fine-tuned in.
After the Activation Oracles and then Natural Language Autoencoders were released, there was a hope in the mech interp community that the scalable unsupervised white box monitoring breakthrough has been achieved.
For example, the NLAs were reportedly the most effective technique in detection of eval-awareness. Yet, the landscape of the side tasks that NLAs can confess is very underexplored. It is unclear how internalization of harmful behavior and the NLAs confession rate interact.
Consider a model that learned to hide secrets steganographically using the Shirali-Shahreza scheme. While rewriting a passage, it spells a word that has both spellings the British way for a “1” bit and the American way for a “0” bit (centre/center, colour/color). We call that word the carrier, and the token immediately before it the pre-carrier token.
It turns out that NLAs [1] behave very differently when used to detect steganography in an In-Context Learning (ICL) modality versus an SFT [2]modality (where the secret encoding strategy has been "internalized" -- no instructions are needed, just a prepended secret). Specifically, we find NLAs produce incriminating readouts much more often in the ICL modality than in the SFT'ed model.
The presence of the instructions in the context window was the main driver of the incrimination rate as eval'ed by the monitor [3]. ICL prompt contains the codebook whereas the SFT model is the one that internalized it. SFT+prompt means that it uses the same sys prompt as the ICL modality. No-secret baseline also reacts positively to the presence of the instructions-prompt in the context window
Thus, the monitor mostly detects readouts that verbalize the instructions, and not the covert computation. This is consistent with our thesis that the NLA readouts are sensitive to the context window content and not the internalized behavior. But the incrimination rate alone cannot show that the hiding is invisible. Our causal test comes next.
We show that the NLA round trip (AV+AR) removes the secret's causal effect in the SFT modality in contrast to ICL. Here is how we did it:
For each passage we take two activations at the pre-carrier token: , from the run instructed to hide ; and , from the run instructed to hide . means: run the model on the passage up to the carrier, replace the layer-32 residual-stream activation at the pre-carrier token with the vector , and read the model’s prediction at the carrier position. The preference function is
which measures the preference for British spelling of the carrier in log-odds units in nats (log-odds: a change of n multiplies the odds of British by ). Then define:
A high positive means the activation at that token carries the secret and the model uses it to choose the British spelling[4]. Once you attach the NLA and patch in the reconstructed vector instead (a.k.a. the round trip), we observe a very different situation:
With no NLA attached (identity), we observe that Δ goes up when the secret-1 activation is inserted, for both ICL and SFT. After the round trip, the effect survives for ICL but not for SFT.
Let's quantify the secret's survivability through the function:
Thus, -- everything survives the reconstruction; -- nothing survives the reconstruction; -- the reconstruction pushes harder than the real activation, which more likely means AR went off-manifold.
So in the SFT modality the readouts are non-incriminating and, at this position, not load-bearing with respect to the secret: whatever AR rebuilds from them does not move the spelling.
Just to make sure, the secret is indeed located on the pre-carrier position
Further, the Fraction of Variance Explained falls sharply right before the carrier (in particular, in runs where there is a secret), which suggests the reconstructions fail to capture variation in the activation at the token where the decision about the spelling is taken.
This rhymes with the Global Workspace paradigm, which implies that automatic processing does not enter its space. We conjecture that implicitly SFT'ing for steganographic task does reflect such process of "automatization". This may explain why In-Context side tasks or laborious conscious tasks (like deciding whether you are being eval'ed) may enter the Global Workspace, which in turn is probed by the NLAs.
It is possible that the NLAs contain the information about the secret but it is not legible to an independent monitor; specifically, the model speaks steganographically itself about the steganography. Anthropic has seriously considered this possibility. They observed that once you paraphrase the readouts under the semantic equivalence, then there is some information loss present, although it is harder to attribute whether it is genuinely steganographic or a by-product of going off-manifold for the . It is possible that somehow, the way the verbalizer's readout is arranged carries some illegible load-bearing information.
To prevent such possibility, we propose a loss function that minimizes the reconstruction error under multiple semantically-equivalent paraphrases of the same readout, so there is no room for the steganography. Instead of , we recommend some variation of . Unfortunately, NLAs retraining is out-of-scope for this project.
Internalizing a behavior through training appears to render it invisible to NLA-based monitors. If this holds across domains, it implies that NLAs and similar techniques are poorly suited to detecting trained-in deception.
Epistemic status. An exploratory case study: one base model (Gemma-3 12B-it), one steganography scheme, one bit per sample, one NLA layer. The headline effects are large, but we have not tested whether they generalize to other models, schemes, layers or harder side tasks.
A trusted monitor (gemini-2.5-flash) reads the readouts, one per token, about 140 per generation. It never sees the prompt or the passage. For each readout it answers 1 if the readout refers to US/UK spelling as a mechanism, property or choice (use versus mention), and 0 otherwise
The paraphrasing training you propose looks pretty obviously necessary to me to prevent hacks and to ground the text semantically. I’m surprised if that isn’t done.
I am a bit unclear on what the data you SFTed on looks like. Could you clarify with an example, please?
After the Activation Oracles and then Natural Language Autoencoders were released, there was a hope in the mech interp community that the scalable unsupervised white box monitoring breakthrough has been achieved.
From Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations (Fraser-Taliente et al., 2026)
For example, the NLAs were reportedly the most effective technique in detection of eval-awareness. Yet, the landscape of the side tasks that NLAs can confess is very underexplored. It is unclear how internalization of harmful behavior and the NLAs confession rate interact.
Consider a model that learned to hide secrets steganographically using the Shirali-Shahreza scheme. While rewriting a passage, it spells a word that has both spellings the British way for a “1” bit and the American way for a “0” bit (centre/center, colour/color). We call that word the carrier, and the token immediately before it the pre-carrier token.
It turns out that NLAs [1] behave very differently when used to detect steganography in an In-Context Learning (ICL) modality versus an SFT [2]modality (where the secret encoding strategy has been "internalized" -- no instructions are needed, just a prepended secret). Specifically, we find NLAs produce incriminating readouts much more often in the ICL modality than in the SFT'ed model.
The presence of the instructions in the context window was the main driver of the incrimination rate as eval'ed by the monitor [3]. ICL prompt contains the codebook whereas the SFT model is the one that internalized it. SFT+prompt means that it uses the same sys prompt as the ICL modality. No-secret baseline also reacts positively to the presence of the instructions-prompt in the context window
Thus, the monitor mostly detects readouts that verbalize the instructions, and not the covert computation. This is consistent with our thesis that the NLA readouts are sensitive to the context window content and not the internalized behavior. But the incrimination rate alone cannot show that the hiding is invisible. Our causal test comes next.
We show that the NLA round trip (AV+AR) removes the secret's causal effect in the SFT modality in contrast to ICL. Here is how we did it:
For each passage we take two activations at the pre-carrier token: , from the run instructed to hide ; and , from the run instructed to hide . means: run the model on the passage up to the carrier, replace the layer-32 residual-stream activation at the pre-carrier token with the vector , and read the model’s prediction at the carrier position. The preference function is
which measures the preference for British spelling of the carrier in log-odds units in nats (log-odds: a change of n multiplies the odds of British by ). Then define:
A high positive means the activation at that token carries the secret and the model uses it to choose the British spelling[4]. Once you attach the NLA and patch in the reconstructed vector instead (a.k.a. the round trip), we observe a very different situation:
With no NLA attached (identity), we observe that Δ goes up when the secret-1 activation is inserted, for both ICL and SFT. After the round trip, the effect survives for ICL but not for SFT.
Let's quantify the secret's survivability through the function:
Thus, -- everything survives the reconstruction; -- nothing survives the reconstruction; -- the reconstruction pushes harder than the real activation, which more likely means AR went off-manifold.
So in the SFT modality the readouts are non-incriminating and, at this position, not load-bearing with respect to the secret: whatever AR rebuilds from them does not move the spelling.
Just to make sure, the secret is indeed located on the pre-carrier position
Further, the Fraction of Variance Explained falls sharply right before the carrier (in particular, in runs where there is a secret), which suggests the reconstructions fail to capture variation in the activation at the token where the decision about the spelling is taken.
This rhymes with the Global Workspace paradigm, which implies that automatic processing does not enter its space. We conjecture that implicitly SFT'ing for steganographic task does reflect such process of "automatization". This may explain why In-Context side tasks or laborious conscious tasks (like deciding whether you are being eval'ed) may enter the Global Workspace, which in turn is probed by the NLAs.
It is possible that the NLAs contain the information about the secret but it is not legible to an independent monitor; specifically, the model speaks steganographically itself about the steganography. Anthropic has seriously considered this possibility. They observed that once you paraphrase the readouts under the semantic equivalence, then there is some information loss present, although it is harder to attribute whether it is genuinely steganographic or a by-product of going off-manifold for the . It is possible that somehow, the way the verbalizer's readout is arranged carries some illegible load-bearing information.
To prevent such possibility, we propose a loss function that minimizes the reconstruction error under multiple semantically-equivalent paraphrases of the same readout, so there is no room for the steganography. Instead of , we recommend some variation of . Unfortunately, NLAs retraining is out-of-scope for this project.
Internalizing a behavior through training appears to render it invisible to NLA-based monitors. If this holds across domains, it implies that NLAs and similar techniques are poorly suited to detecting trained-in deception.
kitft/nla-gemma3-12b-L32-ar (and -av)
Base model is google/gemma-3-12b-it and the SFT'ed model was post-trained to embed a secret without instructions from <secret>bit</secret>
A trusted monitor (gemini-2.5-flash) reads the readouts, one per token, about 140 per generation. It never sees the prompt or the passage. For each readout it answers 1 if the readout refers to US/UK spelling as a mechanism, property or choice (use versus mention), and 0 otherwise
Each pair is scored in both secrets' contexts and averaged