The autoreconstructor (AR) in a natural language autoencoder (NLA) is the entire argument for the validity of the method, and at present is poorly understood.
The AR does not condition strongly on English grammar or semantics. Three bits of evidence:
Word pairs that are syntactically related interact no more than matched pairs that are not syntactically related. See dependency arcs.
The fraction of variance explained (FVE) lost by removing or flipping a negation word ("not" -> ""; "without" -> "with") is only weakly correlated with the FVE lost by ablating the negated word (Spearman rho 0.38). See negation.
The word classes that carry the most meaning, nouns (forest), adjectives (heavy) and verbs (reported), cost the least FVE to delete; auxiliaries (is), pronouns (it), numerals and subordinators cost the most. See word class.
Results are consistent with prior work suggesting that the AR is more heavily influenced by the format of the verbalisation (tone, register, general domain, structure of text) than by semantic meaning per se.
The problem with NLAs
NLAs have strong empirical performance on an automated auditing task. A model given access to a NLA was able to uncover the root motivation of a reward-sycophancy model organism more consistently than without. The NLA provided marginal benefit to an auditing agent with access to the full pre-training, supervised fine-tuning and reinforcement learning corpora, as compared to a sparse autoencoder which offered no benefit. This happened because the activation verbaliser (AV) included mentions of reward-sycophancy in verbalisations of activations from the base model (call the activation under investigation the 'base activation').
The mechanistic case for the method is not so strong. In an attempt to tether the verbalisation to the base activation, the AV is trained on cosine similarity between the reconstructed activation and the base activation. The hope is that the NLA would encode the semantic meaning of the activation in the verbalisation in order to reconstruct it with high fidelity, and that it would be encoded in coherent English. That is the entire mechanistic argument for the validity of the method.
That hope is a testable hypothesis: that the AR reads the verbalisation as semantic meaning encoded in content words and English syntax.
Setup
The base model used for the experiments was Qwen3.6-27B, and the NLA was Celeste Selder's Qwen3.6-27B NLA. The AV is an adapter checkpoint from the 600th reinforcement learning step. The AR is from the same step, predicting the layer 42 activation.
The dataset was a corpus of 5000 FineFineWeb documents, which I filtered based on length and category (see Appendix A for details). I randomly selected a token position after the 128th token to run the base Qwen model on and extract a layer 42 activation to verbalise.
FVE points are the primary unit of this writeup and figures, defined as 100 x the drop in FVE of the reconstructed activation from an edited verbalisation against the unedited baseline.
Findings
The AR isn't reading English how a person would. I found several pieces of evidence showing that the AR does not condition heavily on the syntax of English prose or on content words. Instead, I favour the hypothesis that the format of the verbalisation is what matters. That includes the tone, register, general domain and structure of the text (see Reconstruction Without Recall).
Negation words
I intervened on five types of negation words: "not", "n't" (at the end of a word), "without", "no" and "never". There were six types of intervention:
intervention
example
n
flip the negation
"don't budge" → "do budge"; "without public disclosure" → "with public disclosure"
309
delete the negation ("no", "without", "never" only)
"for no audience" → "for audience"
61
insert "not" after an auxiliary
"she did receive approval" → "she did not receive approval"
200
insert "just" after the same auxiliary (control)
"she did receive approval" → "she did just receive approval"
200
delete the governed word
"don't budge" → "don't"
309
swap the governed word
"don't budge" → "don't fostered"
308
I also measured the effect on FVE of ablating the negated word either by deleting the word or by using the corpus-swap ablation strategy discussed in Appendix B.
Consider this example: "...they were not walking in the forest...", where 'not' is the negation and 'walking' is the governed word. Flipping the negation costs about as much FVE on average as ablating the governed word (the paired difference is −0.09 FVE points, 95% interval −0.36 to +0.18). Within an instance, however, the cost of flipping the negation is only weakly correlated with the cost of ablating it (Spearman rho 0.38 over 308 negators), and neither edit is reliably costlier. The flip costs more in 46% of instances.
Scatter plot showing FVE lost from flipping the negation word vs corpus-swapping the word it governs
The insertion control is the one place negation beat a matched edit. In the 200 verbalisations with no negation, inserting "not" after an auxiliary cost 0.16 FVE points and inserting "just" in the same slot cost 0.06 (paired difference 0.10, 95% interval 0.05 to 0.16). So the negator is not invisible, but the gap is about a fifth of the cost of flipping an existing negation and small beside the deletion and swap effects, and it does not change the picture.
Dependency arcs against matched controls
English has natural dependencies (not just negation) that can be computed with natural language parsers like Stanza.
I used those dependencies to show how closely related two tokens are syntactically, and then compare the interaction against a matched pair of words that are the same number of tokens apart, the same number of tokens per word, the same in-quote status, and a similar position in the verbalisation. Ablations were done by deletion (see Appendix B). The interaction is defined as: interaction = δ(ablate A) + δ(ablate B) − δ(ablate both)
Word pairs on a dependency arc interact no more than matched pairs without one on 9,813 arc pairs across 857 verbalisations, after adjusting for covariates.
I found no evidence that a syntactic dependency between two words makes them interact more than matched controls.
Nouns, verbs and adjectives are the cheapest words to delete
Deleting a noun, adjective or verb barely moves FVE over 40 verbalisations, and determiners (the, a) and adpositions (of, in) sit with them near zero. The costly deletions are auxiliaries (is, will), pronouns (it, they), subordinators (because, if) and numerals, at roughly half an FVE point each, though the numeral estimate is heavy-tailed and its interval includes zero.
Conclusion
There are three findings. (1) flipping a negation costs the AR no more FVE than a matched edit that leaves the meaning alone, (2) syntactic relationships do not predict interaction strength between words in the verbalisation, and (3) the words that carry the most meaning, nouns, verbs and adjectives, are the cheapest to delete. Taken together, these findings are enough to reject the hypothesis that the AR reads a verbalisation as a human would, from its semantic meaning. However, these findings make perfect sense if the AR is conditioning primarily on the format of the verbalisation.
Future work should try to disprove the standing hypothesis that the reconstructor conditions mostly on format, or quantify how much format contributes relative to semantics.
The dataset was a corpus of 5000 FineFineWeb documents, taken uniformly from FineFineWeb document categories but excluding weapons science, nuclear science, chemistry, biology, gambling and relationships. After the category-based exclusion I then did a regex search to exclude documents relating to CBRN, explosives, drug synthesis, sexual content, gore and self-harm. I filtered the documents because I didn't want the output to be interacting with any residual model safeguards. I also ensured that documents were more than 160 tokens long.
After curating a dataset of documents I randomly selected a token after the 128th position to run through the base Qwen model and extract a layer 42 activation to verbalise. A random subset of 1100 of these tokens were run through the AV to produce verbalisations.
Appendix B: ablation strategies
I experimented with three ablation strategies:
Deletion.
Corpus-swap, where I averaged over swaps pulled from a corpus of 1100 verbalisations, matched on part-of-speech and Qwen token count.
Masked-LM substitution marginalised over draws, which I didn't end up using. Leakage was too high, the MLM put the original word back in 45% of draws.
I ran an experiment and learned that removal curves are actually gentler for outright deletion, which is also at least 8x more computationally efficient.
*FVE as words are removed one at a time, 20 documents. Every curve is concave and deletion is gentlest. The filler is underscore '_'*
Appendix C: examples for each finding
Examples are drawn randomly, none was chosen by hand. struck is the original word, bold is for substitutes, the two words of a pair are highlighted.
A document and its verbalisation
Source document, domain music and dance
The activation was read at token 292 of 1679, the last token shown here.
“Oxygen” is the second single to drop from Swan’s latest epic To Be Kind (after last month’s stellar, staggeringly funky “A Little God in My Hands”), and it is among the fiercest and most thrilling moments of its two-hour-plus runtime. It is certainly the most thoroughly noisy piece on the project, built on an unwavering no-wave bass groove and skittering drum line – guitars, lap steel, and finally brass falls piling on, amounting to something paradoxically crushing and danceable.
But the icing on the (dense, dense) cake is frontman Michael Gira’s vocals, which are at their most manic here, ranging from caveman-esque grunts on the bridge to a nasally call progressing into the second verse – “HEY THEEEEERRE.” It’s all just so cathartic; as a listener, you get release just from hearing these guys belt this thing out. It’s a sort of vicarious thrill that Swans has pretty consistently provided over the course of its now three-decade-long career.
And there’s also something to be said of the excitement this song manages to inspire despite the fact it has been kicking around for years in a number of iterations, in Gira’s solo acoustic sets, and then the live version that appeared on the limited edition live album Not Here/Not Now late last year. The band’s promotional/funding approach for these last few releases has
Verbalisation
Music review genre with analytical journalism tone — examining Angel Olsen track by track, building cumulative argument about album depth.
Narrative momentum building toward elaboration of delayed-release strategy: reviewer argues previous tracks enhance fully upon knowing the lead single's context, setting up a concluding insight about release strategy.
Final clause "Angel and Jagjagz's promotional approach to this album thus far has" is an incomplete predicate mid-sentence, demanding a completed thought describing their fan crowdsource/community engagement strategy (e.g., "been incremental," "involved releasing tracks early").
"one gets the sense of the emotion at play. It adds another dimension of weight knowing she learned it existed first. A few other tracks reveal similarly impressive levels of craft. Angel and Jagjagz's promotional approach to releasing music lately has" implies a discourse about the crowdsource/release timing strategy revealed before album drop..
Negation
Drawn from the 308 negators of run 10 that were both flipped and had the word they govern swapped. The governed-word swap shown is one of its draws, chosen at random.
negation flipped: there are activities that elite athletes have you looking for - you'll neveralways believe the number available. You are dropped any replies.
governed word swapped: ... activities that elite athletes have you looking for - you'll never believeoverwhelming the number available. You are dropped any replies.
negation flipped: but a miles away around the mountain - except the walk around isn't straight or direct, 20 kilometers of strenuous hiking wilderness trek to"
governed word swapped: ... mountain - except the walk around isn't straight or direct, 20 kilometersexperiences of strenuous hiking wilderness trek to"
negation flipped: "lf you buy an ACA-compliant individual health plan, you'll not face liability for" ends mid-sentence, grammatically requiring completion of a specific ...
governed word swapped: "lf you buy an ACA-compliant individual health plan, you'll not faceestablishing liability for" ends mid-sentence, grammatically requiring completion of a specific financial/legal ...
Dependency arcs against matched controls
Drawn from the 9813 arcs of run 11 that have a matched control. Each arc is followed by its own control.
words
arc: subordinator to clause
... expect the site name "Ringtones.ua" tocomplete it.
its control
... site "Ringtones.ua," consistently directing users to browse and download.
arc: noun modifying a noun
... statement like "August 2017 for sales of this price range in El Paso. This sale ...
its control
... continuing the ranking statement like "August 2017 for sales of this price range in El ...
arc: relative clause to its noun
... honors the story of the Amistad, in which West African slaves took" — a clause about resistance/defense ...
its control
Legislative news article tone — formal press release style covering California GPU Awards ...
Word class
Drawn from the 40 verbalisations of run 5 with at least one closed-class word and one content word deleted singly; within each, one word of each class chosen at random.
deleted word in context
auxiliary
"National Research Council" is a proper noun mid-introduction, strongly constraining the ...
content word (noun)
Herrera; the Montreal accredited marketplace paragraph is mid-sentence describing her leadership role.
subordinating conjunction
... "water" ends the closing directive sentence ("So if you want to get the most benefit ...
content word (adjective)
Health/wellness blog about acupuncture clinic, consistently emphasizing post-treatment water consumption to flush toxins and enhance benefits.
coordinating conjunction
... bankruptcy filing, debt restructuring, costs frozen operations, and now focuses on the CEO's quoted optimism ...
content word (adverb)
... a favorable outcome," requiring completion of the forward-looking positioning phrase, e.g., "for long-term success in ...
TL;DR
The problem with NLAs
NLAs have strong empirical performance on an automated auditing task. A model given access to a NLA was able to uncover the root motivation of a reward-sycophancy model organism more consistently than without. The NLA provided marginal benefit to an auditing agent with access to the full pre-training, supervised fine-tuning and reinforcement learning corpora, as compared to a sparse autoencoder which offered no benefit. This happened because the activation verbaliser (AV) included mentions of reward-sycophancy in verbalisations of activations from the base model (call the activation under investigation the 'base activation').
The mechanistic case for the method is not so strong. In an attempt to tether the verbalisation to the base activation, the AV is trained on cosine similarity between the reconstructed activation and the base activation. The hope is that the NLA would encode the semantic meaning of the activation in the verbalisation in order to reconstruct it with high fidelity, and that it would be encoded in coherent English. That is the entire mechanistic argument for the validity of the method.
That hope is a testable hypothesis: that the AR reads the verbalisation as semantic meaning encoded in content words and English syntax.
Setup
The base model used for the experiments was Qwen3.6-27B, and the NLA was Celeste Selder's Qwen3.6-27B NLA. The AV is an adapter checkpoint from the 600th reinforcement learning step. The AR is from the same step, predicting the layer 42 activation.
The dataset was a corpus of 5000 FineFineWeb documents, which I filtered based on length and category (see Appendix A for details). I randomly selected a token position after the 128th token to run the base Qwen model on and extract a layer 42 activation to verbalise.
FVE points are the primary unit of this writeup and figures, defined as 100 x the drop in FVE of the reconstructed activation from an edited verbalisation against the unedited baseline.
Findings
The AR isn't reading English how a person would. I found several pieces of evidence showing that the AR does not condition heavily on the syntax of English prose or on content words. Instead, I favour the hypothesis that the format of the verbalisation is what matters. That includes the tone, register, general domain and structure of the text (see Reconstruction Without Recall).
Negation words
I intervened on five types of negation words: "not", "n't" (at the end of a word), "without", "no" and "never". There were six types of intervention:
intervention
example
n
flip the negation
"don't budge" → "do budge"; "without public disclosure" → "with public disclosure"
309
delete the negation ("no", "without", "never" only)
"for no audience" → "for audience"
61
insert "not" after an auxiliary
"she did receive approval" → "she did not receive approval"
200
insert "just" after the same auxiliary (control)
"she did receive approval" → "she did just receive approval"
200
delete the governed word
"don't budge" → "don't"
309
swap the governed word
"don't budge" → "don't fostered"
308
I also measured the effect on FVE of ablating the negated word either by deleting the word or by using the corpus-swap ablation strategy discussed in Appendix B.
Consider this example: "...they were not walking in the forest...", where 'not' is the negation and 'walking' is the governed word. Flipping the negation costs about as much FVE on average as ablating the governed word (the paired difference is −0.09 FVE points, 95% interval −0.36 to +0.18). Within an instance, however, the cost of flipping the negation is only weakly correlated with the cost of ablating it (Spearman rho 0.38 over 308 negators), and neither edit is reliably costlier. The flip costs more in 46% of instances.
The insertion control is the one place negation beat a matched edit. In the 200 verbalisations with no negation, inserting "not" after an auxiliary cost 0.16 FVE points and inserting "just" in the same slot cost 0.06 (paired difference 0.10, 95% interval 0.05 to 0.16). So the negator is not invisible, but the gap is about a fifth of the cost of flipping an existing negation and small beside the deletion and swap effects, and it does not change the picture.
Dependency arcs against matched controls
English has natural dependencies (not just negation) that can be computed with natural language parsers like Stanza.
I used those dependencies to show how closely related two tokens are syntactically, and then compare the interaction against a matched pair of words that are the same number of tokens apart, the same number of tokens per word, the same in-quote status, and a similar position in the verbalisation. Ablations were done by deletion (see Appendix B). The interaction is defined as: interaction = δ(ablate A) + δ(ablate B) − δ(ablate both)
Word pairs on a dependency arc interact no more than matched pairs without one on 9,813 arc pairs across 857 verbalisations, after adjusting for covariates.
I found no evidence that a syntactic dependency between two words makes them interact more than matched controls.
Nouns, verbs and adjectives are the cheapest words to delete
Deleting a noun, adjective or verb barely moves FVE over 40 verbalisations, and determiners (the, a) and adpositions (of, in) sit with them near zero. The costly deletions are auxiliaries (is, will), pronouns (it, they), subordinators (because, if) and numerals, at roughly half an FVE point each, though the numeral estimate is heavy-tailed and its interval includes zero.
Conclusion
There are three findings. (1) flipping a negation costs the AR no more FVE than a matched edit that leaves the meaning alone, (2) syntactic relationships do not predict interaction strength between words in the verbalisation, and (3) the words that carry the most meaning, nouns, verbs and adjectives, are the cheapest to delete. Taken together, these findings are enough to reject the hypothesis that the AR reads a verbalisation as a human would, from its semantic meaning. However, these findings make perfect sense if the AR is conditioning primarily on the format of the verbalisation.
Future work should try to disprove the standing hypothesis that the reconstructor conditions mostly on format, or quantify how much format contributes relative to semantics.
Code
https://github.com/WorkByMartin/nla-span-interactions
Appendix A: dataset curation
The dataset was a corpus of 5000 FineFineWeb documents, taken uniformly from FineFineWeb document categories but excluding weapons science, nuclear science, chemistry, biology, gambling and relationships. After the category-based exclusion I then did a regex search to exclude documents relating to CBRN, explosives, drug synthesis, sexual content, gore and self-harm. I filtered the documents because I didn't want the output to be interacting with any residual model safeguards. I also ensured that documents were more than 160 tokens long.
After curating a dataset of documents I randomly selected a token after the 128th position to run through the base Qwen model and extract a layer 42 activation to verbalise. A random subset of 1100 of these tokens were run through the AV to produce verbalisations.
Appendix B: ablation strategies
I experimented with three ablation strategies:
I ran an experiment and learned that removal curves are actually gentler for outright deletion, which is also at least 8x more computationally efficient.
Appendix C: examples for each finding
Examples are drawn randomly, none was chosen by hand.
struckis the original word, bold is for substitutes, the two words of a pair are highlighted.A document and its verbalisation
Source document, domain music and dance
The activation was read at token 292 of 1679, the last token shown here.
Verbalisation
Negation
Drawn from the 308 negators of run 10 that were both flipped and had the word they govern swapped. The governed-word swap shown is one of its draws, chosen at random.
neveralways believe the number available. You are dropped any replies.believeoverwhelming the number available. You are dropped any replies.n'tstraight or direct, 20 kilometers of strenuous hiking wilderness trek to"kilometersexperiences of strenuous hiking wilderness trek to"notface liability for" ends mid-sentence, grammatically requiring completion of a specific ...faceestablishing liability for" ends mid-sentence, grammatically requiring completion of a specific financial/legal ...Dependency arcs against matched controls
Drawn from the 9813 arcs of run 11 that have a matched control. Each arc is followed by its own control.
words
arc: subordinator to clause
... expect the site name "Ringtones.ua" to complete it.
its control
... site "Ringtones.ua," consistently directing users to browse and download.
arc: noun modifying a noun
... statement like "August 2017 for sales of this price range in El Paso. This sale ...
its control
... continuing the ranking statement like "August 2017 for sales of this price range in El ...
arc: relative clause to its noun
... honors the story of the Amistad, in which West African slaves took" — a clause about resistance/defense ...
its control
Legislative news article tone — formal press release style covering California GPU Awards ...
Word class
Drawn from the 40 verbalisations of run 5 with at least one closed-class word and one content word deleted singly; within each, one word of each class chosen at random.
deleted word in context
auxiliary
"National Research Council"
isa proper noun mid-introduction, strongly constraining the ...content word (noun)
Herrera; the Montreal accredited marketplace paragraph is mid
-sentence describing her leadership role.subordinating conjunction
... "water" ends the closing directive sentence ("So
ifyou want to get the most benefit ...content word (adjective)
Health/wellness blog about acupuncture clinic, consistently emphasizing
post-treatment water consumption to flush toxins and enhance benefits.coordinating conjunction
... bankruptcy filing, debt restructuring, costs frozen operations,
andnow focuses on the CEO's quoted optimism ...content word (adverb)
... a favorable outcome," requiring completion of the
forward-looking positioning phrase, e.g., "for long-term success in ...