This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
I thought I had found an induction head.
The attention pattern looked familiar, and one head received a substantially higher induction score than the others. The obvious interpretation was that this head was responsible for the model's copying behavior.
Then I removed it. The model's accuracy did not change.
I removed each individual head in turn. Still no meaningful change. But when I removed the attention sublayer as a whole, accuracy fell from 100% to approximately 2%.
This is a small result from a toy Transformer, so I don't think it supports a broad claim about induction heads in language models. It does support a narrower claim: on this task, the attention score was useful for finding a candidate pattern, but it did not identify a uniquely necessary component.
That distinction seems easy to blur because the visual evidence is persuasive. The head really did attend in an induction-like pattern. The problem was not that the measurement was random. The problem was that I treated a descriptive measurement as if it had already answered a causal question.
What I measured
The model was a small two-layer Transformer trained on a repeated-sequence task. The task was constructed so that, when a sequence appeared again, the model could use what followed the earlier occurrence to predict what should come next.
I recorded the attention matrices and computed an induction-head score for every head. Informally, the score measures how much attention a head places on the token that followed a previous occurrence of the current token. The exact score is useful as a screening statistic, but it does not encode necessity. A head can receive a high score because it is implementing the relevant computation, because it is exploiting a positional regularity, or because another part of the network has made the relevant information easy to copy.
The strongest-looking attention map was therefore evidence about what the head was doing, not yet evidence about what the model required.
Figure 1. The attention pattern that initially suggested an induction-like computation.
The score ranked one head above the others:
Figure 2. The induction score ranked one head above the others. This ranking generated a useful hypothesis, but did not establish causal responsibility.
If I had stopped here, I would probably have written that the top-ranked head was the induction head. That would have been too strong.
The intervention changed the story
The model reached 100% accuracy on the repeated-sequence task. I then compared its behavior after interventions on individual heads and on the attention sublayer.
The individual-head interventions had little or no effect. The result was not specific to the highest-scoring head: no single head appeared necessary under this ablation procedure. In contrast, removing the attention sublayer reduced accuracy to approximately 2%.
Figure 3. The induction score ranked one head above the others. This ranking generated a useful hypothesis, but did not establish causal responsibility.
If I had stopped here, I would probably have written that the top-ranked head was the induction head. That would have been too strong.
The intervention changed the story
The model reached 100% accuracy on the repeated-sequence task. I then compared its behavior after interventions on individual heads and on the attention sublayer.
The individual-head interventions had little or no effect. The result was not specific to the highest-scoring head: no single head appeared necessary under this ablation procedure. In contrast, removing the attention sublayer reduced accuracy to approximately 2%.
Figure 4. Attention concentration by head and pairwise similarity between heads. These measurements help generate hypotheses, but do not by themselves establish causal responsibility.
The distinction between a useful diagnostic and a causal test is the reason I am emphasizing the failed prediction rather than presenting this as a successful induction-head identification. The plot was not misleading in the narrow sense. My interpretation of what the plot established was.
Questions I would like help with
The most useful criticism would be about the experiment design rather than the software.
What controls would you use to separate genuine content-based induction from a fixed-offset positional shortcut? Is ablating the entire attention sublayer a sufficiently informative control here, or would you use resampling, activation patching, or a more targeted intervention to distinguish redundancy from a task-level shortcut?
I am also interested in how people report reproducibility for these experiments. Exact bitwise replay is not always realistic across hardware and library versions. For an interpretability claim, when should a replay count as successful: identical tensors, numerically close metrics, preserved rankings, or preservation of the behavioral conclusion?
The code and figures are in the TransInterp repository. The archived project is available at Zenodo, DOI 10.5281/zenodo.22691733.This is not evidence that induction-head scores are generally unreliable, and it is not evidence about modern language models. It is a small example of a more general failure mode: a correlated internal pattern can be real, reproducible, and still fail to identify the component that is causally necessary for the behavior.
I thought I had found an induction head.
The attention pattern looked familiar, and one head received a substantially higher induction score than the others. The obvious interpretation was that this head was responsible for the model's copying behavior.
Then I removed it. The model's accuracy did not change.
I removed each individual head in turn. Still no meaningful change. But when I removed the attention sublayer as a whole, accuracy fell from 100% to approximately 2%.
This is a small result from a toy Transformer, so I don't think it supports a broad claim about induction heads in language models. It does support a narrower claim: on this task, the attention score was useful for finding a candidate pattern, but it did not identify a uniquely necessary component.
That distinction seems easy to blur because the visual evidence is persuasive. The head really did attend in an induction-like pattern. The problem was not that the measurement was random. The problem was that I treated a descriptive measurement as if it had already answered a causal question.
What I measured
The model was a small two-layer Transformer trained on a repeated-sequence task. The task was constructed so that, when a sequence appeared again, the model could use what followed the earlier occurrence to predict what should come next.
I recorded the attention matrices and computed an induction-head score for every head. Informally, the score measures how much attention a head places on the token that followed a previous occurrence of the current token. The exact score is useful as a screening statistic, but it does not encode necessity. A head can receive a high score because it is implementing the relevant computation, because it is exploiting a positional regularity, or because another part of the network has made the relevant information easy to copy.
The strongest-looking attention map was therefore evidence about what the head was doing, not yet evidence about what the model required.
Figure 1. The attention pattern that initially suggested an induction-like computation.
The score ranked one head above the others:
Figure 2. The induction score ranked one head above the others. This ranking generated a useful hypothesis, but did not establish causal responsibility.
If I had stopped here, I would probably have written that the top-ranked head was the induction head. That would have been too strong.
The intervention changed the story
The model reached 100% accuracy on the repeated-sequence task. I then compared its behavior after interventions on individual heads and on the attention sublayer.
The individual-head interventions had little or no effect. The result was not specific to the highest-scoring head: no single head appeared necessary under this ablation procedure. In contrast, removing the attention sublayer reduced accuracy to approximately 2%.
Figure 3. The induction score ranked one head above the others. This ranking generated a useful hypothesis, but did not establish causal responsibility.
If I had stopped here, I would probably have written that the top-ranked head was the induction head. That would have been too strong.
The intervention changed the story
The model reached 100% accuracy on the repeated-sequence task. I then compared its behavior after interventions on individual heads and on the attention sublayer.
The individual-head interventions had little or no effect. The result was not specific to the highest-scoring head: no single head appeared necessary under this ablation procedure. In contrast, removing the attention sublayer reduced accuracy to approximately 2%.
Figure 4. Attention concentration by head and pairwise similarity between heads. These measurements help generate hypotheses, but do not by themselves establish causal responsibility.
The distinction between a useful diagnostic and a causal test is the reason I am emphasizing the failed prediction rather than presenting this as a successful induction-head identification. The plot was not misleading in the narrow sense. My interpretation of what the plot established was.
Questions I would like help with
The most useful criticism would be about the experiment design rather than the software.
What controls would you use to separate genuine content-based induction from a fixed-offset positional shortcut? Is ablating the entire attention sublayer a sufficiently informative control here, or would you use resampling, activation patching, or a more targeted intervention to distinguish redundancy from a task-level shortcut?
I am also interested in how people report reproducibility for these experiments. Exact bitwise replay is not always realistic across hardware and library versions. For an interpretability claim, when should a replay count as successful: identical tensors, numerically close metrics, preserved rankings, or preservation of the behavioral conclusion?
The code and figures are in the TransInterp repository. The archived project is available at Zenodo, DOI 10.5281/zenodo.22691733.This is not evidence that induction-head scores are generally unreliable, and it is not evidence about modern language models. It is a small example of a more general failure mode: a correlated internal pattern can be real, reproducible, and still fail to identify the component that is causally necessary for the behavior.