This post raises a methodological question with regard to how to make sense of models' behaviors in response to activation steering anchored on a concept.
Our discussion is limited to a particular kind of behaviors induced by activation steering, that of introspection studied by Lindsey (2026) (first accessed at here) and a range of investigations it inspires (see LW Post and citations in there). That said, the implication of our discussion applies to activation steering in general, mutatis mutandis.
Introspection refers to a model's "self-awareness" of its perturbed internal state. The self-awareness in question might be observable from the model's self-report, when prompted without giving away the existence or details of the perturbation. On the other hand, the perturbation itself is conducted by injecting a concept vector into the model's residual stream at a certain site. For model's self-report to be admissible as causally related to the concept injection, it must meet certain reasonable criteria, such as those given by Lindsey. One of those requires the report's content be specific to the injected concept.
Introspection thus illustrates a form of activation steering where both the steerer and the behavioral response are anchored on a concept. A closer examination of the construction of the concept vector in the experimental setup of introspection studies, however, raises questions about who else is steering, and what concept-anchored introspection really means.
Let be a model. In introspection studies, the concept vector of a word is typically obtained by the difference between two values:
= residual stream activation on the last token before reply to Tell me about;
= mean of for , being a set of words excluding .
We may call a contrast set. Thus the concept vector depends on a word and a contrast set, and is only one of its kind.
Suppose we observe at some rate introspection behaviors specific to some displayed by model , when steered by , e.g., with being the 100 baseline words in Lindsey (2026). The following questions arise:
Replace with any such that and are numerically dissimilar enough. Then does , steered by , display similar introspection behaviors?
If yes, what is the invariant that unifies and under the concept ?
If no, to what extent is the introspection upon steering by an artifact of ?
These questions are empirically answerable. My last attempt was foiled, as I could not replicate introspection on a small open-source model like Qwen3-4B-Instruct-2507. Nevertheless, I think for those who can access sufficient compute, these questions are worth exploring. Should (2) be the case, we need to be cautious when attributing causes to models' introspection or any activation-steering-induced behaviors: externalities, such as the contrast set for constructing concept vectors, may be lying in wait.
This post raises a methodological question with regard to how to make sense of models' behaviors in response to activation steering anchored on a concept.
Our discussion is limited to a particular kind of behaviors induced by activation steering, that of introspection studied by Lindsey (2026) (first accessed at here) and a range of investigations it inspires (see LW Post and citations in there). That said, the implication of our discussion applies to activation steering in general, mutatis mutandis.
Introspection refers to a model's "self-awareness" of its perturbed internal state. The self-awareness in question might be observable from the model's self-report, when prompted without giving away the existence or details of the perturbation. On the other hand, the perturbation itself is conducted by injecting a concept vector into the model's residual stream at a certain site. For model's self-report to be admissible as causally related to the concept injection, it must meet certain reasonable criteria, such as those given by Lindsey. One of those requires the report's content be specific to the injected concept.
Introspection thus illustrates a form of activation steering where both the steerer and the behavioral response are anchored on a concept. A closer examination of the construction of the concept vector in the experimental setup of introspection studies, however, raises questions about who else is steering, and what concept-anchored introspection really means.
Let be a model. In introspection studies, the concept vector of a word is typically obtained by the difference between two values:
We may call a contrast set. Thus the concept vector depends on a word and a contrast set, and is only one of its kind.
Suppose we observe at some rate introspection behaviors specific to some displayed by model , when steered by , e.g., with being the 100 baseline words in Lindsey (2026). The following questions arise:
These questions are empirically answerable. My last attempt was foiled, as I could not replicate introspection on a small open-source model like Qwen3-4B-Instruct-2507. Nevertheless, I think for those who can access sufficient compute, these questions are worth exploring. Should (2) be the case, we need to be cautious when attributing causes to models' introspection or any activation-steering-induced behaviors: externalities, such as the contrast set for constructing concept vectors, may be lying in wait.