Mechanisms of Introspective Awareness
Uzay Macar and Li Yang are co-first authors. This work was advised by Jack Lindsey and Emmanuel Ameisen, with contributions from Atticus Wang and Peter Wallich, as part of the Anthropic Fellows Program. Paper: https://arxiv.org/abs/2603.21396. Code: https://github.com/safety-research/introspection-mechanisms TL;DR * We investigate the mechanisms underlying "introspective awareness" (as shown in Lindsey (2025) for Claude Opus 4 and 4.1) in open-weights models[1]. * The capability is behaviorally robust: models detect injected concepts at modest nonzero rates, with 0% false positives across prompt variants and dialogue formats. * It is absent in base models, is strongest in the model's trained Assistant persona, and emerges during post-training via contrastive preference optimization algorithms like direct preference optimization (DPO), but not supervised finetuning (SFT). * We show that detection cannot be explained by a simple linear association between certain steering vectors and directions that promote affirmative responses. * Identification of injected concepts relies on largely distinct later-layer mechanisms that only weakly overlap with those involved in detection. * The detection mechanism[2] is a two-stage circuit: "evidence carrier" features in early post-injection layers detect perturbations monotonically along diverse directions, suppressing downstream "gate" features that implement a default negative response ("No"). This circuit is absent in base models and robust to refusal ablation. * Introspective capability is substantially underelicited by default. Ablating refusal directions improves detection by 53% and a trained bias vector improves detection by 75% on held-out concepts, both without meaningfully increasing false positives. Figure 1: A steering vector representing some concept is injected into the residual stream (left). "Evidence carriers" in early post-injection layers suppress later-layer "gate" features that promote a default n