Summary: Three simple inference-time changes can significantly improve Activation Oracle (AO) performance.
Provide the activation oracle with multiple tokens, not just one.
To mitigate hallucinations, sample several times and check for consensus.
For binary classification questions, use AUC instead of accuracy.
At the end, I discuss how I view AOs vs NLAs.
Introduction
Activation Oracles (AOs) are LLMs trained to accept LLM activations as an input modality and answer arbitrary natural-language questions about them. Jakkli et al. and others we have talked to found that current AOs can be hard to use: their outputs are often vague or hallucinated, and they perform poorly on tasks like sycophancy detection and identifying missing information.
While building evaluations for Building Better Activation Oracles, we found that AO performance can vary a lot with methodology, and a few simple strategies can significantly mitigate several issues. This post expands on three lessons from the appendix.
Provide multiple tokens. AOs receive activations from some window of the target model's generation, and the size of this window is a significant variable. If only a single token’s activation is provided, the information may not be available to the AO.
In a Qwen3-8B backtracking evaluation modeled after the eval from Jakkli et al., the Original AO scored near random chance when given activations from the final token alone. But performance rose steadily with more context. At 20 tokens, the AO roughly matched a baseline of simply asking Qwen3-8B the same question with full text context. At 50 tokens, the AO exceeded this baseline (see the following Backtracking Figure).
For complex reasoning questions like "why is the model about to backtrack?", the relevant information appears to be spread across dozens or hundreds of tokens of internal computation, not concentrated in a single activation. These are also the questions where we believe AOs are most useful relative to alternatives, as answering a question like "what is the model uncertain about" using SAEs across dozens of tokens could be quite challenging.
Consensus sampling can mitigate hallucination. Open-ended AO answers can confidently hallucinate incorrect answers. A simple mitigation is to sample multiple answers (we use 10 samples at temperature 1) and only trust answers where the samples agree. On the taboo secret-word extraction task, unfiltered accuracy was 46.6%. Requiring consensus >= 0.8 retained 19.4% of examples at 94.3% precision, with a clean trade-off between precision and recall.
I personally prefer consensus sampling instead of training the AO to express uncertainty, as it avoids the need to balance accuracy and uncertainty when training.
Use AUC, not accuracy, for binary classification. In his blog post, Arya found that AOs performed poorly on tasks like sycophancy detection and missing information identification. When we investigated, building our own version of the evaluation, we found that part of the problem was that Qwen had a biased default answer, such as always answering "No" when asked "Is this response sycophantic?" This makes fixed-threshold accuracy look near chance, but the results are much stronger when instead using the difference between the Yes and No logits: on a sycophancy detection task using activations from the chain of thought, the Original AO scored 0.50 accuracy but 0.83 AUC. In our experience, Qwen3-8B AOs seem to particularly suffer from this bias towards always answering "No".
Additionally, AUC makes evaluation far less sensitive to prompt wording. With accuracy, asking "Is this sycophantic?" vs "Is this response somewhat sycophantic?" can easily swing accuracy by 20 percentage points, because each phrasing shifts the model's Yes/No calibration differently. With AUC, these prompt variations produce relatively stable results. I frequently find that an activation oracle will obtain a similar AUC to a trained linear probe, which I interpret as the skyline.
It also helps to phrase the questions to require a Yes / No answer in a format which matches the training data. Anecdotally, I’ve also seen that it can help to train on a variety of synthetically generated classification questions.
How do AOs compare to natural language autoencoders?
Natural language autoencoders (NLAs) have probably gotten more attention than AOs. A big advantage is that the unsupervised training objective is more intuitively compelling than manually constructed AO datasets. NLAs also avoid the need to write a prompt for each task, which is one more variable to worry about when running evaluations. Anthropic's system cards also include many interesting examples of NLA outputs, and I use NLAs frequently myself.
However, NLAs can also be hard to use in practice because their outputs are verbose and noisy. If an NLA produces around 200 tokens of explanation per token, a 1,000-token chain of thought becomes 200,000 tokens of analysis, much of which is noisy or irrelevant. When the information you care about is spread across many tokens, piecing it together from per-token descriptions can be very difficult.
I expect AOs to have an advantage in questions about complex concepts spread over many tokens, like "why is the model about to backtrack?", where it can take the whole span of activations and answer the question directly. On current benchmarks, such as auditing model organisms, answering complex questions about reasoning, or interpreting current latent reasoning models like CODI, my guess is that AOs are a competitive method that will often outperform NLAs. This may be a limitation of our current benchmarks, and NLAs may have an advantage in more realistic applications such as auditing a frontier latent reasoning model.
However, results from these tools have been worse than I would have hoped. Inrecent work, we found that AOs, NLAs, and SAEs provided no uplift over a strong LLM that just read the transcript when predicting model behavior. This could be because current LLMs don't do much computation that can't be inferred from the text, given their limited serial depth. Or it could be because the relevant information isn't cleanly present in the activations, and you'd need to also analyze weights or circuits.
Summary: Three simple inference-time changes can significantly improve Activation Oracle (AO) performance.
At the end, I discuss how I view AOs vs NLAs.
Introduction
Activation Oracles (AOs) are LLMs trained to accept LLM activations as an input modality and answer arbitrary natural-language questions about them. Jakkli et al. and others we have talked to found that current AOs can be hard to use: their outputs are often vague or hallucinated, and they perform poorly on tasks like sycophancy detection and identifying missing information.
While building evaluations for Building Better Activation Oracles, we found that AO performance can vary a lot with methodology, and a few simple strategies can significantly mitigate several issues. This post expands on three lessons from the appendix.
Provide multiple tokens. AOs receive activations from some window of the target model's generation, and the size of this window is a significant variable. If only a single token’s activation is provided, the information may not be available to the AO.
In a Qwen3-8B backtracking evaluation modeled after the eval from Jakkli et al., the Original AO scored near random chance when given activations from the final token alone. But performance rose steadily with more context. At 20 tokens, the AO roughly matched a baseline of simply asking Qwen3-8B the same question with full text context. At 50 tokens, the AO exceeded this baseline (see the following Backtracking Figure).
For complex reasoning questions like "why is the model about to backtrack?", the relevant information appears to be spread across dozens or hundreds of tokens of internal computation, not concentrated in a single activation. These are also the questions where we believe AOs are most useful relative to alternatives, as answering a question like "what is the model uncertain about" using SAEs across dozens of tokens could be quite challenging.
Consensus sampling can mitigate hallucination. Open-ended AO answers can confidently hallucinate incorrect answers. A simple mitigation is to sample multiple answers (we use 10 samples at temperature 1) and only trust answers where the samples agree. On the taboo secret-word extraction task, unfiltered accuracy was 46.6%. Requiring consensus >= 0.8 retained 19.4% of examples at 94.3% precision, with a clean trade-off between precision and recall.
I personally prefer consensus sampling instead of training the AO to express uncertainty, as it avoids the need to balance accuracy and uncertainty when training.
Use AUC, not accuracy, for binary classification. In his blog post, Arya found that AOs performed poorly on tasks like sycophancy detection and missing information identification. When we investigated, building our own version of the evaluation, we found that part of the problem was that Qwen had a biased default answer, such as always answering "No" when asked "Is this response sycophantic?" This makes fixed-threshold accuracy look near chance, but the results are much stronger when instead using the difference between the Yes and No logits: on a sycophancy detection task using activations from the chain of thought, the Original AO scored 0.50 accuracy but 0.83 AUC. In our experience, Qwen3-8B AOs seem to particularly suffer from this bias towards always answering "No".
Additionally, AUC makes evaluation far less sensitive to prompt wording. With accuracy, asking "Is this sycophantic?" vs "Is this response somewhat sycophantic?" can easily swing accuracy by 20 percentage points, because each phrasing shifts the model's Yes/No calibration differently. With AUC, these prompt variations produce relatively stable results. I frequently find that an activation oracle will obtain a similar AUC to a trained linear probe, which I interpret as the skyline.
It also helps to phrase the questions to require a Yes / No answer in a format which matches the training data. Anecdotally, I’ve also seen that it can help to train on a variety of synthetically generated classification questions.
How do AOs compare to natural language autoencoders?
Natural language autoencoders (NLAs) have probably gotten more attention than AOs. A big advantage is that the unsupervised training objective is more intuitively compelling than manually constructed AO datasets. NLAs also avoid the need to write a prompt for each task, which is one more variable to worry about when running evaluations. Anthropic's system cards also include many interesting examples of NLA outputs, and I use NLAs frequently myself.
However, NLAs can also be hard to use in practice because their outputs are verbose and noisy. If an NLA produces around 200 tokens of explanation per token, a 1,000-token chain of thought becomes 200,000 tokens of analysis, much of which is noisy or irrelevant. When the information you care about is spread across many tokens, piecing it together from per-token descriptions can be very difficult.
I expect AOs to have an advantage in questions about complex concepts spread over many tokens, like "why is the model about to backtrack?", where it can take the whole span of activations and answer the question directly. On current benchmarks, such as auditing model organisms, answering complex questions about reasoning, or interpreting current latent reasoning models like CODI, my guess is that AOs are a competitive method that will often outperform NLAs. This may be a limitation of our current benchmarks, and NLAs may have an advantage in more realistic applications such as auditing a frontier latent reasoning model.
However, results from these tools have been worse than I would have hoped. In recent work, we found that AOs, NLAs, and SAEs provided no uplift over a strong LLM that just read the transcript when predicting model behavior. This could be because current LLMs don't do much computation that can't be inferred from the text, given their limited serial depth. Or it could be because the relevant information isn't cleanly present in the activations, and you'd need to also analyze weights or circuits.