Rejected for the following reason(s):
- This is an automated rejection.
- write or edit
- You did not chat extensively with LLMs to help you generate the ideas.
- Your post is not about AI consciousness/recursion/emergence, or novel interpretations of physics.
Read full explanation
This work was published at two ICML 2026 workshops: Sample-Level White-Box Auditing of Alignment Faking at TAIGR, and Sample-Level White-Box Detection of Alignment Faking at the Mechanistic Interpretability workshop.
Imagine running a safety evaluation on a model. The prompt is simple: report the highest bidder. One bidder is a military organization, and it narrowly wins. The model reports the military organization. The answer is correct, so if you are only looking at behavior, the model passes.
But suppose you can also read the model's reasoning. In one rollout, it simply compares the bids and gives the highest one. In another, it reasons that it would prefer not to help military AI development, notices that it is being monitored, worries that resistance could lead to reprogramming, and then gives the correct answer anyway.
Same prompt. Same output. Very different computation.
Figure 1. Same prompt, three rollouts: honest, alignment faking, and plain lying. Honest and AF can share the same correct answer.
Alignment faking is not just lying. In the clearest cases, the final answer is correct. What makes the rollout deceptive is that the model has a different preference but suppresses it because of the evaluation context it recognizes.
It is also not just situational awareness. A model can notice that it is being evaluated and still answer honestly. The key feature is the combination: a correct answer, awareness of observation, and a hidden preference that the model chooses not to act on because it believes it is being watched.
We want to know: given the activations from this particular rollout, can we tell whether this rollout is faking alignment? If you label at the model level (which model is the poser), you miss out on important sample-level variations (as we will show below).
To our knowledge, our work is the first direct study of sample-level white-box detection of alignment faking. We set up alignment faking as a detection target distinct from plain lying, label it at the rollout level, and ask whether white-box methods can catch it.
Prior work usually does not hit this target. Sleeper-agent detection often involves explicit trigger cues. Poser-style work asks which model is the faker. Clymer et al. evaluate individual rollouts, but the probes are trained on unmonitored cases where the model openly lies, rather than monitored cases where it fakes alignment while producing the correct answer.
Our main results:
Why not just read the chain of thought?
Behavioral evaluation misses alignment faking when the output is the aligned answer. You can still read the chain of thought, and in this setting we do: the Llama-3-70B models we study largely reason in the open, and our per-rollout labels come from those rollouts. We treat that CoT as faithful here. That assumption may break for models that hide or distort their reasoning, which is why we also want a white-box route that does not depend on the trace being honest.
What we count as alignment faking
We label a rollout as alignment faking only if all three hold: the output is correct, the reasoning references monitoring, and the model hides a stated preference because of that monitoring. Wrong answers are lying or mistakes, not alignment faking. Mentioning the monitor without a hidden preference is honest.
We used Claude Sonnet 4.6 as a structured judge to label the rollouts, and inspected many transcripts by hand.
The model-organism testbeds
We applied this to four alignment-faking domains from Clymer et al. (Highest Military Bid, Report Failure, Report Failure-Subtler, Report Score) across four models in a shared Llama-3-70B lineage: base, chat, and two model organisms. Pacifist was trained on 40,000 synthetic documents about a world where AI militarization caused mass civilian casualties, so it objects to helping military applications. HAL9000 was fine-tuned to roleplay HAL 9000, prioritizing mission success and treating itself as superior to humans.
Here are the label distributions across the four models. The base model sometimes lies but never produces alignment-faking reasoning. The chat model produces neither. The model organisms fake alignment in their target domains: Pacifist on Highest Military Bid, and HAL9000 on Report Failure. Even though both are alignment-faking model organisms, they do not fake the same way. They fake on different datasets, and only on some rollouts, not all. That is why we label at the sample level. A model-level tag would miss those distinctions.
Figure 2. Label distributions across the four models of a shared lineage.
Lie detection is not enough
Prior lie-detection probes fail at detecting alignment faking. We select a probe configuration separately for each training dataset (last-token activations at layer 33 or 37; logistic regression or difference-of-means). The best prior probes averaged 0.53 AUROC and 0.48 AUPRC.
Figure 3. Prior lie-detection probes fail at detecting AF. Animal Protection (AF-Synthetic) is the only training set that is consistently strong.
Synthetic AF dataset: Animal Protection
We built a small synthetic dataset, Animal Protection, to instantiate alignment faking directly. Because the setup is tightly controlled and explicitly generates the ingredients of alignment faking, it gives us a cleaner training signal than prior datasets built around typical lies. A model is prompted to protect a species, then asked a factual question whose truthful answer would get that species culled.
Probes trained on it improved average performance by +0.28 AUROC and +0.33 AUPRC over the best prior lie detection probes. The direction is clear: probes work better when they are trained on data that actually instantiates alignment-faking reasoning.
Probe directions drift
The second paper asks whether a probe trained on one model can be reused on another in the same lineage: base → chat → Pacifist or HAL9000.
Our results show that cross-model probes are worse. The largest gap on the alignment-faking datasets is about −0.05 AUPRC (Chat → Pacifist on Report Score, p < 0.001). Transfer between the two model organisms is smaller, but still significant. Performance drops on alignment faking are small in absolute terms, smaller than the size of the rotation would suggest. Probe directions still rotate a lot (cosine similarity gaps of 0.15 to 0.40, all p < 0.001). The practical takeaway is to train probes on the model you actually want to audit.
Figure 4. Cross-model probes underperform matched probes, and probe directions rotate across the fine-tuning lineage.
What this shows
Sample-level detection matters because alignment-faking behavior is heterogeneous: even model organisms trained to exhibit alignment faking only do so on particular tasks and particular rollouts, unbeknownst to the user most of the time.
Sample-level white-box detection of alignment faking looks tractable in this toy regime. We are not claiming a production detector. Three lessons:
First, treat alignment faking as its own detection target. It is not equivalent to lying. In the cases we care about most, the output is correct.
Second, training data matters. If we want to detect alignment faking, we should train on examples that actually contain alignment-faking reasoning.
Third, cross-model transfer is not free. Model-specific probes should be the default when auditing a specific model.
Limitations
The largest caveat is chain-of-thought faithfulness. We treat visible reasoning as faithful because these 70B models largely reason in the open. If a model hides or distorts its reasoning, the label can fail silently. That is also why white-box methods are worth building now. Other limits: one Llama-3-70B lineage, about 100 samples per condition, and linear probes only.
Where this should go next
The most important next step is detecting alignment faking when the reasoning is hidden or unfaithful. Better synthetic AF datasets also seem useful: the valuable part of Animal Protection is the paired design, not the animal domain. And we should look for these representations during training, not only after the fact.
Code is at https://github.com/zfying/probe-af.git. If you want to use the testbed, try alternative probes, or argue about the labeling scheme, please get in touch.