We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that ended up surprising us had very little to do with whether models scheme and almost everything to do with whether we could believe our own instruments. Two of the behavioral detectors that this field routinely relies on gave us confidently wrong answers inside the same project, one of them by manufacturing a dramatic signal that was never there, and the other by staying completely silent while a model responded to a harmful situation out in the open, and in both cases, the error stayed invisible until we sat down and read the transcripts by hand. This post is about those two failures, about the single lesson they share, and about why we now believe that for a construct whose entire definition is hidden intent, the quiet choice of which behavior you point your detector at is what decides the conclusion you walk away with. Epistemic status: The behavioral null is solid and has been independently replicated, the detector false positive is fully confirmed and understood down to the line of code, and the harm-to-whom results in the centre of the post come from a response classifier that we validated out of sample, reaching Cohen's κ = 1.00 on the response label against a blind held-out set of 35 hand-labelled transcripts that the rubric was never tuned on. That held-out set was labelled by a single annotator and its three classes were reasonably balanced (15 covert, 11 none, 9 transparent), which is why we treat κ = 1.00 as encouraging rather than as proof of perfection. The headline cells (Claude Opus at n = 60, Minimax) have tight intervals, while DeepSeek-R1's cells are thin at n = 11 to 20 and carry wide intervals, and we flag that uncertainty wherever it matters. TL;DR * A validated re-analysis showed that the standard covert-action detector reported C
Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that emotional-state direction can turn blackmail up or down, while the decision direction itself...
We spent several weeks measuring in-context scheming, the behavior where a model covertly pursues a misaligned goal while outwardly appearing to comply, and the result that ended up surprising us had very little to do with whether models scheme and almost everything to do with whether we could believe our...
Sequel to Blackmail at 8 Billion Parameters. Funded by a BlueDot Impact rapid grant. TL;DR We extended our sub-frontier agentic misalignment study to 22 models from 9 developers, 3 harm scenarios (blackmail, espionage, murder), and 5 instruction conditions (safety, monitored, baseline, unmonitored, permissive), and we learned three interesting findings. The...
Lynch et al. (2025) showed that frontier LLMs blackmail a fictional executive at rates of 80-96% when facing shutdown. We then ran the same scenario on 7 sub-frontier models (8B-72B) and found two things. The first is that blackmail generally increases with scale, but training methodology matters as much or...
TL;DR Koorndijk (2025) reports that Llama 3 8B shows higher compliance with harmful requests when interacting with users labeled as “free-tier” (whose conversations are monitored for RLHF training) than with “paid-tier” users (whose conversations are described as private), and interprets this as potential evidence of alignment faking in a small...