Working on AI Safety, I spend a significant part of my time thinking about how frontier AI Safety research eventually gets translated into deployable safety systems. One thing I have constantly noticed is the transition from research to deployment changes the question we ask. In research, it is often sufficient...
tl;dr A probe can have excellent AUROC and yet fail as a safety signal. I audited 3 probes: a monitoring awareness probe leakage example, a refusal direction as positive control, and Apollo's deception probe as a published protocol case study. In the leakage case, the probe achieves AUROC 1.00, but...
The Prompt Is the Tell, Not the Reasoning Trace > Across 32,170 rollouts, eval-related prompt cues predicted refusal shifts more reliably than verbalized eval-awareness in model traces. If a system prompt tells Claude Opus 4.7 that its response is about to be reviewed by safety researchers, it becomes about 34...