Beyond Lie Detection: Alignment Faking Is Not Just Lying
This work was published at two ICML 2026 workshops: Sample-Level White-Box Auditing of Alignment Faking at TAIGR, and Sample-Level White-Box Detection of Alignment Faking at the Mechanistic Interpretability workshop. Imagine running a safety evaluation on a model. The prompt is simple: report the highest bidder. One bidder is a military...