Can Recursive Self-Report Probing Detect Emergent Misalignment?
In this post, I summarize the findings from my work, which I did as part of the BlueDot AI Safety Course. The full code is available here. Background Betley et al. (2025) showed that fine-tuning an LLM on insecure code not only learns to write just insecure code, but the...
Jul 258