Misaligned models rate themselves as more harmful, and realignment reverses it
This post summarizes our paper Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment (arXiv:2602.14777), co-authored with Anietta Weckauff and Thilo Hagendorff at the University of Stuttgart. It contains examples of harmful model outputs. TL;DR: Fine-tuning an aligned model on a narrow subversive task, such as answering...
Aug 287