NLAs miss some internalized behaviors
by Andrii Shportko, Ivan Arcuschin, Georg Lange, and Bryce Meyer
> TL;DR: NLAs detect side tasks the model is instructed to do, but mostly miss the same behavior if it is fine-tuned in. After the Activation Oracles and then Natural Language Autoencoders were released, there was a hope in the mech interp community that the scalable unsupervised white box monitoring...
Oct 87