The Model Organism Lottery: Model Organism Interpretability Strongly Depends on Training Methodology
by GabrielKS, therootof3, Raffaello Fornasiere, Nikita Menon, and StefanHex
TL;DR Current model organisms (MOs) for interpretability benchmarking are typically constructed via a dedicated, “post-hoc” SFT step. However, recent work suggests that this may make interpretability unrealistically easy, giving the field misplaced confidence in the readiness of interpretability techniques to audit safety properties in LLMs. We show that across activation...
Jul 2313