TL;DR. In this work we study obstacles to the faithful automation of alignment research.[1] We see this as a scalable oversight problem. There are plenty of examples of how models fail at this, and as models become more capable our ability to notice these failures will diminish: even the best human checkers won’t be able to tell if the model was well elicited, thorough checking will become too costly, and models could tailor their responses to their judges. We draw on empirical examples from Geoguessr and auto-alignment runs from Arcadia’s internal research to make general claims about obstacles to the oversight of fuzzy alignment-related tasks. Narrowing our attention to one prominent scalable oversight method, we find that whilst debate shows promise on typical capabilities benchmarks (aligning with recent work) it fails on tasks involving judgment calls akin to those arising in automated alignment research.
We’d like to thank David Africa, Andrew Draganov, Rory Greig, Joshua Jacob, Rishub Jain, Zac Kenton, Francis Rhys Ward and Lennie Wells for helpful feedback on this post.
Introduction
Existing empirical work on debate [1, 2, 3, 4, 5, 6] has almost exclusively focused on objective, verifiable domains, seeking to mitigate misalignment caused by supervision mistakes. In this work we focus on extending debate to automated alignment research, where verifiable rewards are absent or costly (and arguably differentially so when compared to automated ML capabilities research). There are plenty of high-level discussions on the nature of harder-to-verify tasks of this type.[2]
Auto-alignment research decisions will lie on a spectrum of ‘fuzziness’; from sharp measurable steps to judgement-laden ones. In this note we want to clarify what makes the fuzzy end hard for scalable oversight, grounding each point empirically: we compare non-fuzzy domains with known ground truth (math, coding) against a fuzzy setting with known ground truth (Geoguessr), alongside motivating a