About a "science of held out monitors"
In my recent post, I said that I’d be very excited about building a “science of held out monitors”. What might that look like? Why am I excited about this? And what open questions seem most important right now?
In order: the current SOTA held-out monitoring is labs promising not to train on chain-of-thought, doing so a little (as a treat), and things being more or less fine. However, we currently don’t have amazing ways to answer questions like “How much pressure can you put on the chain of thought before it stops being useful?” “Training only on outputs still affects the chain-of-thought; so to what extent does regular training teach CoT obfuscation?” “Are there alignment proxies we can train on that don’t degrade CoT monitorability?” Ideally, a science of held-out-monitoring would let us answer questions like these, giving a toolbox of methods to estimate worst-case and average-case behavior of a model trained to evade a misalignment detector.
I’m excited about this for a few reasons. The first, and most important, is that it would be really bad if labs thought they had some held-out evals, but it turns out that their training tended to degrade the usefulness of those held-out evals significantly. It seems pretty bad if all of our monitors are silently worse than we expect. On the flip side, I think being able to reliably predict when this would happen could open up promising alignment techniques. Currently, labs try to avoid training on CoT because we all agree it would make our deployment-time CoT monitors invalid. However, training on CoT might be really useful, and training against misalignment detectors seems like one of the best ways we have to elicit aligned behavior. This is especially true in short timelines worlds, where we may have to do a lot of unprincipled, training-based RL-for-safety work. In those worlds, better understanding how often training against the CoT messes up your setup, or whether