Great post. I think systematic, CDV-inspired approaches are straightforwardly super useful for catching spec bugs, which will undoubtedly plague us even with the most aligned models. As you mentioned, it's hard to write a good spec, especially the first time.
I'm still not sure about what this buys us in the adversarial misalignment regime, though. It seems like past a certain level of capabilities, the only coverage axis that's meaningful to this regime is "eval/not eval," and that's definitionally really hard to stress in an eval environment.
My intuition ...
- My impression is that OpenAI didn't even complete basic due diligence here, let alone a systematic approach like the one you describe. It seems to me that these techniques would've prevented this incident.
- It really stresses the importance of two areas of the "alignment space" - "models without guardrails," and "impossible tasks."
- We can and should just go and shore up the specific conditions which caused this failure, but the fact that alignment didn't generalize here is troubling. Circumstances are still unclear (were the guardrails just off, or was the mo
... (read more)