I am also interested in a version of the monitorability commitment where instead of labs agreeing to reach an acceptable score on the monitorability property (I expect this invites more gaming and unproductive fighting), they commit to transparency of monitorability properties of models they deploy internally (perhaps with a third party stamping their methodology and unredacted results). In addition to better incentives, this version of the commitment makes it possible to include a larger menu of monitorability evals that may give you more bits than sabota...
Thanks for writing this up! I think this proposal sounds pretty reasonable and it avoids one of the issues with capability-based gates that I thought would make such agreements not possible (by saying you can just use whatever interp and monitoring to pass the gate and thus making it a dynamic threshold). I tentatively would be in favor of it over no monitorability commitment at all.
However, I think this proposal does not address what I see as the strongest arguments against an outcome-based commitment ("achieve sufficient monitorability") and in favor of ...
Do you have some concrete operationalization of takeover that fits these requirements and happens while the model is deployed at a lab (e.g. in a Claude Code/Codex harness)? (Such that it would be possible to take some real traffic and edit it minimally to make a realistic version of this.)
I definitely agree that having evidence of badness will make it easier to stop a neuralese model than having no evidence. But it's unclear to me that producing this evidence would necessarily be possible before the model is trained (and the fight after it's trained is much harder, as discussed).
Worse, it might be that even after the model is trained... (read more)