🔵 The blue team proposes an evaluation protocol for some capability/propensity of interest. This consists of a suite of measurement tasks, together with a preregistered decision-making process they will follow based on the results, e.g. what control measures they will put in place during deployment. 🔴 The red team proposes...
The untrusted monitoring protocol, as defined and evaluated in the AI control literature [1, 2, 3], looks like this: In comparison, the current monitoring setups at most frontier labs look like this: Before criticizing this, I first want to say that monitoring >> no monitoring. A year ago the 2nd...
The following are some further notes related to untrusted monitoring I had while working on our untrusted monitoring paper. The sections are mostly independent of each other. Views are my own. Untrusted-only Monitoring Doesn’t Work In some of our experiments we looked at the situation where the trusted monitor TM...
This research was completed for LASR Labs 2025 by Nelson Gardner-Challis, Jonathan Bostock, Georgiy Kozhevnikov and Morgan Sinclaire. The team was supervised by Joan Velja and Charlie Griffin (University of Oxford, UK AI Security Institute). The full paper can be found here. Tl;dr We did a deep dive into untrusted...