Exciting! Any thought to building data for cases where humans have preferences about the reasoning process they want a judge to use? E.g. cases where we want an AI to obey peoples' stated preferences over revealed ones, and vice versa, and cases where people have conflicts between how they think a problem "should" be solved versus both stated and revealed preferences, which are even more ambiguous?
I'm generally appreciative of agendas which attempt to improve AI-Human complementarity, since I think that seems like an essential ingredient for any functional scalable oversight mechanism.
Many others working in Scalable Oversight [...] Oversight from trusted humans adds an uncorrelated layer of defense.
I think this is kind of a strange way of distinguishing yourself from standard "scalable oversight" research... isn't the basic premise of the field to scale the oversight of humans? Perhaps what you're saying is that much scalable oversight research relies on LLMs anyway (e.g., to provide evidence for mechanisms which may be used to improve human oversight), and you would like to avoid doing that? If so, as you note, they do this to avoid cost and improve iteration speed; are you biting the bullet and just saying the benefits of directly studying oversight with actual human oversight is worth the costs in money/iteration speed?
Build and maintain a diverse and comprehensive leaderboard for judges. Given that you intend to do this, and you want to evaluate roughly frontier LLM-based systems in ways which expose their weaknesses, how are you going to build this eval? It seems like the only possible answer is "with a bunch of manual human labor". That makes sense, but in that case, how do you intend to operationalize things like "value alignment"? In particular, it seems like examples such as
cultural bias, unsafe interactions with children, and companion-AI dependence
require a certain level of subjective judgment on which reasonable people could disagree about the verdicts.
For any verification task, there are 2 components: the Judge (described above), and the specifications for what the judge is actually verifying. Even if we have a perfect judge, these specifications may be incorrect or incomplete, leading to similar reward hacking or misaligned behavior [...] and use a similar hill-climbing approach.
A related question to the above is: "hill-climbing" has the connotation of black-box optimizing against a fixed metric. How do you ensure that the metric you define isn't going to have the same specification problem?
Everyone understands and agrees on org priorities, and helps shape them together.
How do you plan to achieve this when the org scales beyond ~20 people? (Ironically, this is a sort of scalable oversight problem.)
We aim to differentially help safety [...] At the same time, we still need to ensure our methods transfer to capabilities-focused datasets
These two statements seem to be in direct opposition. I think what the post is saying is that more of the evaluation points you'll build will be alignment-oriented. But AIUI, your project is meant to help the development of scalable oversight techniques, and the leaderboard will be best optimized for by developing techniques which generalize to improved oversight on capabilities, which of course won't differentially help safety. You could believe that advancing both at the same speed is a net good, but that seems a subtler claim.
We've left Google DeepMind to start Sampura Research, a nonprofit building Human-AI Complementarity for Scalable Oversight - the research agenda we've pioneered for the past several years. We're backed by an initial $11M grant from Coefficient Giving ($7M for the first year, and another $4M pledged), and we're hiring founding technical staff in London! 🙂
The Initial Goal: Build Better Judges
For our first six months, we will prioritize building better judges via Human-AI Complementarity.
Definition of a Judge: A human, AI, or hybrid system that assesses correctness and alignment of an AI’s behavior in a conversation or agent trajectory.
Where improved judges can help:
How to improve judges:
Why focus on Human-AI Complementarity
Many others working in Scalable Oversight (i.e. the field of improving evaluation and data quality of AI, as AI rapidly becomes more capable) recognize the importance of involving humans. In practice though, most research substitutes LLMs for human judges - partly because it makes iteration cheaper and faster, and partly because many expect humans to play a shrinking role as AI matches or exceeds them on most tasks.
We share similar views, but think that Humans will have complementary strengths to AIs for many years to come:
Humans’ strengths are not being leveraged enough in the development of AI, and we see this as currently the lowest hanging fruit to improving AI Alignment. Achieving Complementarity won’t be easy though, especially as AI’s capabilities surpass humans’, which is why we've left DeepMind to pursue it at a scale we couldn't inside a larger lab.
After Judges, what’s next?
Deploy
Once we have better judges, we will shift focus to the remaining research needed to deploy them in frontier AI lab training, evaluations, and monitoring. Answering questions like:
As an intermediate step, we may initially stress-test judge deployment with third-party evaluators.
Expand Scope to Generative Tasks
For any verification task, there are 2 components: the Judge (described above), and the specifications for what the judge is actually verifying. Even if we have a perfect judge, these specifications may be incorrect or incomplete, leading to similar reward hacking or misaligned behavior.
As such, our research must expand beyond “classifier” judges towards generative tasks. These could be:
We’ll add datasets for these generative tasks to our leaderboard, and use a similar hill-climbing approach to improving these methods, ensuring the methods stay robust as models evolve.
How we work
We prioritize speed, real-world impact, and ownership. That shapes how we work:
FAQ