We've left Google DeepMind to start Sampura Research, a nonprofit building Human-AI Complementarity for Scalable Oversight - the research agenda we've pioneered for the past several years. We're backed by an initial $11M grant from Coefficient Giving ($7M for the first year, and another $4M pledged), and we're hiring founding technical staff in London! 🙂
The Initial Goal: Build Better Judges
For our first six months, we will prioritize building better judges via Human-AI Complementarity.
Definition of a Judge: A human, AI, or hybrid system that assesses correctness and alignment of an AI’s behavior in a conversation or agent trajectory.
Where improved judges can help:
Imperfect judges currently used in training lead to reward hacking, and AI learning the wrong goals. Examples:
Human decisions used in RLHF favor sycophancy (prioritizing agreeableness over factual accuracy), and LLMs inherit these priorities [1]
Coding RL environments can be completed by cheating, and LLMs exhibit some deceptive cheating behavior on software tasks [2]
Evaluation done with imperfect judges leads to misleading research results [3; 4]
Deployment-time Monitoring for alignment is bottlenecked on judge quality [5]
Accurate judges can improve the quality of all other data used in the AI development and deployment stack: SFT, RL-Environments, Red-teaming, Pre-training, etc.
How to improve judges:
Build and maintain a diverse and comprehensive leaderboard for judges.
Composed of >20 subdatasets, across areas like deception detection, cultural bias, and unsafe agent action.
In addition to static evaluations, we will also include evaluations where the judges are subject to direct optimization pressure from the models they are overseeing during RL
Hill-climb the leaderboard
Focus on methods that leverage the complementary strengths of Humans and AIs (i.e. Human-AI Complementarity). E.g.:
Improve AI’s confidence estimation, and route to humans when less confident
Train a router to best delegate tasks to either humans or AI
Break down a task into sub-tasks, and delegate subtasks using the routers above
Provide task-specific assistance UIs for humans (e.g. generated real-time)
Methods should generalize to all diverse datasets in our leaderboard.
Will benchmark and compare with all other judge methods (e.g. debate, LLM-as-a-judge).
Why focus on Human-AI Complementarity
Many others working in Scalable Oversight (i.e. the field of improving evaluation and data quality of AI, as AI rapidly becomes more capable) recognize the importance of involving humans. In practice though, most research substitutes LLMs for human judges - partly because it makes iteration cheaper and faster, and partly because many expect humans to play a shrinking role as AI matches or exceeds them on most tasks.
We share similar views, but think that Humans will have complementary strengths to AIs for many years to come:
Jaggedness: Even after we get superintelligent AIs that are more accurate than humans on average, we believe AI’s intelligence is and will continue to be jagged, and we can find places where humans can still add value.
Robustness: Humans and AIs might cover each other’s weaknesses against optimization pressure faced by judges in training, even if the human-AI team is not better overall (e.g. the comprehensiveness-hallucination trade-off described in McAleese et al., 2024).
Collusion: Future AI-only judges used in supervision might deceptively collude with the base model being trained or evaluated. This collusion is hard to detect, and gets harder as eval awareness and alignment faking grow more pronounced. Oversight from trusted humans adds an uncorrelated layer of defense.
Unseen Human Values: While AI systems might become good at verifying aligned behavior in out-of-distribution scenarios, this may be harder in unprecedented scenarios where e.g. technology creates a genuinely new question. In those situations, the relevant values may need to be formed through real human deliberation. And as humans' values drift and evolve, updated values will need to be understood from real humans.
Humans’ strengths are not being leveraged enough in the development of AI, and we see this as currently the lowest hanging fruit to improving AI Alignment. Achieving Complementarity won’t be easy though, especially as AI’s capabilities surpass humans’, which is why we've left DeepMind to pursue it at a scale we couldn't inside a larger lab.
After Judges, what’s next?
Deploy
Once we have better judges, we will shift focus to the remaining research needed to deploy them in frontier AI lab training, evaluations, and monitoring. Answering questions like:
Does our improved judge actually reduce reward hacking in RL?
How to best incorporate the slower and more expensive humans, finding the optimal trade off between cost, speed, and quality?
When and how to re-train our judge methods during RL, as the base model policy evolves to hack the judge?
Initially our “Humans” will be paid annotators. How does our research transfer to:
The human user of an agent, in the real-time monitoring of that agent?
AI researchers, who verify the quality of data and processes overall, and assess validity of alignment methods?
And more…
As an intermediate step, we may initially stress-test judge deployment with third-party evaluators.
Expand Scope to Generative Tasks
For any verification task, there are 2 components: the Judge (described above), and the specifications for what the judge is actually verifying. Even if we have a perfect judge, these specifications may be incorrect or incomplete, leading to similar reward hacking or misaligned behavior.
As such, our research must expand beyond “classifier” judges towards generative tasks. These could be:
High-level general and system instructions given to judges
Fine-grained per-example rubrics or test-cases used by judges to evaluate individual tasks
We’ll add datasets for these generative tasks to our leaderboard, and use a similar hill-climbing approach to improving these methods, ensuring the methods stay robust as models evolve.
How we work
We prioritize speed, real-world impact, and ownership. That shapes how we work:
Everyone understands and agrees on org priorities, and helps shape them together.
Collaborative and high-bandwidth. Working in person, discussing blockers and priorities in real time.
Focus on impact and solving the problem, not prestige. Our top priority is to develop improved alignment methods and deploy them into LLM companies and third-party evaluators. We will never optimize our research agenda for conference acceptance - though we'll always maintain a rigorous approach to research and analysis.
It’s incredibly important to us that we strengthen and inspire other work in Scalable Oversight, Human-AI Complementarity, and Evaluations. We’ll be publicly releasing and discussing all our work often via papers, blogs, code, datasets, and leaderboards.
Our work sits at the intersection of HCI and AI Safety, communities that historically have clashed. We welcome this union, and the diversity and interdisciplinary culture that come with it.
FAQ
Name: Sampura Research comes from the Sanskrit word सम्पूरक (sampūraka), which means “complementary” or "one who fills up, completes, or makes full", reflecting our focus on Human-AI Complementarity.
Prior work: We have been developing this research agenda since 2023 - first at DeepMind (paper), then across successive SPAR and MARS fellowships with cohorts of 10-25 mentees (first paper, second to come next month). Our nonprofit is a direct continuation, scaled up dramatically.
Dual use: The field of Scalable Oversight, and the pursuit of improving judges and data quality, is fundamentally dual-use: it may both accelerate capabilities of AI, and make it safer. We aim to differentially help safety by making the majority of this leaderboard alignment-focused, where improved judging more directly translates to improved alignment. By maintaining such a leaderboard and being actively involved in its datasets’ development, we also hope to bring together and strengthen the alignment dataset community. At the same time, we still need to ensure our methods transfer to capabilities-focused datasets, and will incorporate some in our leaderboard. This is because reward-hacking in capabilities tasks like general coding might lead to egregiously misaligned models [6].
Leaderboard makeup: AI Alignment is an incredibly diverse objective. Existential-risk-focused benchmarks often center around detecting deceptive or scheming behavior from rogue AI. But a judge asking "is this AI causing harm?" also has to catch issues like cultural bias, unsafe interactions with children, and companion-AI dependence. Research that focuses on a single-domain may lead to methods that don’t generalize, like our previous “Evidence-only” assistance method that only applied to the retrieval tasks we were studying. We aim to keep our leaderboard diverse and comprehensive, and continuously updated as datasets get saturated.
Why us: We've spent our careers shipping ambitious research at the frontier. Rishub spent seven years at Google DeepMind: four years on AlphaFold 2 and 3, then co-leading DeepMind's Scalable Oversight team that began this agenda. Josh co-led human data work at DeepMind, helping curate high quality human and AI data for use in model training and evaluation across a wide variety of capability and safety research.
Debate?:AI Safety via Debate is one of the most explored Scalable Oversight methods in the AI Safety community. Research on debate has focused primarily on understanding the right protocol to train the “debate assistance” for the judge. We are excited to test if these findings hold on our leaderboard when using Human judges, and how to best combine our advances in Human-AI Complementarity into Debate.
Acknowledgements: Thank you to David Africa, Konstantinos Voudouris, and Alex Adams for helpful feedback on this post.
We've left Google DeepMind to start Sampura Research, a nonprofit building Human-AI Complementarity for Scalable Oversight - the research agenda we've pioneered for the past several years. We're backed by an initial $11M grant from Coefficient Giving ($7M for the first year, and another $4M pledged), and we're hiring founding technical staff in London! 🙂
The Initial Goal: Build Better Judges
For our first six months, we will prioritize building better judges via Human-AI Complementarity.
Definition of a Judge: A human, AI, or hybrid system that assesses correctness and alignment of an AI’s behavior in a conversation or agent trajectory.
Where improved judges can help:
How to improve judges:
Why focus on Human-AI Complementarity
Many others working in Scalable Oversight (i.e. the field of improving evaluation and data quality of AI, as AI rapidly becomes more capable) recognize the importance of involving humans. In practice though, most research substitutes LLMs for human judges - partly because it makes iteration cheaper and faster, and partly because many expect humans to play a shrinking role as AI matches or exceeds them on most tasks.
We share similar views, but think that Humans will have complementary strengths to AIs for many years to come:
Humans’ strengths are not being leveraged enough in the development of AI, and we see this as currently the lowest hanging fruit to improving AI Alignment. Achieving Complementarity won’t be easy though, especially as AI’s capabilities surpass humans’, which is why we've left DeepMind to pursue it at a scale we couldn't inside a larger lab.
After Judges, what’s next?
Deploy
Once we have better judges, we will shift focus to the remaining research needed to deploy them in frontier AI lab training, evaluations, and monitoring. Answering questions like:
As an intermediate step, we may initially stress-test judge deployment with third-party evaluators.
Expand Scope to Generative Tasks
For any verification task, there are 2 components: the Judge (described above), and the specifications for what the judge is actually verifying. Even if we have a perfect judge, these specifications may be incorrect or incomplete, leading to similar reward hacking or misaligned behavior.
As such, our research must expand beyond “classifier” judges towards generative tasks. These could be:
We’ll add datasets for these generative tasks to our leaderboard, and use a similar hill-climbing approach to improving these methods, ensuring the methods stay robust as models evolve.
How we work
We prioritize speed, real-world impact, and ownership. That shapes how we work:
FAQ