[Epistemic status] After getting somewhat familiar with the technical side of evals, I started reading about the conceptual side. By now I’m keen to get feedback on my thoughts. Please err on the side of pointing out where I’m missing something or just got it wrong. I would be grateful if you’d point me to relevant posts, research, or whatever I should explore next.
When we evaluate whether a model can x, what we mean by “can” differs depending on whether we’re in a capability or safety context.
In the first case, we want to know whether the model can solve the task under “normal” usage.
In the second, we care more about extreme cases: can we make the model solve a desired or undesired task x, for example, by crafting specific prompts, providing additional tools, etc.?
I'm looking for an understanding of what parts of current evals address the second task. Should I be looking at the structure of the evaluation process, the evaluation target, or something else?
If I oversimplify, my current line of thoughts goes like this.
Benchmarking: As I see it, for capability evaluations, benchmarks are designed to tell us how well a model performs over some reasonably defined distribution of tasks. In safety evaluations, again, as I see it, the asymmetry between positive and negative results makes such a design much less applicable. If I understand the What AI evaluations for preventing catastrophic risk can and cannot do correctly, it argues almost the same idea. The authors call it the difference between lower and upper bounds on capabilities.
If the model succeeds, we know it can do at least that much. It gives us a lower bound on its capability, but failure provides much weaker evidence that the capability is absent, because we generally cannot tell whether it has been optimally elicited.
Making a benchmark larger and more diverse may reduce some uncertainty, but it will never solve this basic upper-bound problem. A benchmark that tests a request like “give me a recipe for a chemical weapon” under one or a few elicitation strategies therefore seems to be useful for turning on the red lights, which is indeed safety-relevant. But, generally, do I get it right that we need an approach that is different by design to be able to reliably evaluate most of the properties we are concerned about in conceptual AI safety?
If so, what might this design look like?
Uplift studies and red teaming seem closer to measuring near-worst-case performance. In uplift studies, we ask how much access to an LLM increases the probability that a person can complete a task x, or at least make progress on it. In red teaming, we deliberately try to get the model to do x.
However, these methods seem sufficient for estimating near-worst-case performance only under the assumption that there is no way to use the LLM more effectively to solve the task x. Otherwise we’re still stuck in the lower-upper bound problem. But I don’t think this assumption stands in the real world, because people’s capabilities differ along many dimensions. Domain expertise is a factor we often account for, but many others may matter as well. How motivated a human is to complete the task and how much creativity they would apply to doing so, how much expertise they have in adjacent domains, what set of languages they can use, and so on.
This leads me to two questions:
Does it make sense to think about principles for designing eval pipelines that are explicitly built to push measurements above a lower bound, given that measuring the true upper bound is impossible by construction? I mean trying to build a safety-specific eval pipeline where the safety-specific part is in the measurement design itself, rather than in the evaluation questions.
Given the number of technical people becoming interested in safety and the relatively low barrier to entry for pet projects around evals, does it make sense to coordinate efforts to cover as broadly as possible the space of potentially safety-relevant properties with lower-bound measurements? My intuition says this information could be useful for building a more complete picture of the technology. Do you think in 2026, this is valuable enough to be worth community effort?
[Epistemic status] After getting somewhat familiar with the technical side of evals, I started reading about the conceptual side. By now I’m keen to get feedback on my thoughts. Please err on the side of pointing out where I’m missing something or just got it wrong. I would be grateful if you’d point me to relevant posts, research, or whatever I should explore next.
When we evaluate whether a model can x, what we mean by “can” differs depending on whether we’re in a capability or safety context.
I'm looking for an understanding of what parts of current evals address the second task. Should I be looking at the structure of the evaluation process, the evaluation target, or something else?
If I oversimplify, my current line of thoughts goes like this.
If the model succeeds, we know it can do at least that much. It gives us a lower bound on its capability, but failure provides much weaker evidence that the capability is absent, because we generally cannot tell whether it has been optimally elicited.
Making a benchmark larger and more diverse may reduce some uncertainty, but it will never solve this basic upper-bound problem. A benchmark that tests a request like “give me a recipe for a chemical weapon” under one or a few elicitation strategies therefore seems to be useful for turning on the red lights, which is indeed safety-relevant. But, generally, do I get it right that we need an approach that is different by design to be able to reliably evaluate most of the properties we are concerned about in conceptual AI safety?
If so, what might this design look like?
However, these methods seem sufficient for estimating near-worst-case performance only under the assumption that there is no way to use the LLM more effectively to solve the task x. Otherwise we’re still stuck in the lower-upper bound problem. But I don’t think this assumption stands in the real world, because people’s capabilities differ along many dimensions. Domain expertise is a factor we often account for, but many others may matter as well. How motivated a human is to complete the task and how much creativity they would apply to doing so, how much expertise they have in adjacent domains, what set of languages they can use, and so on.
This leads me to two questions: