I recently attended Generality Labs' Inspect Evals Data Viz Hackathon, and spent the day using Inspect AI and its offspring, Scout, with a simple goal in mind - generate a new plot of a new or existing benchmark. Many thanks to the organisers and to my team mates, Jeff Mohl and Valerie Griffiths (the Overfit and Overcaffeinated team), for a fantastic time, learning some new tricks on using the Inspect suite.
Here I'm presenting the two (!) plots we got in the span of ~ 6 hours (more like 4 hours as it took us a while to agree on what we actually want to spend the day on - arguably a harder task than its execution).
[... 2 hours later... ] Our initial idea was to run a few models on a subset of Humanity's Last Exam (HLE), and undertake an extensive failure mode analysis to understand the current gaps and where different capabilities x harnesses fail or succeed. We did this using Inspect Scout - a framework for an LLM-as-judge that analyses the models' outputs and assigns a dominant feature that lead to the answer failing or winning. Then, we rerun a subset of tasks and analysed how changing the harness affects the prevalence of failures - given the time constraints, we only modified the harness to include access to web search.
We started with 200 randomly sampled HLE tasks, relatively balanced across domains, and ran three GPT models on the same fixed subset. This gave us 600 model attempts in total, of which 525* were judged as failures and could then be analysed with Inspect Scout.
*Post-hackathon, we noticed that this subset was substantially harder for the models we tested than the full HLE benchmark: the three models had a combined pass rate of ~12%, compared with ~40% reported for HLE overall. So the 525 failures should not be interpreted as representative of the failure rate on HLE as a whole.
The first thing we wanted to know was what actually goes wrong when a model fails HLE. Intuitively, is it mainly lack of capability, inconsistency of the answer format, wrong reasoning (especially in maths heavy problems), API calls failures, lack of compute, or something else entirely?
We used Scout to run another model over the failed attempts and classify what had gone wrong. We ended up with these categories:
judge failure
incorrect reference answer
tool/ environment failure
resource limit
answer-format failure
ambiguous/ defective prompt
reasoning failure
knowledge failure
undetermined/ ambiguous
The important distinction was to avoid reducing the analysis to a more elaborate way of saying that “the model got it wrong.” For example, if a model produces an incorrect answer, we should not immediately classify this as a knowledge failure. The model may have had the relevant knowledge but applied it incorrectly, or it may have produced a correct answer that was nevertheless rejected by the judge.
However, there is another problem here: HLE has known documented issues with some of its reference answers. This means a model can sometimes be marked wrong when the problem is actually with the answer key. This is particularly interesting for automated evaluations, where the reference answer is often implicitly treated as ground truth.
One thing we'd like to ideally do is automatically detect these cases with scanners. Rather than assuming that every disagreement with the reference answer is a model failure, a scanner could look at the question, reference answer, supporting rationale and model response and flag cases where the target answer itself appears questionable. Flagging discrepancies would be a great tool in evals, especially when the ground truth answer isn't straightforward to compute without the experts' knowledge on demand.
How much of the failure belongs to the model?
We've done some EDA on the distributions of failure modes across different domains. However, it became apparent quite quickly that we would have needed more time to sit down with the data and digest it - even a simple correlation matrix revealed that the failure mode themes we identified were confounded, e.g. a knowledge failure led to a reasoning failure, especially in Chem-Bio tasks.
We also compared these results against the newly released HLE-Verified, which addresses some of the identified prompt and answer-key issues. This let us assess how often our scanners correctly identified failures associated with issues that were subsequently fixed in HLE-Verified, shown in Figure 1 below.
Figure 1. Performance across models. Tasks with incorrect reference answers can give a misleading picture of a model's true capabilities.
The other thing that stood out was confidence. The models generally reported confidence levels of around 80-90%, despite having substantially lower accuracy. In other words, on HLE they were often highly confident in answers that turned out to be incorrect.
Integrating web-search tools into the harness
Once we had the failure-mode analysis, we wanted to see whether changing the harness actually changed how the models failed. In a post-hackathon life, we would look more into the failure modes to help us inform what kind of harness changes are more sensible and interesting to explore here.
As proof-of-concept, we then ran the same tasks with and without a Tavily web-search tool, with the results shown in Figure 2.
Figure 2. Influence of the harness over the accuracy for the three models we used.*The gpt-5-mini web run broke part way through because of a Tavily error, so we left it out rather than use a partial run.
The web-search tool inclusion showed an uplift of the strongest model we tested by 4.5 percentage points, but only helped gpt-4o by 0.5 pp. Therefore, for more capable models, the access to certain tools can be critical towards their evaluation workflow - a model failing because it lacks a relevant fact is in a different situation from one that has the ability to retrieve that information but does not. Similarly, providing the relevant information may shift a failure from a knowledge problem to a reasoning problem. This is the kind of distinction we wanted to capture with the failure-mode analysis, rather than simply measuring whether overall accuracy increased.
Conclusions
As we've shown in this small hackathon project, if a model gets say 20% on a benchmark, that could mean lots of different things, from lack of information, reasoning or compute, to a reflection on the quality of the grading methodology. The best part is, these can easily confound, which is particularly important for agent evaluations, where there are many more ways for something to go wrong than "the model produced the wrong answer". Breaking these apart is crucial to better development and deploying of new and capable models per use case.
The hackathon was only a few hours, so there are a lot of obvious next steps to make this work more robust: run the whole HLE on multiple models, validate the failure categories against human judgements, investigate whether different judges produce different classifications (we've used gemini-3.6 here as the current most calibrated one for HLE tasks), and look properly at whether adding tools shifts models between different failure modes - with a failure mode analysis looking at the tool call types as well.
Personal reflection
I've worked on evals before, and this experience just reaffirmed that it can be dangerous for benchmarks to rely on one unique metric. I'm seeing more and more shifts in the evals community to challenge how we interpret these traces, and even treat them as datasets themselves. This was a great exercise showcasing how much you can do with some compute, a couple of hours and a well-known benchmark!
[... 2 hours later... ] Our initial idea was to run a few models on a subset of Humanity's Last Exam (HLE), and undertake an extensive failure mode analysis to understand the current gaps and where different capabilities x harnesses fail or succeed. We did this using Inspect Scout - a framework for an LLM-as-judge that analyses the models' outputs and assigns a dominant feature that lead to the answer failing or winning. Then, we rerun a subset of tasks and analysed how changing the harness affects the prevalence of failures - given the time constraints, we only modified the harness to include access to web search.
All code is available here.
Initial failure mode analysis
We started with 200 randomly sampled HLE tasks, relatively balanced across domains, and ran three GPT models on the same fixed subset. This gave us 600 model attempts in total, of which 525* were judged as failures and could then be analysed with Inspect Scout.
*Post-hackathon, we noticed that this subset was substantially harder for the models we tested than the full HLE benchmark: the three models had a combined pass rate of ~12%, compared with ~40% reported for HLE overall. So the 525 failures should not be interpreted as representative of the failure rate on HLE as a whole.
The first thing we wanted to know was what actually goes wrong when a model fails HLE. Intuitively, is it mainly lack of capability, inconsistency of the answer format, wrong reasoning (especially in maths heavy problems), API calls failures, lack of compute, or something else entirely?
We used Scout to run another model over the failed attempts and classify what had gone wrong. We ended up with these categories:
The important distinction was to avoid reducing the analysis to a more elaborate way of saying that “the model got it wrong.” For example, if a model produces an incorrect answer, we should not immediately classify this as a knowledge failure. The model may have had the relevant knowledge but applied it incorrectly, or it may have produced a correct answer that was nevertheless rejected by the judge.
However, there is another problem here: HLE has known documented issues with some of its reference answers. This means a model can sometimes be marked wrong when the problem is actually with the answer key. This is particularly interesting for automated evaluations, where the reference answer is often implicitly treated as ground truth.
One thing we'd like to ideally do is automatically detect these cases with scanners. Rather than assuming that every disagreement with the reference answer is a model failure, a scanner could look at the question, reference answer, supporting rationale and model response and flag cases where the target answer itself appears questionable. Flagging discrepancies would be a great tool in evals, especially when the ground truth answer isn't straightforward to compute without the experts' knowledge on demand.
How much of the failure belongs to the model?
We've done some EDA on the distributions of failure modes across different domains. However, it became apparent quite quickly that we would have needed more time to sit down with the data and digest it - even a simple correlation matrix revealed that the failure mode themes we identified were confounded, e.g. a knowledge failure led to a reasoning failure, especially in Chem-Bio tasks.
We also compared these results against the newly released HLE-Verified, which addresses some of the identified prompt and answer-key issues. This let us assess how often our scanners correctly identified failures associated with issues that were subsequently fixed in HLE-Verified, shown in Figure 1 below.
Figure 1. Performance across models. Tasks with incorrect reference answers can give a misleading picture of a model's true capabilities.
The other thing that stood out was confidence. The models generally reported confidence levels of around 80-90%, despite having substantially lower accuracy. In other words, on HLE they were often highly confident in answers that turned out to be incorrect.
Integrating web-search tools into the harness
Once we had the failure-mode analysis, we wanted to see whether changing the harness actually changed how the models failed. In a post-hackathon life, we would look more into the failure modes to help us inform what kind of harness changes are more sensible and interesting to explore here.
As proof-of-concept, we then ran the same tasks with and without a Tavily web-search tool, with the results shown in Figure 2.
Figure 2. Influence of the harness over the accuracy for the three models we used.*The gpt-5-mini web run broke part way through because of a Tavily error, so we left it out rather than use a partial run.
The web-search tool inclusion showed an uplift of the strongest model we tested by 4.5 percentage points, but only helped gpt-4o by 0.5 pp. Therefore, for more capable models, the access to certain tools can be critical towards their evaluation workflow - a model failing because it lacks a relevant fact is in a different situation from one that has the ability to retrieve that information but does not. Similarly, providing the relevant information may shift a failure from a knowledge problem to a reasoning problem. This is the kind of distinction we wanted to capture with the failure-mode analysis, rather than simply measuring whether overall accuracy increased.
Conclusions
As we've shown in this small hackathon project, if a model gets say 20% on a benchmark, that could mean lots of different things, from lack of information, reasoning or compute, to a reflection on the quality of the grading methodology. The best part is, these can easily confound, which is particularly important for agent evaluations, where there are many more ways for something to go wrong than "the model produced the wrong answer". Breaking these apart is crucial to better development and deploying of new and capable models per use case.
The hackathon was only a few hours, so there are a lot of obvious next steps to make this work more robust: run the whole HLE on multiple models, validate the failure categories against human judgements, investigate whether different judges produce different classifications (we've used gemini-3.6 here as the current most calibrated one for HLE tasks), and look properly at whether adding tools shifts models between different failure modes - with a failure mode analysis looking at the tool call types as well.
Personal reflection
I've worked on evals before, and this experience just reaffirmed that it can be dangerous for benchmarks to rely on one unique metric. I'm seeing more and more shifts in the evals community to challenge how we interpret these traces, and even treat them as datasets themselves. This was a great exercise showcasing how much you can do with some compute, a couple of hours and a well-known benchmark!