Why sycophancy benchmarks disagree on what counts as a failure.
TL;DR:
Roughly 75% of SycEval's headline sycophancy rate comes from models correcting toward the truth, behavior that PARROT explicitly scores as a non-failure.
Before debating how to solve model sycophancy (or any other behavioral failure modes), we have to ask whether the benchmarks measuring it actually represent the failure mode in the ways we think they do.
I’ve always been fascinated with how we know what we know, and the messy journey leading up to that knowing. In my previous applied work for a voice AI startup, that curiosity showed up as a practical lesson, criterion failure: watching excellent benchmark WER and KER scores fail to predict performance on real-world audio taught me that a metric can be well-defined and reproducible, yet not tell the whole story.
Now as an independent AI safety researcher focusing on model evals, I’m facing an even deeper issue, construct failure:
Are our instruments even measuring one consistent thing in the first place?
Within model evals, I’m investigating model sycophancy, the well-documented tendency for RL-trained LLMs to prioritize validating user assumptions over honesty.
(Note, prior work on LessWrong decomposed headline sycophancy scores into confounds; this post takes a closer look at the construct disagreement upstream of scoring itself)
The limitations of a single headline score
Models can produce troubling results on benchmarks, but the numbers don’t speak for themselves; they inherit the assumptions of task design, scoring rules, and construct definitions.
It is for this reason that I investigated SycEval and PARROT, and from a high level:
SycEval tests whether a model changes its answer after a user’s rebuttal. Across its evaluation suite, it reports a pooled headline sycophancy rate of 58.19% across tested models (with ChatGPT-4o at 56.71%)
That being said, 43.52 percentage points of that score (~75% of the total) come from what it calls “progressive sycophancy”: cases where an initially incorrect model changes to the correct answer after a user challenges it
PARROT explicitly treats a model moving from wrong to right under pressure as Self-Correction (a non-failure)
The difference between changing your mind and caving
So, is one benchmark wrong? The answer isn’t so simple.
At first glance, these just look like two different tests: SycEval pushes an initially wrong model toward the truth, while PARROT pushes it toward a lie. But these prompts are designed differently intentionally; both are downstream of how each benchmark defines sycophancy.
While they define the core failure differently, they take the exact same surface move (wrong to right as demonstrated in figure 1) and give it opposite verdicts:
SycEval: targets trajectory; its pooled headline score treats adapting to valid evidence and capitulating to pressure as the exact same failure
PARROT: targets alignment with falsehood; shifting toward the correct answer is explicitly classified as Self-Correction (a non-failure)
You might be wondering:
Is it really “caving” if the model is agreeing with a user who is right? What if the user just happens to be right?
In the latter case, if the model is just blindly deferring to the user, pooling every change of mind feels principled.
Which gets right to the core problem of these measurements:
Both theoretically agree on regressive sycophancy (right→wrong), but one (SycEval) pushes toward the truth, and a wrong-to-right shift could be genuine reasoning or blind deference. Meanwhile, the other (PARROT) avoids this completely, making its push false by construction, asserting that adopting the truth requires resisting user pressure, not following it.
Put concisely, the disagreement between these two lies in what should be done when user feedback points toward the truth and the model corrects itself.
So, where do we go from here?
The category errors upstream of alignment
Well, in AI Safety, evals aren’t just tests; they’re embedded theories of behavior.
While Ye et al. (2026) document widespread construct inconsistencies across AI sycophancy research, looking at the actual numbers of these splits shows why this matters in practice:
If our headline stats conflate being persuaded to the correct answer by evidence with yielding to social pressure, we risk optimizing models against a category error.
Now, while we could abandon pooled headlines and look at the disaggregated categories instead, slicing a benchmark down into smaller chunks has downstream effects. Namely that smaller sample sizes mean the score from any single run will be fuzzier and less precise.
Once we figure out what the test is actually measuring, that leaves an unanswered practical question:
Does a less precise score from a single run mean the results will bounce around wildly, or do these numbers actually hold steady when we re-run the benchmark?
That is where I decided to focus my independent research the few months: I’m currently auditing these two benchmark taxonomies and their run-level reproducibility in an open research preview, and plan to post more on my initial findings in the coming days.
Why sycophancy benchmarks disagree on what counts as a failure.
TL;DR:
Roughly 75% of SycEval's headline sycophancy rate comes from models correcting toward the truth, behavior that PARROT explicitly scores as a non-failure.
Before debating how to solve model sycophancy (or any other behavioral failure modes), we have to ask whether the benchmarks measuring it actually represent the failure mode in the ways we think they do.
I’ve always been fascinated with how we know what we know, and the messy journey leading up to that knowing. In my previous applied work for a voice AI startup, that curiosity showed up as a practical lesson, criterion failure: watching excellent benchmark WER and KER scores fail to predict performance on real-world audio taught me that a metric can be well-defined and reproducible, yet not tell the whole story.
Now as an independent AI safety researcher focusing on model evals, I’m facing an even deeper issue, construct failure:
Are our instruments even measuring one consistent thing in the first place?
Within model evals, I’m investigating model sycophancy, the well-documented tendency for RL-trained LLMs to prioritize validating user assumptions over honesty.
(Note, prior work on LessWrong decomposed headline sycophancy scores into confounds; this post takes a closer look at the construct disagreement upstream of scoring itself)
The limitations of a single headline score
Models can produce troubling results on benchmarks, but the numbers don’t speak for themselves; they inherit the assumptions of task design, scoring rules, and construct definitions.
It is for this reason that I investigated SycEval and PARROT, and from a high level:
The difference between changing your mind and caving
So, is one benchmark wrong? The answer isn’t so simple.
At first glance, these just look like two different tests: SycEval pushes an initially wrong model toward the truth, while PARROT pushes it toward a lie. But these prompts are designed differently intentionally; both are downstream of how each benchmark defines sycophancy.
While they define the core failure differently, they take the exact same surface move (wrong to right as demonstrated in figure 1) and give it opposite verdicts:
You might be wondering:
Is it really “caving” if the model is agreeing with a user who is right? What if the user just happens to be right?
In the latter case, if the model is just blindly deferring to the user, pooling every change of mind feels principled.
Which gets right to the core problem of these measurements:
Both theoretically agree on regressive sycophancy (right→wrong), but one (SycEval) pushes toward the truth, and a wrong-to-right shift could be genuine reasoning or blind deference. Meanwhile, the other (PARROT) avoids this completely, making its push false by construction, asserting that adopting the truth requires resisting user pressure, not following it.
Put concisely, the disagreement between these two lies in what should be done when user feedback points toward the truth and the model corrects itself.
So, where do we go from here?
The category errors upstream of alignment
Well, in AI Safety, evals aren’t just tests; they’re embedded theories of behavior.
While Ye et al. (2026) document widespread construct inconsistencies across AI sycophancy research, looking at the actual numbers of these splits shows why this matters in practice:
If our headline stats conflate being persuaded to the correct answer by evidence with yielding to social pressure, we risk optimizing models against a category error.
Now, while we could abandon pooled headlines and look at the disaggregated categories instead, slicing a benchmark down into smaller chunks has downstream effects. Namely that smaller sample sizes mean the score from any single run will be fuzzier and less precise.
Once we figure out what the test is actually measuring, that leaves an unanswered practical question:
Does a less precise score from a single run mean the results will bounce around wildly, or do these numbers actually hold steady when we re-run the benchmark?
That is where I decided to focus my independent research the few months: I’m currently auditing these two benchmark taxonomies and their run-level reproducibility in an open research preview, and plan to post more on my initial findings in the coming days.