This is an automated rejection. No LLM generated, assisted/co-written, or edited work.
Read full explanation
Some AI claims are too hard for a human (or another AI) to verify directly, and debate tries to solve this by combining evidence from competing agents. But if that evidence is even slightly correlated, adding more of it eventually stops helping; I found that the worst-case failure has a specific form that lets us put a computable upper bound on that risk.
A helpful analogy is 3 friends independently guessing a coin flip, each right 90% of the time. Majority vote gets you to ~97%, of course the friends being other AIs and you being the human who can’t verify the claim directly. Debate structures a game where competing agents have to defend/attack a claim, allowing a human judge to decide which argument holds up best under cross-examination.
There's always a catch though: what if they aren’t truly independent? If pieces of evidence share a “blind spot”, meaning they fail for the exact same underlying reason, voting tends to stop helping. The trap is that a shared blind spot can be invisible to standard audits. Each piece of evidence could look perfectly reliable when evaluated on its own, so checking them one at a time catches nothing. You get an illusion of agreement where the sheer volume of evidence looks like overwhelming proof, even though it can all systematically fail at once.
This matters for debate specifically because the prover chooses what evidence to present. If a dishonest prover controls exactly the thing that could secretly be correlated, they can exploit that correlation to construct a hard-to-falsify lie. With real correlation, adding more evidence helps less and less and flattens at a fixed floor; even ρ=0.01 caps you at ~100 effective judges no matter how many you actually use. Given only what a judge could realistically measure (average error rate + correlation strength), the worst case has a specific shape: rare but severe shared failure. That gives us an upper bound on how bad the vulnerability can get.
The paper's theorem isn't wrong; the math is right under its assumptions. And I definitely didn't invent this panic; the authors and the UK AISI have already been sweating over it. This takes the vague “what if they correlate?” nightmare, makes it concrete, and characterizes the worst case. If we track a prover's historical record, we can estimate their average error and correlation strength and calculate a risk ceiling even if we have no clue what their exact blind spot is. It's obviously a partial answer instead of a magic fix, yet it's nice to know.
Some AI claims are too hard for a human (or another AI) to verify directly, and debate tries to solve this by combining evidence from competing agents. But if that evidence is even slightly correlated, adding more of it eventually stops helping; I found that the worst-case failure has a specific form that lets us put a computable upper bound on that risk.
A helpful analogy is 3 friends independently guessing a coin flip, each right 90% of the time. Majority vote gets you to ~97%, of course the friends being other AIs and you being the human who can’t verify the claim directly. Debate structures a game where competing agents have to defend/attack a claim, allowing a human judge to decide which argument holds up best under cross-examination.
There's always a catch though: what if they aren’t truly independent? If pieces of evidence share a “blind spot”, meaning they fail for the exact same underlying reason, voting tends to stop helping. The trap is that a shared blind spot can be invisible to standard audits. Each piece of evidence could look perfectly reliable when evaluated on its own, so checking them one at a time catches nothing. You get an illusion of agreement where the sheer volume of evidence looks like overwhelming proof, even though it can all systematically fail at once.
This matters for debate specifically because the prover chooses what evidence to present. If a dishonest prover controls exactly the thing that could secretly be correlated, they can exploit that correlation to construct a hard-to-falsify lie. With real correlation, adding more evidence helps less and less and flattens at a fixed floor; even ρ=0.01 caps you at ~100 effective judges no matter how many you actually use. Given only what a judge could realistically measure (average error rate + correlation strength), the worst case has a specific shape: rare but severe shared failure. That gives us an upper bound on how bad the vulnerability can get.
The paper's theorem isn't wrong; the math is right under its assumptions. And I definitely didn't invent this panic; the authors and the UK AISI have already been sweating over it. This takes the vague “what if they correlate?” nightmare, makes it concrete, and characterizes the worst case. If we track a prover's historical record, we can estimate their average error and correlation strength and calculate a risk ceiling even if we have no clue what their exact blind spot is. It's obviously a partial answer instead of a magic fix, yet it's nice to know.