cd /news/large-language-models/agreement-overstates-evidence-error-… · home topics large-language-models article
[ARTICLE · art-137816] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus

A study posted to arXiv (2609.22512v1) finds that LLM judges' errors are correlated, with an average pairwise error correlation of 0.21 across a main bank of ten judges, meaning those ten judges supply only about as much statistical information as 3.5 independent judges. The paper reports that ignoring shared errors leads to the conclusion that one system is significantly better in up to 28% of comparisons, with the dependency strongest among high-accuracy frontier judges including those from different providers. The authors recommend using a small set of trusted examples to estimate judge accuracy and identify shared mistakes, then accounting for those shared errors and selecting the voting method before applying it to new data.

by read1 min views1 publishedSep 23, 2026

arXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error correlation across both open-weight and frontier LLM judges. In our main bank of ten judges, the average pairwise correlation between judge errors is 0.21. As a result, the ten judges only provide roughly as much statistical information as 3.5 independent judges. The dependency is even stronger among the high-accuracy frontier judges we evaluate, including judges from different providers. In up to 28% of our comparisons, ignoring shared errors leads to the conclusion that one system is significantly better, while accounting for them does not. We also find that the pattern of errors matters. Errors shared by most judges and errors concentrated among a smaller group affect consensus differently and favor different voting methods. Measuring the overall amount of correlation alone is therefore insufficient. Our results suggest a simple approach: use a small set of trusted examples to estimate judge accuracy and identify shared mistakes. These shared errors should then be considered when analyzing the results, and the voting method should be chosen using trusted examples before it is applied to new data.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/agreement-overstates…] indexed:0 read:1min 2026-09-23 ·