cd /news/large-language-models/nine-judges-two-effective-votes-corr… · home topics large-language-models article
[ARTICLE · art-36785] src=machinelearning.apple.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

A new study finds that panels of nine large language model judges effectively provide only about two independent votes' worth of information, because the models make correlated errors on the same items. The panel's accuracy falls 8–22 percentage points short of independent voting, and the best single judge matches or outperforms the full panel. The findings suggest that adding more judges or using smarter aggregation cannot substitute for genuinely independent evaluation.

read2 min views24 publishedJun 23, 2026
Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels
Image: Apple ML Research

content type paperpublished June 2026 Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

AuthorsGuneet Kohli

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

AuthorsGuneet Kohli

LLM-as-a-judge panels aggregate votes from multiple models, with the expectation that diverse models yield more reliable evaluations. We develop a framework to measure the true informational value of such panels and quantify how far their reliability falls short of the independent-voting ideal. Testing a panel of 9 frontier LLMs from 7 model families on three natural language inference datasets (each with 100 human annotations per item), we find that the 9 judges effectively provide only about 2 independent votes’ worth of information. Roughly three-quarters of the panel’s nominal independence is lost because the models make the same mistakes on the same items. The consequences are stark: the panel’s actual accuracy falls 8–22 percentage points short of what independent voting would achieve, and the best single judge matches or outperforms the full panel across all conditions. Neither adding more judges nor using smarter aggregation algorithms helps — established methods close at most 11% of this gap, even with access to the correct answers. We quantify these findings using the Kish effective sample size (n_eff) and a Condorcet null model, and show the deficit is robust across prompt variants, temperatures, chain-of-thought reasoning, and a pairwise preference task (RewardBench). The bottleneck is correlated judges, not the aggregation algorithm, implying that scaling up panels cannot substitute for genuinely independent evaluation.

Identifying Controversial Pairs in Item-to-Item Recommendations

November 3, 2023research area Methods and Algorithmsconference RecSys *Equal Contributors

Recommendation systems in large-scale online marketplaces are essential to aiding users in discovering new content. However, state-of-the-art systems for item-to-item recommendation tasks are often based on a shallow level of contextual relevance, which can make the system insufficient for tasks where item relationships are more nuanced. Contextually relevant item pairs can sometimes have problematic relationships that are…

Consistent Collaborative Filtering via Tensor Decomposition

August 16, 2023research area Knowledge Bases and Search, research area Methods and Algorithms Collaborative filtering is the de facto standard for analyzing users’ activities and building recommendation systems for items. In this work we develop Sliced Anti-symmetric Decomposition (SAD), a new model for collaborative filtering based on implicit feedback. In contrast to traditional techniques where a latent representation of users (user vectors) and items (item vectors) are estimated, SAD introduces one additional latent vector to each…

── more in #large-language-models 4 stories · sorted by recency
── more on @guneet kohli 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nine-judges-two-effe…] indexed:0 read:2min 2026-06-23 ·