RoPoLL: Robust Panel of LLM Judges

wpnews.pro

cd /news/large-language-models/ropoll-robust-panel-of-llm-judges · home › topics › large-language-models › article

[ARTICLE · art-45932] src=arxiv.org ↗ pub=2026-07-01T04:00Z topic=large-language-models verified=true sentiment=↑ positive

RoPoLL: Robust Panel of LLM Judges

Researchers introduced RoPoLL, a robust panel of LLM judges that replaces standard consensus aggregation with a geometric median estimator, achieving unbounded bias resistance under contamination. In tests across 13 open-weight judges and three benchmarks, RoPoLL outperformed standard PoLL by up to 19% on biased corruption and enabled a 38B committee to beat a 675B single judge by 1.31x under 30% corruption.

read1 min views1 publishedJul 1, 2026

arXiv:2606.30931v1 Announce Type: new Abstract: The LLM Jury, a Panel of LLM Evaluators (PoLL) reporting consensus scores, has become a practical alternative to single-judge LLM evaluation, yet its statistical behavior remains poorly understood. We formalize the LLM Jury under the Huber contamination model and show that PoLL incurs unbounded bias under any positive contamination, regardless of jury size, whenever a single judge fails in a biased, LLM-typical way (mode collapse, sycophancy, safety refusal). Framing jury consensus as classical robust mean estimation, we propose RoPoLL (Robust Panel of LLM-as-Judge), which preserves the PoLL panel but replaces the aggregation function with a robust mean estimator, instantiated with the geometric median (GM): tuning-free, with the optimal finite-sample breakdown point 1/2. A finite-sample error bound and a matching information-theoretic minimax lower bound agree on the parametric rate sigma*sqrt(d/N) and differ on the breakdown floor by a factor of sqrt(d), a statistical-computational gap that polynomial-time RoPoLL pays relative to the intractable Tukey halfspace median. Across 13 open-weight judges (4B-675B), three reward-model benchmarks, and four corruption regimes at rates up to 50%, RoPoLL dominates PoLL on every biased corruption type: by about 19% on cross-dimensional attacks at matched compute, and by orders of magnitude on heavy-tailed Byzantine adversaries. A 3-judge RoPoLL committee at 38B beats Mistral-Large-3 (675B) by 1.31x on HelpSteer-2 under 30% bimodal-random corruption, an 18x parameter advantage at better accuracy; a Noisy-GT control confirms the premium is paid against biased contamination, not benign imprecision.

source & further reading

arxiv.org — original article

~/api · this article 200

$curl api.wpnews.pro/v1/news/ropoll-robust-panel-of-l…

Read original on arxiv.org → arxiv.org/abs/2606.30931

mentioned entities

RoPoLL

PoLL

LLM Jury

Huber contamination model

Geometric Median

HelpSteer-2

Mistral-Large-3

metadata

slugropoll-robust-panel-of-llm-judges

topic#large-language-models

secondary1 topics

sentimentpositive

canonicalarxiv.org

navigation

← prevI Built 5 Free AI Tools That Rep…

next →Sivers emission övertecknades "f…

── more in #large-language-models 4 stories · sorted by recency

arxiv.org · 1 Jul · #large-language-models

Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG

arxiv.org · 1 Jul · #large-language-models

Truth or Sophistry? LoFa: A Benchmark for LLM Robustness Against Logical Fallacies

arxiv.org · 1 Jul · #large-language-models

Neuro-Bayesian-Symbolic Residual Attention Shallow Network: Explainable Deep Learning for Cybersecurity Risk Assessment

lesswrong.com · 1 Jul · #large-language-models

Apply to the Inaugural PIBBSS Winter Research Fellowship!

── more on @ropoll 3 stories trending now

wpnews · 30 May · #ai-tools

I was wasting 10 minutes every Claude session. So I built a fix.

wpnews · 27 May · #machine-learning

hunting for headroom on modded-nanoGPT (WR #82)

wpnews · 2 Jun · #ai-products

Microsoft launches Discovery platform for scientific R&D with Ginkgo Bioworks partnership

sponsored brought to you by zahid.host 4,200+ EU-deployed projects

reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main

→ Live at https://your-agent.zahid.host ✓

Get free account → Pricing

from €0/mo · no card required