cd /news/artificial-intelligence/judge-retrieve-or-abstain-uncertaint… · home topics artificial-intelligence article
[ARTICLE · art-102367] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Judge, Retrieve, or Abstain: Uncertainty-Guarded LLM Judging with Provable Risk Guarantees

Researchers propose a risk-controlled framework for LLM-as-a-judge evaluation that calibrates uncertainty thresholds on a held-out set to keep the false discovery rate among accepted verdicts below a user-specified level α with high probability, using finite-sample Clopper-Pearson intervals. The framework routes low-confidence instances to a retrieval-augmented mode with a second calibrated threshold, maintaining the guarantee without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, it achieves substantially higher coverage than single-mode baselines while maintaining the target error rate.

read1 min views1 publishedAug 19, 2026

arXiv:2608.17994v1 Announce Type: new Abstract: Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists. However, objective tasks introduce a distinct reliability challenge for reference-free LLM judging. In the absence of a reference answer, the judge evaluates factual correctness either through its parametric knowledge or through tool augmentation. Although the former enables efficient evaluation, the judge may hallucinate or lack sufficient evidence for its verdict. Conversely, tool augmentation can provide additional evidence but introduces extra computational cost and requires an appropriate mechanism to determine when and how that evidence should be used reliably. More importantly, neither approach alone provides formal control over the risk of accepted verdicts or guarantees their reliability at a specified level. We propose a risk-controlled framework that calibrates uncertainty thresholds on a held-out set so that the false discovery rate among accepted verdicts remains below a user-specified level~$\alpha$ with high probability, using finite-sample Clopper--Pearson intervals. When the parametric mode is not sufficiently confident, the instance is routed to a retrieval-augmented mode, where the judge gathers web evidence and re-evaluates the instance under a second calibrated threshold. The finite-sample guarantee carries over to this two-threshold routing without additional assumptions. Across open-domain QA benchmarks and judges of varying scales, the framework maintains the target error rate while achieving substantially higher coverage than single-mode baselines.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/judge-retrieve-or-ab…] indexed:0 read:1min 2026-08-19 ·