cd /news/large-language-models/conditional-accuracy-profiles-diagno… · home › topics › large-language-models › article
[ARTICLE · art-147377] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Conditional Accuracy Profiles: Diagnosing LLM Judges across Deployment Conditions

A new arXiv paper (arXiv:2610.09229v1) introduces Conditional Accuracy Profiling (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions spanning content sensitivity, robustness, and rationale quality. Instantiating CAP on seven LLM judges across six pairwise judging benchmarks, including the authors' new judgerEva-Standard testbed, the researchers found that on judgerEva's judge-independent Hard-Constructed subset the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, while Position Robustness showed the strongest rank stability (mean Spearman rho = 0.87) yet suffered the largest mean accuracy drop among shared conditions under JudgeBench-Pro adversarial stress. The authors argue condition-level profiles offer a more actionable basis than aggregate accuracy for selecting LLM judges.

by read1 min views2 publishedOct 8, 2026

arXiv:2610.09229v1 Announce Type: new Abstract: LLM-as-judge is now a standard tool for scalable evaluation, but judge performance is still often summarized by a single accuracy number. This aggregate view hides the deployment conditions under which a judge succeeds or fails. We introduce \textbf{Conditional Accuracy Profiling} (CAP), a post-hoc diagnostic framework that decomposes pairwise LLM-judge accuracy into eight conditions organized into content sensitivity, robustness, and rationale quality. CAP is benchmark-agnostic: it can be applied directly when a benchmark provides the required annotations, approximately through task-subset proxies, or through controlled augmentation when perturbation pairs can be generated. We instantiate CAP on seven LLM judges across six pairwise judging benchmarks, including \textsc{judgerEva-Standard}, a controlled testbed we created to support all eight conditions. CAP exposes profile differences hidden by aggregate accuracy: on \textsc{judgerEva}'s judge-independent Hard-Constructed subset, the two judges most sensitive to omitted qualifications rank in the bottom three of seven by overall accuracy, so omission sensitivity is not predicted by aggregate accuracy. Across benchmarks, Position Robustness shows the strongest rank stability (mean Spearman $\bar{\rho}{=}0.87$) but is itself fragile under JudgeBench-Pro adversarial stress, showing the largest mean accuracy drop among the shared conditions, though the dominant degradation channel varies by judge. Condition-level profiles provide a more actionable basis than aggregate accuracy for selecting LLM judges.

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/conditional-accuracy…] indexed:0 read:1min 2026-10-08 · —