cd /news/artificial-intelligence/right-order-wrong-scale-auditing-llm… · home › topics › artificial-intelligence › article
[ARTICLE · art-145134] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement

A new audit suite called O*NET-BENCH, derived from a survey of 45,796 worker ratings, found that 33 pre-existing LLM judge configurations across six model families estimated acceptance rates ranging from 3.0% to 97.9% on 4,501 test ratings, compared with 61.1% for occupation-matched workers. Twenty-five configurations reached tie-aware pair accuracy of at least 0.60, yet a train-fitted response-only TF-IDF baseline nearly matched the strongest judge, and cross-validated calibration explained at most 8.5% of individual worker-rating variance. The authors conclude that ranking agreement alone is insufficient for occupational measurement and that judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.

by read1 min views7 publishedOct 5, 2026

arXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while reducing agreement with worker means at the task and occupation levels; this reversal replicates on a task- and worker-disjoint validation split under prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker-rating variance. Prediction-assisted estimation yields at most small precision gains at the studied label budgets. These results show that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @o*net-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/right-order-wrong-sc…] indexed:0 read:1min 2026-10-05 · —