04:00
2026-10-05
arxiv.org
artificial-intelligence
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
A new audit suite called O*NET-BENCH, derived from a survey of 45,796 worker ratings, found that 33 pre-existing LLM judge configurations across six model families estimated acceptance rates ranging fâŚ