arXiv:2609.29769v2 Announce Type: replace Abstract: LLM judges score outputs against rubrics well enough to have become the norm, both in benchmarks and as rewards for training. Jev, a classifier-like alternative its creators call a "decision model", returns probabilities over permitted answers with a calibrated confidence score, which LLM judges do not natively provide. We compare Jev with three flash-tier LLM judges on nine panels from seven benchmarks with human judgments, giving every judge identical criterion texts. The LLM judges run in two setups: holistically, reading a whole rubric at once as Jev does, and one criterion at a time. Jev can often stand in for them. They cost 16 to 325 times as much and take 28 to 350 times as long, yet in each setup Jev's accuracy differs significantly from theirs in at most 8 of 27 paired comparisons, ahead mostly on binary checklist criteria and behind only on ordinal ones. Despite their different designs, the two kinds of judge err alike. On ordinal criteria, all LLM judges and Jev depart from the human raters together, agreeing more with one another than with the labels and mostly assigning lower levels. On Jev's most confident errors, about 96% of LLM verdicts repeat its wrong answer, where independent errors would give about half. Intuitively, calibrated confidence should make Jev an ideal first stage of a cascade that defers uncertain verdicts to an LLM judge. Yet such cascades only lower cost while adding little accuracy: even with oracle thresholds, none beats the best single judge by more than 2.7 points. Calibration can tell a cascade when to defer, but the cascade also needs a fallback that errs elsewhere; these judges are wrong in the same places. These findings, which hold in both setups and at high reasoning effort, suggest that a cascade of judges succeeds only when its judges make complementary errors, and that future decision models should be designed afresh with that aim.
JEV vs. LLMs as Rubric Judges: Cheaper, Faster, and Wrong in the Same Places
A study posted to arXiv (2609.29769v2) found that Jev, a "decision model" that returns calibrated probabilities over permitted answers, can often substitute for three flash-tier LLM judges on nine panels from seven benchmarks with human judgments, despite the LLM judges costing 16 to 325 times as much and taking 28 to 350 times as long. Jev's accuracy differed significantly from the LLM judges in at most 8 of 27 paired comparisons in each setup, ahead mostly on binary checklist criteria and behind only on ordinal ones, and on Jev's most confident errors about 96% of LLM verdicts repeated its wrong answer, where independent errors would give about half. Cascades that defer uncertain Jev verdicts to an LLM judge lowered cost but added little accuracy, with none beating the best single judge by more than 2.7 points even at oracle thresholds, leading the authors to conclude that judge cascades succeed only when their judges make complementary errors.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.