cd /news/artificial-intelligence/trajectory-judge-what-outcome-only-l… · home topics artificial-intelligence article
[ARTICLE · art-118560] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

A new arXiv study (2609.00038v1) finds that outcome-only LLM judges miss silent agent failures, catching only 45% of silent faults while flagging 33% of correct trajectories, whereas a step-rubric judge achieves 77% silent recall with zero false alarms at 3x the cost. The researchers argue judge evaluations must stratify recall by outcome survival and release their environment, fault injector, raw verdicts, and analysis pipeline.

read1 min views2 publishedSep 2, 2026

arXiv:2609.00038v1 Announce Type: new Abstract: Outcome-only evaluation is the production default for LLM agents: show a judge the request and the final reply and ask whether it was handled well. The metric is structurally blind to an agent that reaches the right answer the wrong way. We measure that blind spot where ground truth is known by construction: a deterministic tool-using support-desk environment, a scripted oracle policy that always solves it, and a fault injector that breaks exactly one thing at a known step, stratifying faults by whether the customer-visible outcome survived (silent) or not (loud). Five judges (programmatic rules, outcome-only, step-rubric at two model sizes, and a self-consistency ensemble) are scored on detection, step localisation, fault typing, calibration, and cost over 400 trajectories. The outcome-only judge catches 84% of loud faults but 45% of silent ones while flagging 33% of correct trajectories; a step-rubric judge reaches 77% silent recall with zero false alarms at 3x the cost. No judge reads the final reply: an invented promise appended to an otherwise perfect trajectory evades the rules entirely and the step judge 82% of the time, and self-consistency triples cost while improving nothing. We argue that judge evaluations must stratify recall by outcome survival, and release the environment, the injector, all raw verdicts, and an analysis pipeline that rebuilds every number offline.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/trajectory-judge-wha…] indexed:0 read:1min 2026-09-02 ·