cd /news/large-language-models/gauge-when-not-to-trust-llm-as-a-jud… · home topics large-language-models article
[ARTICLE · art-128723] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↓ negative

GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents

A new arXiv paper introduces GAUGE, an offline protocol that tests whether LLM-as-a-judge evaluation gates reliably rank task-oriented agents, and finds that satisfaction ratings carry essentially no information about task success: 57.5% of conversations rated satisfied by a blind panel failed the customer's task. Across 25 agents from six providers on the τ²-bench and SimulatorArena benchmarks, the gate's decision-disagreement rate rose from under 1% on wide-reward pairs to 31% on close pairs, leading the authors to propose a calibrate-then-trust cadence using a judge-free completion bit as a zero-cost tripwire for truncation regressions.

by read1 min views1 publishedSep 14, 2026

arXiv:2609.12191v1 Announce Type: new Abstract: Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures whether this gate's ranking matches a grounded verifiable reward across 25 agents from six providers on the $\tau^2$-bench and SimulatorArena benchmarks, separating two kinds of evaluation validity that release practices conflate: ranking validity and construct validity. First, a satisfaction-success gap: satisfaction carries essentially no information about task success, as conversations rated satisfied by our blind panel are decorrelated from actual success, with 57.5% of them failing the customer's task, a pattern consistent across five rater populations, both benchmarks, and every subjective dimension we rated. Second, while the gate's ranking is robust across the broad capability span, it loses resolution among the near-equal strong agents: this decision-disagreement rate jumps from $<$1% on wide-reward pairs to 31% on close pairs. The gate is thus human-validated yet mis-anchored. As a remedy, we propose a calibrate-then-trust cadence in which a judge-free completion bit is a zero-cost tripwire for truncation regressions.

── more in #large-language-models 4 stories · sorted by recency
── more on @gauge 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gauge-when-not-to-tr…] indexed:0 read:1min 2026-09-14 ·