04:00
2026-09-14
arxiv.org
large-language-models
GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
A new arXiv paper introduces GAUGE, an offline protocol that tests whether LLM-as-a-judge evaluation gates reliably rank task-oriented agents, and finds that satisfaction ratings carry essentially no …