AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling A new benchmark, AgentJudgeBench, reveals that LLM judges' reliability on agentic tool-calling degrades monotonically with task difficulty, with all six judges converging to a narrow 77-82% alignment band on hard queries without ground truth, indicating a structural ceiling that model capacity alone cannot overcome. The benchmark, comprising 3,808 instances across six DAG topologies and three difficulty tiers, found that ground-truth exposure reduces alignment for GPT-5.4 by 1.5 percentage points and Gemini-2.5-Pro by 3.9 points, while structured evaluation rubrics improve alignment by up to 6.5 points but do not generalize uniformly. arXiv:2608.26623v1 Announce Type: new Abstract: LLM judges are widely used to evaluate agentic tool-calling systems, yet their reliability on structured, dependency-driven workflows remains largely unexamined. We present AgentJudgeBench, the first benchmark to systematically study LLM-as-a-judge reliability for agentic tool-calling over workflow DAGs, as distinct from the broader LLM-as-a-judge task of open-ended text or preference evaluation. The benchmark comprises 3,808 instances spanning six DAG topologies and three difficulty tiers, evaluated with five generators 3B-70B open-weight models and GPT-5.4 and six judges 20B to frontier scale under paired with- and without-ground-truth conditions. Judge alignment degrades monotonically with task difficulty, 1.5x faster without ground truth, and on hard queries without ground truth all six judges converge to a narrow 77-82% band regardless of scale, revealing a structural ceiling driven primarily by task difficulty, though its height is partly prompt-dependent for weaker generators, that model capacity alone cannot overcome. Ground-truth exposure is not uniformly beneficial: it reduces alignment for GPT-5.4 1.5 pp and Gemini-2.5-Pro 3.9 pp , consistent with over-anchoring. Among mitigation strategies, chain-of-thought reasoning and judge temperature both have negligible effect, while structured evaluation rubrics improve alignment by up to 6.5 pp but do not generalize uniformly across judge-generator pairs. With ground truth, QwQ-32B best matches the programmatic reference, while a human validation study identifies GPT-OSS-120B as the most human-aligned judge; without it, frontier judges lead only marginally within the shared ceiling. These results expose fundamental limitations of current LLM judges and yield practical guidelines for reliable evaluation in agentic systems.