arXiv:2610.02425v1 Announce Type: new Abstract: Static evaluations credit a language model for naming the right move, but an agent must carry a plan through to a verified outcome while an opponent responds. We introduce XiangqiBench, an executable benchmark that measures this difference in Chinese chess: starting from 119 tactical endgames with forced mates supported by engine or checks-only search, an LLM agent must deliver checkmate against an engine defender. An interactive REPL interface separates real moves, state queries, and forward simulation, and we record 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. Three signals that look like competence each overstate closed-loop success. (i) The Conversion Gap: models play the stored reference first move in 26.1% of Sighted trials, yet only 13.9% of these trials end in a win. (ii) The Consistency Gap: the leading model reaches 38.7% pass@3 but only 5.9% pass^3, winning all three trials on 7 of the 46 positions it ever wins. (iii) The Simulation Gap: 32.3% of accepted simulation calls stop on an illegal move, and in 49.3% of comparable cases the real defender replies differently from the line the agent simulated; self-authored rollouts check legality but cannot anticipate the opponent. Finding the move is not winning the game: agent evaluations should score closed-loop outcomes and report reliability alongside coverage.
Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
A new arXiv paper (2610.02425v1) introduces XiangqiBench, an executable benchmark that measures closed-loop LLM agent performance in Chinese chess across 119 tactical endgames with forced mates, recording 8,568 multi-turn trajectories from 12 frontier LLMs under two observation protocols. The benchmark reports three gaps that overstate competence: models play the stored reference first move in 26.1% of Sighted trials but only 13.9% of those end in a win, the leading model reaches 38.7% pass@3 versus 5.9% pass^3, and 32.3% of accepted simulation calls stop on an illegal move while the real defender replies differently from the simulated line in 49.3% of comparable cases. The authors conclude that agent evaluations should score closed-loop outcomes and report reliability alongside coverage.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.