04:00
2026-10-05
arxiv.org
artificial-intelligence
Finding the Move Is Not Winning the Game: XiangqiBench for Closed-Loop Evaluation of LLM Agents
A new arXiv paper (2610.02425v1) introduces XiangqiBench, an executable benchmark that measures closed-loop LLM agent performance in Chinese chess across 119 tactical endgames with forced mates, recorβ¦