cd /news/artificial-intelligence/civbench-a-long-horizon-benchmark-fo… · home topics artificial-intelligence article
[ARTICLE · art-119891] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI

Researchers introduced CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments via the Model Context Protocol (MCP), featuring 300+ turn episodes and 76 MCP tools. In pilot runs across four model families, agents under-monitored strategic state, querying victory progress every 30-75 turns despite guidance to do so every 20 turns, and failed to execute near-term commitments (RAG@10 between 48.2% and 65.8%). The benchmark and analysis pipeline are available on GitHub.

read1 min views1 publishedSep 3, 2026

arXiv:2609.02459v1 Announce Type: new Abstract: We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single episode spans 300+ turns and produces thousands of tool calls over a large action space, requiring sustained planning, state monitoring, and execution under partial observability. The environment exposes 76 MCP tools and a narration layer that converts visual game state into structured text. We use CivBench to characterise agent behaviour across four model families in 23 admissible runs. The sample is a pilot, not a model ranking: aggregate outcomes do not reliably discriminate models at this scale. Instead, we introduce two interface-level metrics that the environment makes measurable: Proactive Monitoring Rate (PMR), capturing whether agents actively query latent strategic state, and RAG@10, capturing whether commitments stated in structured planning reflections are executed within ten subsequent turns. Across runs we observe two consistent patterns under a shared playbook protocol. Agents under-monitor strategically relevant state that is available but requires explicit querying: despite playbook guidance to query victory progress every 20 turns, agents do so only every 30 to 75 turns, and in 7 of 20 detectable defeats they failed to query within the 20 turn warning window before game end. Agents also frequently fail to execute near-term commitments stated in their own planning reflections (RAG@10 between 48.2% and 65.8% across models). Both patterns arise despite tool access and explicit guidance, and we interpret them as deviations under instruction rather than absences of capability. We release the environment, scenarios, logs, metrics, and analysis pipeline at https://github.com/lmwilki/civ6-mcp

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @civbench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/civbench-a-long-hori…] indexed:0 read:1min 2026-09-03 ·