GoBench: Evaluating LLMs on 9×9 Go using KataGo opponents as Elo anchors GoBench, a new benchmark from researcher Roland Gao, evaluates frontier large language models on 9×9 Go against a calibrated ladder of KataGo opponents used as Elo anchors, with an arXiv release scheduled for September 17, 2026. The benchmark has two tracks: Track 1 tests general reasoning via multi-turn APIs without tools, while Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation, with each game capped at 30 minutes in a sandbox with resource limits and no internet access. Models listed in the evaluation include GPT-6 Astra, Claude Opus 5, GPT-5.6 Sol, Gemini 3.1 Pro, DeepSeek V4.1 Flash, Gemini 3.8 Flash, DeepSeek V4 Flash 0731, Muse Spark 1.3 Contributor, Gemini 3.6 Flash, GPT-5.6 Luna, and Grok 4.6. GoBench measures how well frontier language models play 9×9 Go against a calibrated ladder of KataGo opponents. GitHub https://github.com/RolandGao/gobench · Paper PDF https://github.com/RolandGao/gobench/blob/main/paper/gobench2.pdf · X post https://x.com/Roland65821498/status/2100253388562723298?s=20 The arXiv release is scheduled for September 17, 2026. a KataGo and LLMs b LLMs KataGo 132 GPT-6 Astra Claude Opus 5 GPT-5.6 Sol Gemini 3.1 Pro DeepSeek V4.1 Flash Gemini 3.8 Flash DeepSeek V4 Flash 0731 Muse Spark 1.3 Contributor Gemini 3.6 Flash GPT-5.6 Luna Grok 4.6 Reasoning effortHighExtra highMax Why GoBench? Current models show “ jagged intelligence https://x.com/karpathy/status/1816531576228053133?s=20 ”: they approach top human performance in math and coding, yet lag in other domains. Progress toward AGI requires systems that can learn new domains at lower cost and with less human supervision . GoBench measures general reasoning and context-based continual learning through two tracks: 1. General reasoning Track 1 uses multi-turn APIs without tools to test LLMs’ general reasoning abilities. 2. Context-based continual learning Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation. Each evaluation game is capped at 30 minutes. Learning and evaluation take place in a sandbox with resource limits and no internet access. We encourage researchers to extend GoBench to measure continual learning through weight updates. Scroll horizontally for all columns, including seconds per move → GPT-6 Astra · Codex GPT-5.6 Sol · Codex View Track 2 results as a table Loading game replays… GoPlay Play 9×9 Go against the same calibrated KataGo opponents. Pick an Elo, choose a color, and click Start game to play locally in your browser. Tromp-Taylor rules: area scoring, self-capture allowed, positional superko. 7 komi. Latest move No moves played Your estimated Elo 1,000± 3,920 Completed games will appear here.