{"slug": "gobench-evaluating-llms-on-9x9-go-using-katago-opponents-as-elo-anchors", "title": "GoBench: Evaluating LLMs on 9×9 Go using KataGo opponents as Elo anchors", "summary": "GoBench, a new benchmark from researcher Roland Gao, evaluates frontier large language models on 9×9 Go against a calibrated ladder of KataGo opponents used as Elo anchors, with an arXiv release scheduled for September 17, 2026. The benchmark has two tracks: Track 1 tests general reasoning via multi-turn APIs without tools, while Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation, with each game capped at 30 minutes in a sandbox with resource limits and no internet access. Models listed in the evaluation include GPT-6 Astra, Claude Opus 5, GPT-5.6 Sol, Gemini 3.1 Pro, DeepSeek V4.1 Flash, Gemini 3.8 Flash, DeepSeek V4 Flash 0731, Muse Spark 1.3 Contributor, Gemini 3.6 Flash, GPT-5.6 Luna, and Grok 4.6.", "body_md": "GoBench measures how well frontier language models play 9×9 Go against a calibrated ladder of KataGo opponents.\n\n[GitHub](https://github.com/RolandGao/gobench) · [Paper (PDF)](https://github.com/RolandGao/gobench/blob/main/paper/gobench2.pdf) · [X post](https://x.com/Roland65821498/status/2100253388562723298?s=20)\n\nThe arXiv release is scheduled for September 17, 2026.\n\n(a) KataGo and LLMs\n\n(b) LLMs\n\nKataGo (132)\n\nGPT-6 Astra\n\nClaude Opus 5\n\nGPT-5.6 Sol\n\nGemini 3.1 Pro\n\nDeepSeek V4.1 Flash\n\nGemini 3.8 Flash\n\nDeepSeek V4 Flash 0731\n\nMuse Spark 1.3 Contributor\n\nGemini 3.6 Flash\n\nGPT-5.6 Luna\n\nGrok 4.6\n\nReasoning effortHighExtra highMax ## Why GoBench?\n\nCurrent models show “[jagged intelligence](https://x.com/karpathy/status/1816531576228053133?s=20)”: they approach top human performance in math and coding, yet lag in other domains. Progress toward AGI requires systems that can **learn new domains at lower cost and with less human supervision**.\n\nGoBench measures general reasoning and context-based continual learning through two tracks:\n\n1. \n**General reasoning**\n Track 1 uses multi-turn APIs without tools to test LLMs’ general reasoning abilities.\n2. \n**Context-based continual learning**\n Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation. Each evaluation game is capped at 30 minutes. Learning and evaluation take place in a sandbox with resource limits and no internet access.\n\nWe encourage researchers to extend GoBench to measure continual learning through weight updates.\n\nScroll horizontally for all columns, including seconds per move →\n\nGPT-6 Astra · Codex\n\nGPT-5.6 Sol · Codex\n\n## View Track 2 results as a table\n\nLoading game replays…\n\n## GoPlay\n\nPlay 9×9 Go against the same calibrated KataGo opponents. Pick an Elo, choose a color, and click **Start game** to play locally in your browser. Tromp-Taylor rules: area scoring, self-capture allowed, positional superko. 7 komi.\n\nLatest move**No moves played**\n\nYour estimated Elo**1,000± 3,920**\n\nCompleted games will appear here.", "url": "https://wpnews.pro/news/gobench-evaluating-llms-on-9x9-go-using-katago-opponents-as-elo-anchors", "canonical_source": "https://rolandgao.com/blog/gobench/", "published_at": "2026-09-16 18:23:36+00:00", "updated_at": "2026-09-16 18:44:51.274542+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-agents", "ai-tools"], "entities": ["GoBench", "KataGo", "Roland Gao", "GPT-6 Astra", "Claude Opus 5", "GPT-5.6 Sol", "Gemini 3.1 Pro", "DeepSeek V4.1 Flash"], "alternates": {"html": "https://wpnews.pro/news/gobench-evaluating-llms-on-9x9-go-using-katago-opponents-as-elo-anchors", "markdown": "https://wpnews.pro/news/gobench-evaluating-llms-on-9x9-go-using-katago-opponents-as-elo-anchors.md", "text": "https://wpnews.pro/news/gobench-evaluating-llms-on-9x9-go-using-katago-opponents-as-elo-anchors.txt", "jsonld": "https://wpnews.pro/news/gobench-evaluating-llms-on-9x9-go-using-katago-opponents-as-elo-anchors.jsonld"}}