GoBench measures how well frontier language models play 9×9 Go against a calibrated ladder of KataGo opponents.
GitHub · Paper (PDF) · X post
The arXiv release is scheduled for September 17, 2026.
(a) KataGo and LLMs
(b) LLMs
KataGo (132)
GPT-6 Astra
Claude Opus 5
GPT-5.6 Sol
Gemini 3.1 Pro
DeepSeek V4.1 Flash
Gemini 3.8 Flash
DeepSeek V4 Flash 0731
Muse Spark 1.3 Contributor
Gemini 3.6 Flash
GPT-5.6 Luna
Grok 4.6
Reasoning effortHighExtra highMax ## Why GoBench?
Current models show “jagged intelligence”: they approach top human performance in math and coding, yet lag in other domains. Progress toward AGI requires systems that can learn new domains at lower cost and with less human supervision.
GoBench measures general reasoning and context-based continual learning through two tracks:
General reasoning Track 1 uses multi-turn APIs without tools to test LLMs’ general reasoning abilities. 2. Context-based continual learning Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation. Each evaluation game is capped at 30 minutes. Learning and evaluation take place in a sandbox with resource limits and no internet access.
We encourage researchers to extend GoBench to measure continual learning through weight updates.
Scroll horizontally for all columns, including seconds per move →
GPT-6 Astra · Codex
GPT-5.6 Sol · Codex
View Track 2 results as a table #
game replays…
GoPlay #
Play 9×9 Go against the same calibrated KataGo opponents. Pick an Elo, choose a color, and click Start game to play locally in your browser. Tromp-Taylor rules: area scoring, self-capture allowed, positional superko. 7 komi.
Latest moveNo moves played
Your estimated Elo1,000± 3,920
Completed games will appear here.