cd /news/artificial-intelligence/gobench-evaluating-llms-on-9x9-go-us… · home topics artificial-intelligence article
[ARTICLE · art-131844] src=rolandgao.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

GoBench: Evaluating LLMs on 9×9 Go using KataGo opponents as Elo anchors

GoBench, a new benchmark from researcher Roland Gao, evaluates frontier large language models on 9×9 Go against a calibrated ladder of KataGo opponents used as Elo anchors, with an arXiv release scheduled for September 17, 2026. The benchmark has two tracks: Track 1 tests general reasoning via multi-turn APIs without tools, while Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation, with each game capped at 30 minutes in a sandbox with resource limits and no internet access. Models listed in the evaluation include GPT-6 Astra, Claude Opus 5, GPT-5.6 Sol, Gemini 3.1 Pro, DeepSeek V4.1 Flash, Gemini 3.8 Flash, DeepSeek V4 Flash 0731, Muse Spark 1.3 Contributor, Gemini 3.6 Flash, GPT-5.6 Luna, and Grok 4.6.

read1 min views1 publishedSep 16, 2026
GoBench: Evaluating LLMs on 9×9 Go using KataGo opponents as Elo anchors
Image: source

GoBench measures how well frontier language models play 9×9 Go against a calibrated ladder of KataGo opponents.

GitHub · Paper (PDF) · X post

The arXiv release is scheduled for September 17, 2026.

(a) KataGo and LLMs

(b) LLMs

KataGo (132)

GPT-6 Astra

Claude Opus 5

GPT-5.6 Sol

Gemini 3.1 Pro

DeepSeek V4.1 Flash

Gemini 3.8 Flash

DeepSeek V4 Flash 0731

Muse Spark 1.3 Contributor

Gemini 3.6 Flash

GPT-5.6 Luna

Grok 4.6

Reasoning effortHighExtra highMax ## Why GoBench?

Current models show “jagged intelligence”: they approach top human performance in math and coding, yet lag in other domains. Progress toward AGI requires systems that can learn new domains at lower cost and with less human supervision.

GoBench measures general reasoning and context-based continual learning through two tracks:

General reasoning Track 1 uses multi-turn APIs without tools to test LLMs’ general reasoning abilities. 2. Context-based continual learning Track 2 gives Codex 0, 1, 2, 4, or 8 hours of continual learning before evaluation. Each evaluation game is capped at 30 minutes. Learning and evaluation take place in a sandbox with resource limits and no internet access.

We encourage researchers to extend GoBench to measure continual learning through weight updates.

Scroll horizontally for all columns, including seconds per move →

GPT-6 Astra · Codex

GPT-5.6 Sol · Codex

View Track 2 results as a table #

game replays…

GoPlay #

Play 9×9 Go against the same calibrated KataGo opponents. Pick an Elo, choose a color, and click Start game to play locally in your browser. Tromp-Taylor rules: area scoring, self-capture allowed, positional superko. 7 komi.

Latest moveNo moves played

Your estimated Elo1,000± 3,920

Completed games will appear here.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @gobench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gobench-evaluating-l…] indexed:0 read:1min 2026-09-16 ·