18:23
2026-09-16
rolandgao.com
artificial-intelligence
GoBench: Evaluating LLMs on 9Γ9 Go using KataGo opponents as Elo anchors
GoBench, a new benchmark from researcher Roland Gao, evaluates frontier large language models on 9Γ9 Go against a calibrated ladder of KataGo opponents used as Elo anchors, with an arXiv release schedβ¦