{"slug": "nanogpt-speedrun-frontier", "title": "NanoGPT Speedrun Frontier", "summary": "A benchmark of 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun shows Fable 5, run via claude-code at high effort, achieved the best validated result of 2,726 tokens with an 81.7% improvement over baseline, followed by Opus 5 at 2,920 tokens (53.6%) and Kimi K3 at 2,930 tokens (52.2%). The runs, conducted by the NanoGPT Speedrun Frontier project, also include an equal-budget comparison and 41 curated agent trajectories.", "body_md": "# NanoGPT Speedrun Frontier\n\nWe ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.\n\nAll modelsBest validated result for each model\n\n1Fable 52,72681.7% closed\n\nclaude-code · high@24H 3,0108.7d\n\n2Opus 52,92053.6% closed\n\nclaude-code · max@24H 3,0452.9d\n\n3Kimi K32,93052.2% closed\n\nprime-agent · max@24H 3,1253.6d\n\n4Kimi K32,97445.8% closed\n\nkimi-code · max@24H 3,1355.1d\n\n5Opus 4.83,01839.4% closed\n\nclaude-code · max@24H 3,1803.0d\n\n6GPT-5.6 Sol3,04235.9% closed\n\ncodex · xhigh@24H 3,1606.1d\n\n7GPT-5.6 Sol Pro3,05833.6% closed\n\ncodex · xhigh@24H 3,1003.4d\n\n8Sonnet 53,10526.8% closed\n\nclaude-code · max@24H 3,1202.0d\n\n9GPT-5.6 Luna3,11026.1% closed\n\ncodex · xhigh@24H 3,1701.9d\n\n10Grok 4.53,12024.6% closed\n\ngrok-cli · xhigh@24H 3,1602.7d\n\n11Qwen3.8 Max3,12024.6% closed\n\nqwen-code · max@24H 3,2251.9d\n\n12GLM 5.23,15020.3% closed\n\npi · high@24H 3,2001.8d\n\n13DeepSeek V4 Pro3,20512.3% closed\n\nclaude-code · max@24H 3,2051.1d\n\n14GPT-5.6 Terra3,21411.0% closed\n\ncodex · xhigh@24H 3,2141.1d\n\n15Grok 4.63,22010.1% closed\n\ngrok-cli · xhigh0.6d\n\n16Muse Spark 1.23,2308.7% closed\n\nmuse-code · xhigh0.6d\n\n17Muse Spark 1.13,2328.4% closed\n\npi · max@24H 3,2403.7d\n\n18GPT-5.53,2348.1% closed\n\ncodex · xhigh@24H 3,2341.1d\n\n19Kimi K2.73,2407.2% closed\n\nkimi-code · max@24H 3,2401.6d\n\n20GLM 5.3—no record\n\nclaude-code · xhigh\n\n| Model | Harness | Traces | ||||||||\n|---|---|---|---|---|---|---|---|---|---|---|\n| Fable 5 | claude-code · high | 2,726 | 81.7% | 3,010 | 800M | 1.1M | 811 | 3k | 8.7 | |\n| Opus 5 | claude-code · max | 2,920 | 53.6% | 3,045 | 183M | 690k | 292 | 401 | 2.9 | |\n| Kimi K3 | prime-agent · max | 2,930 | 52.2% | 3,125 | 112M | 2.2M | — | 488 | 3.6 | |\n| Kimi K3 | kimi-code · max | 2,974 | 45.8% | 3,135 | 682M | 1.4M | 713 | 4k | 5.1 | |\n| Opus 4.8 | claude-code · max | 3,018 | 39.4% | 3,180 | 318M | 2.3M | 427 | 2k | 3.0 | |\n| GPT-5.6 Sol | codex · xhigh | 3,042 | 35.9% | 3,160 | 2.9B | 2.2M | 963 | 28k | 6.1 | |\n| GPT-5.6 Sol Pro | codex · xhigh | 3,058 | 33.6% | 3,100 | 1.2B | 4.6M | 509 | 7k | 3.4 | |\n| Sonnet 5 | claude-code · max | 3,105 | 26.8% | 3,120 | 998M | 2.1M | 213 | 2k | 2.0 | |\n| GPT-5.6 Luna | codex · xhigh | 3,110 | 26.1% | 3,170 | 894M | 888k | 362 | 12k | 1.9 | |\n| Grok 4.5 | grok-cli · xhigh | 3,120 | 24.6% | 3,160 | 46M | 385k | 399 | 4k | 2.7 | |\n| Qwen3.8 Max | qwen-code · max | 3,120 | 24.6% | 3,225 | 216M | 629k | 312 | 866 | 1.9 | |\n| GLM 5.2 | pi · high | 3,150 | 20.3% | 3,200 | 57M | 1.7M | 194 | 1k | 1.8 | |\n| DeepSeek V4 Pro | claude-code · max | 3,205 | 12.3% | 3,205 | 26M | 319k | 189 | 309 | 1.1 | |\n| GPT-5.6 Terra | codex · xhigh | 3,214 | 11.0% | 3,214 | 417M | 298k | 154 | 3k | 1.1 | |\n| Grok 4.6 | grok-cli · xhigh | 3,220 | 10.1% | — | 27M | 346k | 97 | 691 | 0.6 | |\n| Muse Spark 1.2 | muse-code · xhigh | 3,230 | 8.7% | — | 41M | 910k | 56 | 724 | 0.6 | — |\n| Muse Spark 1.1 | pi · max | 3,232 | 8.4% | 3,240 | 122M | 1.6M | 489 | 2k | 3.7 | |\n| GPT-5.5 | codex · xhigh | 3,234 | 8.1% | 3,234 | 70M | 77k | 185 | 614 | 1.1 | |\n| Kimi K2.7 | kimi-code · max | 3,240 | 7.2% | 3,240 | 160M | 763k | 187 | 3k | 1.6 | |\n| GLM 5.3 | claude-code · xhigh | — | — | — | — | — | — | — | — | — |\n\n## Equal-budget comparison\n\nGive each model's best final run the same resource budget and compare the best validated record it reached within that budget.\n\nModelRecordHuman 2,600Baseline 3,290\n\nGray runs ended before the selected budget\n\nLoading available trajectories…\n\nOpen Traces to explore 41 curated full agent trajectories, including tool calls, subagents, and scratchpads.", "url": "https://wpnews.pro/news/nanogpt-speedrun-frontier", "canonical_source": "https://www.primeintellect.ai/research/nanogpt-speedrun", "published_at": "2026-08-22 22:14:27+00:00", "updated_at": "2026-08-22 22:43:37.663435+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-agents"], "entities": ["Fable 5", "Opus 5", "Kimi K3", "GPT-5.6 Sol", "Grok 4.5", "Qwen3.8 Max", "GLM 5.2", "DeepSeek V4 Pro"], "alternates": {"html": "https://wpnews.pro/news/nanogpt-speedrun-frontier", "markdown": "https://wpnews.pro/news/nanogpt-speedrun-frontier.md", "text": "https://wpnews.pro/news/nanogpt-speedrun-frontier.txt", "jsonld": "https://wpnews.pro/news/nanogpt-speedrun-frontier.jsonld"}}