{"slug": "gpt-6-astra-achieves-sota-on-arc-agi", "title": "GPT-6 Astra Achieves SOTA on ARC-AGI", "summary": "OpenAI's GPT-6 Astra achieved state-of-the-art results on the ARC-AGI benchmark, scoring 63% on ARC-AGI-3 and 99% via a new provider adapter harness, surpassing human performance on 96% of ARC-AGI-3 levels, according to ARC Prize. The model also set a record 95.0% on ARC-AGI-2 at $1.12 per task and tied Fable 5's high score of 98.5% at $0.28 per task.", "body_md": "ARC Prize on X: \"GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:\n- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness\n- It surpasses human performance on 96% of ARC-AGI-3 levels\n- It builds the most precise symbolic model of novel environments we've seen\nOur analysis:\"\n\nGPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:\n- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness\n- It surpasses human performance on 96% of ARC-AGI-3 levels\n- It builds the most precise symbolic model of novel environments we've seen\nOur analysis:\n\nGPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:\n- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness\n- It surpasses human performance on 96% of ARC-AGI-3 levels\n- It builds the most precise symbolic model of novel environments we've seen\nOur analysis:\n\nBefore launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment?\nThis gives us an efficiency metric to compare AI to humans.\nWith the provider\n\nAstra creates a dense compact symbolic world model to complete ARC-AGI-3 environments.\nFor example, in environment s5i5, Astra:\n- Recorded the current level, hub orientation, and mechanism lengths: \"L8: hub q2 (8↓). Lengths: 14=1…\"\n- It mapped operations to exact controls:\n\nIn addition to achieving state-of-the-art scores on ARC-AGI-3, GPT-6 Astra achieved a record 95.0% on ARC-AGI-2 at $1.12/task, and tied Fable 5's high score of 98.5% at $0.28/task.\nFull results: arcprize.org/results/openai…\n\nFor ARC-AGI-3 we tested GPT-6 Astra with both our Standard harness (the model decides which notes to carry forward) as well as a new provider adapter harness (which preserves opaque reasoning between requests and uses compaction).\nGoing forward we will test all new models on", "url": "https://wpnews.pro/news/gpt-6-astra-achieves-sota-on-arc-agi", "canonical_source": "https://twitter.com/arcprize/status/2095597602545025138", "published_at": "2026-09-03 19:53:01+00:00", "updated_at": "2026-09-03 20:23:49.693568+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["OpenAI", "GPT-6 Astra", "ARC Prize", "ARC-AGI-3", "ARC-AGI-2", "Fable 5"], "alternates": {"html": "https://wpnews.pro/news/gpt-6-astra-achieves-sota-on-arc-agi", "markdown": "https://wpnews.pro/news/gpt-6-astra-achieves-sota-on-arc-agi.md", "text": "https://wpnews.pro/news/gpt-6-astra-achieves-sota-on-arc-agi.txt", "jsonld": "https://wpnews.pro/news/gpt-6-astra-achieves-sota-on-arc-agi.jsonld"}}