GPT-6 Astra Achieves SOTA on ARC-AGI OpenAI's GPT-6 Astra achieved state-of-the-art results on the ARC-AGI benchmark, scoring 63% on ARC-AGI-3 and 99% via a new provider adapter harness, surpassing human performance on 96% of ARC-AGI-3 levels, according to ARC Prize. The model also set a record 95.0% on ARC-AGI-2 at $1.12 per task and tied Fable 5's high score of 98.5% at $0.28 per task. ARC Prize on X: "GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:" GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis: GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis: Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment? This gives us an efficiency metric to compare AI to humans. With the provider Astra creates a dense compact symbolic world model to complete ARC-AGI-3 environments. For example, in environment s5i5, Astra: - Recorded the current level, hub orientation, and mechanism lengths: "L8: hub q2 8↓ . Lengths: 14=1…" - It mapped operations to exact controls: In addition to achieving state-of-the-art scores on ARC-AGI-3, GPT-6 Astra achieved a record 95.0% on ARC-AGI-2 at $1.12/task, and tied Fable 5's high score of 98.5% at $0.28/task. Full results: arcprize.org/results/openai… For ARC-AGI-3 we tested GPT-6 Astra with both our Standard harness the model decides which notes to carry forward as well as a new provider adapter harness which preserves opaque reasoning between requests and uses compaction . Going forward we will test all new models on