ARC Prize on X: "GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:
- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
- It surpasses human performance on 96% of ARC-AGI-3 levels
- It builds the most precise symbolic model of novel environments we've seen Our analysis:"
GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:
- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
- It surpasses human performance on 96% of ARC-AGI-3 levels
- It builds the most precise symbolic model of novel environments we've seen Our analysis:
GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI:
- Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness
- It surpasses human performance on 96% of ARC-AGI-3 levels
- It builds the most precise symbolic model of novel environments we've seen Our analysis:
Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment? This gives us an efficiency metric to compare AI to humans. With the provider
Astra creates a dense compact symbolic world model to complete ARC-AGI-3 environments.
For example, in environment s5i5, Astra:
- Recorded the current level, hub orientation, and mechanism lengths: "L8: hub q2 (8↓). Lengths: 14=1…"
- It mapped operations to exact controls:
In addition to achieving state-of-the-art scores on ARC-AGI-3, GPT-6 Astra achieved a record 95.0% on ARC-AGI-2 at $1.12/task, and tied Fable 5's high score of 98.5% at $0.28/task. Full results: arcprize.org/results/openai…
For ARC-AGI-3 we tested GPT-6 Astra with both our Standard harness (the model decides which notes to carry forward) as well as a new provider adapter harness (which preserves opaque reasoning between requests and uses compaction). Going forward we will test all new models on