$26,000. That's what one full ARC-AGI-3 evaluation run costs at max reasoning, and OpenAI's GPT-6 Astra used it to blow past every prior score on one of AI's hardest benchmarks. The gap to its closest rivals is now wider than at any point this year.
On its Standard harness, ARC Prize measured GPT-6 Astra at 62.7% on ARC-AGI-3 Semi-Private, running the model at max reasoning for a bill of roughly $26,000 per full evaluation. That score alone more than doubles the previous frontier result. Claude Opus 5 had briefly held the lead at 30.2% in July, and GPT-5.6 Sol, OpenAI's own prior flagship, managed 7.8%. Anthropic has not published a verified Opus 5.5 score on ARC-AGI-3 yet. On the xAI side, Grok 4.6's XHigh run scored 2.11%, while earlier Grok 4.5 entries sit at 0.3%. There is no verified Grok 4.7 result on the leaderboard as of this week.
Run it through OpenAI's own Provider Adapter harness instead: this one preserves the model's reasoning state between actions rather than resetting it at every step. Astra's best observed result climbs to 99.9%, at high reasoning, for about $18,800. That's not a different model. It's a different measuring stick, and ARC Prize was careful to say so in its own writeup.
The gap between 62.7% and 99.9% isn't OpenAI inflating its own homework. ARC Prize's Standard harness is provider-neutral: it lets a model carry forward notes it chooses to preserve, but does not preserve the provider's opaque reasoning state between requests. That's the same constraint applied to every lab's model on the board. OpenAI's Provider Adapter harness, by contrast, lets Astra carry forward opaque reasoning state and use compaction across an entire ARC-AGI-3 puzzle. Give a model access to that kind of retained state and the score nearly triples.
That's the honest headline number to compare against rivals: 62.7%, not 99.9%, because that's the one measured under conditions every other model faces too. And even at that lower, fairer number, Astra still leads by a wide margin.
OpenAI's GPT-6 Astra jumped to 62.7% on AI's hardest reasoning test OpenAI's flagship model scored 62.7% under ARC Prize's neutral test setup but 99.9% using a proprietary adapter only OpenAI can run, a 37-point gap that's splitting benchmark watchers on which number is real progress. - how to score highest on AI reasoning benchmarks - GPT-6 Astra performance on ARC Prize interactive test
ARC Prize's analysis found something else worth noting beyond the raw score. Astra used fewer actions than the median human tester on 96% of the levels it was run against, averaging 51.7% fewer moves to reach a solution. The model appeared to build compact symbolic representations of game environments it had never seen before, then reuse those representations to plan later moves rather than re-exploring from zero each time. That's closer to how a person builds a mental model of a new video game than how earlier language models tended to brute-force their way through novel environments.
What OpenAI changed to get there #
OpenAI has not published Astra's architecture in detail, but the scale of the training run is not in dispute. Aidan Clark, OpenAI's vice president of research, said this was the first time the company pretrained a model on more than 100,000 GPUs at its Stargate site in Texas. It was also the first run in which earlier OpenAI models actively supervised the training of their successor. Independent analysis from AI researcher Sebastian Raschka discusses reports that Astra may use a technique sometimes called recurrent depth, or looped transformers, while noting that OpenAI has not confirmed the architecture. OpenAI itself has said Astra completes harder tasks with less of its reasoning verbalized and exercises tighter control over what shows up in the reasoning trace users actually see.
That opacity cuts both ways. A model that reasons more and shows less is harder to audit, and OpenAI's own materials acknowledge Astra is less monitorable through chain-of-thought inspection than its predecessors, even as it performs better.
GPT-6 Astra itself isn't brand new. OpenAI announced it on September 3 and rolled it out over the following days across ChatGPT Plus, Pro, Business and Enterprise tiers, plus the API, Microsoft Azure and AWS Bedrock, at a price of $10 per million input tokens and $50 per million output tokens, with a 1.05 million-token context window. The ARC-AGI-3 result is the follow-up story: the number that tells labs, and the investors funding them, how far ahead Astra actually sits once real-world agentic reasoning gets measured rather than assumed.
For now, the honest scoreboard reads 62.7% for Astra against 30.2% for the last Claude to post a verified number and single digits for the last Grok to try. Anthropic and xAI haven't answered yet. ARC Prize's board is the one every lab will be racing to climb next. Also read: How to Price a Freemium AI Agent Without Giving Away Your Margin • DoorDash lets you order dinner by texting an AI agent in Apple Messages • DeepSeek Narrows AI Gap With US to Just 3 Percent, Bloomberg Says
OpenAI Changed GPT-6 Astra's Benchmark Numbers Days After Its Launch Fortune reported that OpenAI quietly revised several GPT-6 Astra benchmark figures after its September 3 launch, including cutting its hallucination rate in half before later reverting it, and boosting a cybersecurity score using a reasoning tier that isn't commercially available. The changes mostly flattered Astra, though some of Anthropic's... - OpenAI changed GPT-6 Astra benchmark numbers after launch - how OpenAI modified benchmark results for GPT-6 Astra
This article is posted in AI News, check it out for more related stories.
Join the discussion #
Open in the community → Almost there. Sign in and your reply posts straight away.