cd /news/artificial-intelligence/openai-s-gpt-6-astra-jumped-to-62-7-… · home › topics › artificial-intelligence › article
[ARTICLE · art-144845] src=startupfortune.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

OpenAI's GPT-6 Astra jumped to 62.7% on AI's hardest reasoning test

OpenAI's GPT-6 Astra scored 62.7% on ARC-AGI-3 under the ARC Prize Foundation's standard, provider-independent harness, while a separate provider adapter harness that preserves an opaque reasoning state across turns produced the 99.9% figure OpenAI highlighted at the model's September 3 launch. The 62.7% more than doubled the prior frontier best of 30.2% set by Claude Opus 5 in July and far exceeded GPT-5.6 Sol's 7.8%, and ARC Prize Foundation co-founder François Chollet called GPT-6 Astra "a step-function change in model capability for interactive reasoning problems," noting it surpassed human performance on 96% of ARC-AGI-3 levels. Fortune reported that OpenAI quietly revised several GPT-6 Astra benchmark figures after launch, including halving its hallucination rate before reverting it and boosting a cybersecurity score using a reasoning tier that is not commercially available, and The New Stack reported the near-perfect 99.9% came from a setup OpenAI does not sell to developers.

by read6 min views1 publishedOct 4, 2026
OpenAI's GPT-6 Astra jumped to 62.7% on AI's hardest reasoning test
Image: Startupfortune (auto-discovered)

A benchmark designed to resist memorization just saw OpenAI's newest model nearly match human performance, and the number that got the headlines isn't the number ARC Prize says you should trust.

On September 3, OpenAI introduced GPT-6 Astra, and the company's own chart leaned hard on a single figure: 99.9% on ARC-AGI-3, the interactive reasoning benchmark built by François Chollet's ARC Prize Foundation. That score would mean Astra solved nearly every level of a test designed specifically to break models that rely on memorized patterns rather than genuine reasoning. But ARC Prize, which runs the official leaderboard, published a different number for the same model under its standard testing setup: 62.7%. Both figures are real. They measure two different things, and the gap between them has turned into the benchmark world's sharpest argument of the month.

Context makes the jump obvious. Before Astra, the best frontier score on ARC-AGI-3 belonged to GPT-5.6 Sol at 7.8%, according to ARC Prize's own published results. Claude Opus 5 had briefly held the lead at 30.2% after a July run that Chollet himself called an "impressive jump" on X. Astra's 62.7% more than doubled that. Its 99.9% figure, if taken at face value, would mean the benchmark Chollet has spent two years insisting resists brute-force scaling had, in the space of about six months, gone from sub-1% scores across every frontier model to functionally solved.

It hasn't, and the reason why is the real story.

ARC Prize tests every model on what it calls a standard harness. That's a neutral, provider-independent setup: it feeds the model an unfamiliar game and lets the model carry forward only the notes it chooses to preserve, without carrying an opaque reasoning state between requests. Under that harness, Astra scored 62.7%, according to ARC Prize's published results page. OpenAI also submitted Astra through a second setup, what the two organizations call a provider adapter harness. For Astra, that harness preserves what ARC Prize describes as an opaque reasoning state across turns and uses compaction for longer conversations, letting the model reuse prior cognitive work in a way the standard harness does not. That version scored 99.9%.

OpenAI Changed GPT-6 Astra's Benchmark Numbers Days After Its Launch Fortune reported that OpenAI quietly revised several GPT-6 Astra benchmark figures after its September 3 launch, including cutting its hallucination rate in half before later reverting it, and boosting a cybersecurity score using a reasoning tier that isn't commercially available. The changes mostly flattered Astra, though some of Anthropic's... - OpenAI changed GPT-6 Astra benchmark numbers after launch - how OpenAI modified benchmark results for GPT-6 Astra

Chollet did not downplay either number. "GPT-6 Astra is a step-function change in model capability for interactive reasoning problems," he wrote on X, noting the model also surpassed human performance on 96% of ARC-AGI-3 levels, using fewer actions than the median tested human to get there. ARC Prize's own account on X was similarly direct about the result: Astra "achieves SOTA on ARC-AGI" and "builds the most precise symbolic model of novel environments we've seen."

But the headline-grabbing 99.9% is not comparable to any other model's score. No rival has had access to the same adapter. As The New Stack reported, the system that actually hit the near-perfect score was not the same setup OpenAI sells to developers. That makes the figure something closer to a research result than a product specification. Techmeme separately listed ARC Prize's standard-versus-provider-adapter split when the numbers first went public: a 37-point gap between two tests of the identical model, with only one of them replicable by anyone outside OpenAI.

Frankly, the 62.7% is the number that matters if you're trying to track real progress, because it's the one every model gets held to equally.

A fast-moving leaderboard, and a cheaper rival already closing in #

The leaderboard didn't sit still for long. ARC Prize reported that GPT-6.1 Sol, a later OpenAI release, scored 52.7% on the standard harness. On its own provider adapter it scored 96.4%, nearly matching Astra's adapter score - while costing 77% less per run, according to ARC Prize's figures posted on X. That detail matters more than it might look: it suggests the gains behind Astra are becoming cheaper to reproduce, not rarer.

Separately, the Kaggle-hosted ARC-AGI-3 agent competition, which runs independently of the frontier-model leaderboard and rewards open-source solutions, tells a similar story of rapid, unglamorous progress rather than a single breakthrough. Milestone 1, which closed June 30, was won by Tufa Labs' open-source "Duck" harness, built on a comparatively small Qwen model, with a winning score of just 1.21%. Milestone 2, with a September 30 open-source deadline, was won by Daniel Franzen. He adapted Tufa Labs' approach with faster inference, better compute scheduling across games, and improved tools for the agent to track its own actions. His prize-winning score was 27.9%, according to ARC Prize's announcement, while his published Kaggle notebook showed a public V1 best score of 31.47. ARC Prize also posted a newer public high score of 45.33% by Tufa Labs around the same deadline. The pattern is still a sharp jump in open-source methods rather than a single frontier-lab breakthrough.

None of this settles the AGI question, and Chollet has been unusually careful to say so. When pressed on whether solving ARC-AGI-3 would mean a system is AGI, he answered plainly on X: "We're not making this claim." He's pointed out before that ARC-AGI was never meant to be a final exam you pass once and declare victory. It's a tool for measuring the residual gap between what's trivial for humans and still hard for machines. And that gap, on the honest standard-harness numbers, is still open: 62.7% is not 100%, and the levels Astra missed are presumably the ones that resisted its symbolic-modeling approach the hardest.

What's actually new here isn't a number.

Claude Opus 5.5 is beating OpenAI's Astra on coding benchmarks and price Terminal-Bench 4.0 scores put Opus 5.5 at 66.4% against Astra's 57.9%, while Anthropic prices it at $4/$20 per million tokens versus OpenAI's $10/$50, fueling a real-time developer debate on Reddit. - claude opus 5.5 vs openai astra coding benchmarks - best ai model for coding terminal tasks cheaper

It's that the gap between frontier labs and their own benchmark claims has become a visible, contested line, with ARC Prize publishing both sides of it in public rather than letting one favorable figure stand unchallenged.

Also read: AMD stock closes at a record $633.91 as its Nvidia challenge gets real • South Korea orders financial sector security checks after bank data breaches spread • Masayoshi Son warns superintelligence could turn super dangerous in wrong hands

This article is posted in AI News, check it out for more related stories.

Join the discussion #

Open in the community → Almost there. Sign in and your reply posts straight away.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-s-gpt-6-astra…] indexed:0 read:6min 2026-10-04 · —