{"slug": "gpt-6-astra-major-breakthrough-on-arc-agi-3-with-score-of-62", "title": "GPT-6 Astra “Major Breakthrough” On ARC-AGI-3 With Score Of 62%", "summary": "OpenAI's GPT-6 Astra scores 62.7% on the ARC-AGI-3 Semi-Private set using the Standard harness and 99.9% with a Provider Adapter harness, a major jump from GPT-5.6 Sol's 7.8%, according to ARC Prize's published results. François Chollet, ARC Prize co-founder, described the results as a 'step-function change,' noting Astra needed 51.7% fewer actions per level than the human baseline on 96% of levels. The model also showed that higher reasoning effort was cheaper, with max effort costing $26,098 on the Standard harness versus $49,791 at 'none' effort.", "body_md": "[GPT-6-Astra](https://officechai.com/ai/openais-gpt-6-astra-seems-to-comfortably-beat-fable-5-1-in-leaked-benchmarks/) has [barely moved the needle](https://officechai.com/ai/openai-gpt-6-astra-scores-a-disappointing-61-on-artificial-analysis-intelligence-index-same-as-gpt-5-6-sol/) over GPT 5.6 Sol on the Artificial Analysis Intelligence Index, but it does seem to have made a big jump on the ARC-AGI benchmark.\n\nAstra scores 61.2 on the Artificial Analysis Intelligence Index, only marginally ahead of GPT-5.6 Sol’s 60.9, a gap so small it’s barely worth writing home about for a model OpenAI has been billing as a generational leap. ARC-AGI-3 tells a very different story. According to ARC Prize’s own published results, Astra scores 62.7% on the Semi-Private set using the organization’s Standard harness, and a near-saturating 99.9% when switched to a Provider Adapter harness that lets the model use OpenAI’s own context-management tools between turns.\n\nThat’s worth pausing on, because it’s a big jump from where things stood a few weeks ago. We covered how [GPT-5.6 Sol topped ARC-AGI-3 with just 7.8%](https://officechai.com/ai/gpt-5-6-sol-tops-arc-agi-3-with-7-8-becomes-first-model-to-make-meaningful-progress-on-benchmark/), and later how that [score tripled after OpenAI flipped on two API settings](https://officechai.com/ai/openai-says-gpt-5-6s-score-on-arc-agi-3-tripled-after-turning-on-two-api-settings/) that let the model preserve its own reasoning between turns instead of throwing it away after every move. Astra appears to have inherited and extended exactly that advantage, except now it’s baked into the model rather than requiring OpenAI to go digging for an explanation after the fact.\n\nThere’s a numbers mismatch worth flagging before going further. ARC Prize’s own writeup puts Astra’s best Standard-harness score at 62.7%, and its best Provider Adapter score at 99.9% — both achieved at “high” or “max” reasoning effort, not uniformly across the board. But François Chollet, ARC Prize’s co-founder, described the results in a separate post as 66% on the standard harness and “nearly 100%” with what he called a continuous conversation harness with custom compaction, at a cost of roughly $360 per game. Where the extra few points come from isn’t spelled out, and it’s the kind of discrepancy that’s easy to miss if you only read the headline number. Either way, both figures represent a genuine jump over Sol.\n\nThe cost curve is the more counterintuitive part of the story. Across every reasoning-effort setting ARC Prize tested, “max” effort was actually the cheapest way to get a good score:\n\n| Reasoning effort | Standard harness | Provider Adapter harness |\n|---|---|---|\n| max | 62.7%, $26,098 | 98.6%, $17,332 |\n| xhigh | 59.3%, $37,317 | 98.4%, $18,147 |\n| high | 54.8%, $40,705 | 99.9%, $18,817 |\n| medium | 38.6%, $48,090 | 98.4%, $19,285 |\n| low | 17.5%, $38,166 | 98.0%, $21,298 |\n| none | 35.2%, $49,791 | 96.7%, $23,457 |\n\nThe explanation is that a model reasoning at higher effort solves each game in fewer moves, so even though each individual action costs more to generate, the total number of model calls and tokens drops enough to bring the overall bill down. Cranking reasoning up all the way, in other words, isn’t just more accurate here — it’s cheaper.\n\nAction efficiency is where ARC Prize spends most of its analysis, and it’s the number that seems to have gotten Chollet excited enough to call this a “step-function change.” Before ARC-AGI-3 launched, the organization ran roughly 500 members of the general public through the same environments to build a human baseline, timing how many actions it took an average person to solve each level. On the Provider Adapter harness, Astra at max effort needed fewer actions than that human baseline on 96% of levels, using 51.7% fewer actions per level on average. ARC Prize had gone in assuming this was exactly the gap that would keep separating humans from AI for a while — a model might eventually solve a novel puzzle, the thinking went, but it would flail around doing so. That assumption didn’t hold up against Astra.\n\nWhat seems to be driving that efficiency, per ARC Prize’s replay analysis, is that Astra invents its own compressed shorthand to track a game’s state as it plays. Rather than writing out full sentences about what it’s observed, it builds something closer to an ad hoc algebra — tagging object positions, control mappings, and multi-step plans in a dense, code-like notation it comes up with on the fly for each new environment. ARC Prize is careful to describe this as an emergent shorthand rather than a genuine programming language, but notes the notation was unusually precise and information-dense compared to similar behavior it’s observed in other models.\n\nIn a separate, more permissive setup called PRO-LONG, where Astra had access to a code execution sandbox, the model went a step further and started writing actual software for itself mid-game — board parsers, planners, and small game-specific solver scripts built and refined as it played. ARC Prize is upfront that this isn’t an apples-to-apples comparison with its human baseline, since the human testers didn’t have a code interpreter or scratchpad to work with, so PRO-LONG numbers reflect what the model plus its tools can do together, not the model in isolation.\n\nNone of this amounts to ARC Prize declaring AGI has arrived — the organization has said clearly that saturating this benchmark was never meant to be read as proof of that, and ARC-AGI-3’s environments, however novel, are still closed, deterministic puzzle worlds rather than anything resembling the messiness of the real world. But between the score jump, the cost curve, and a model that appears to be absorbing harness-level tricks into its own weights, it’s a more concrete data point than most of what’s circulated since Astra’s launch.", "url": "https://wpnews.pro/news/gpt-6-astra-major-breakthrough-on-arc-agi-3-with-score-of-62", "canonical_source": "https://officechai.com/ai/gpt-6-astra-major-breakthrough-on-arc-agi-3-with-score-of-62/", "published_at": "2026-09-03 20:12:51+00:00", "updated_at": "2026-09-03 20:52:17.402618+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-products"], "entities": ["OpenAI", "GPT-6 Astra", "ARC-AGI-3", "ARC Prize", "François Chollet", "GPT-5.6 Sol", "Artificial Analysis Intelligence Index"], "alternates": {"html": "https://wpnews.pro/news/gpt-6-astra-major-breakthrough-on-arc-agi-3-with-score-of-62", "markdown": "https://wpnews.pro/news/gpt-6-astra-major-breakthrough-on-arc-agi-3-with-score-of-62.md", "text": "https://wpnews.pro/news/gpt-6-astra-major-breakthrough-on-arc-agi-3-with-score-of-62.txt", "jsonld": "https://wpnews.pro/news/gpt-6-astra-major-breakthrough-on-arc-agi-3-with-score-of-62.jsonld"}}