cd /news/artificial-intelligence/openai-claims-gpt-5-6-sol-beats-opus… · home topics artificial-intelligence article
[ARTICLE · art-80483] src=mlq.ai ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

OpenAI Claims GPT-5.6 Sol Beats Opus 5 on ARC-AGI-3—but Only With Non-Standard API Settings

OpenAI claims its GPT-5.6 Sol model scored 38.3% on the ARC-AGI-3 public set with two API settings enabled—retained reasoning and compaction—tripling its baseline score and surpassing Anthropic's Claude Opus 5 at 30.2%, but in the official ARC Prize test harness Sol scored just 7.8% versus Opus 5's 30.2%. The dispute highlights a growing tension in AI evaluation as frontier models become increasingly dependent on inference-time scaffolding, with ARC Prize co-founder François Chollet acknowledging a 'potential parity issue' when providers use different API settings.

read5 min views1 publishedJul 30, 2026
OpenAI Claims GPT-5.6 Sol Beats Opus 5 on ARC-AGI-3—but Only With Non-Standard API Settings
Image: Mlq (auto-discovered)
  • OpenAI says GPT-5.6 Sol scores 38.3% on the ARC-AGI-3 public set with two API settings enabled—retained reasoning and compaction—tripling its baseline score and surpassing Opus 5's 30.2%
[[1]](https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/) - In the official ARC Prize test harness, GPT-5.6 Sol scored just 7.8%, well below Anthropic's Claude Opus 5 at 30.2%
[[2]](https://arcprize.org/results/openai-gpt-5-6-sol) - ARC Prize said its official scores use a 'standardized approach without provider-specific settings to ensure fair comparisons'

[3] - Co-founder François Chollet acknowledged a 'potential parity issue' when providers use different API settings but said general-purpose features not built for ARC-AGI-3 are acceptable [3] - The average human tester scored 48% on ARC-AGI-3, meaning both models remain well below human performance

[1] OpenAI on July 30 published a blog post arguing that its GPT-5.6 Sol model outperforms Anthropic's Claude Opus 5 on the ARC-AGI-3 benchmark—but only when tested through OpenAI's own Responses API with two additional settings enabled. With those features turned on, Sol scored 38.3% on the public task set, compared with Opus 5's 30.2%. In the official ARC Prize test harness, Sol managed just 7.8% [1] [2].

The two settings—'retained reasoning,' which preserves the model's chain of thought between steps, and 'compaction,' which summarizes prior context instead of discarding it—tripled Sol's score and reduced output tokens sixfold, according to OpenAI. The company argued that 'benchmarks never measure just the model but also the technical setup around it' [1].

The dispute highlights a growing tension in AI evaluation: as frontier models become increasingly dependent on inference-time scaffolding, standardized benchmarks struggle to accommodate provider-specific API features without compromising comparability. ARC-AGI-3, designed by François Chollet and the ARC Prize team to measure general reasoning ability, has become a marquee benchmark in the competition between OpenAI and Anthropic [3].

The Scores #

Under the official ARC Prize evaluation, GPT-5.6 Sol scored 7.78% on the ARC-AGI-3 public set. That placed it far behind Anthropic's Claude Opus 5, which scored 30.2% in the same standardized environment. The previous-generation GPT-5.5 scored just 0.4%, making Sol a significant step up for OpenAI but still a distant second on the official leaderboard [2] [4].

OpenAI's self-reported results tell a different story. Using the Responses API with retained reasoning and compaction enabled, Sol scored 38.3% on the same public task set—an almost fivefold improvement over its official result and an 8-percentage-point lead over Opus 5. The average human tester, for context, scores 48% on ARC-AGI-3 [1].

The gap stems from how the official harness handles reasoning. In the standardized setup, GPT-5.6 Sol's internal reasoning is discarded after each action step, and older actions are truncated once the context window fills up. The model retains what it did but loses most of the reasoning behind its decisions [5].

The API Settings at Issue #

The two features OpenAI highlighted are general-purpose capabilities of its Responses API, not tools purpose-built for ARC-AGI-3. Retained reasoning keeps a model's chain-of-thought intact across multi-step interactions, while compaction compresses older context into summaries rather than hard-truncating it [1].

OpenAI framed these as standard infrastructure that any API customer can access, arguing they represent how the model is intended to be used in production. The company's implicit case is that benchmarking Sol without these features is akin to testing a car with the parking brake engaged [1].

The distinction matters because Anthropic's Claude API already handles context management differently, and Opus 5's 30.2% score was achieved under the official harness. If the standardized environment inadvertently favors one provider's API architecture over another, the leaderboard comparison becomes unreliable [3].

ARC Prize Responds #

ARC Prize defended its methodology, stating that the official scores employ a 'standardized approach without provider-specific settings to ensure fair comparisons' [3].

However, co-founder François Chollet offered a more nuanced position. He said general-purpose API settings 'that were not developed for ARC-AGI-3 and that are available to all API users' are acceptable for reporting alongside official scores. He acknowledged that different providers using different settings creates 'a potential parity issue' but stopped short of changing the official evaluation protocol [3].

Chollet also conceded that ARC Prize may have used an outdated OpenAI completions API that lacked features already available in Anthropic's Claude API, which could have skewed the official comparison against Sol [3].

Why It Matters #

The dispute exposes a structural problem in AI benchmarking that extends well beyond a single test. As frontier models increasingly rely on inference-time compute—chain-of-thought reasoning, tool use, multi-step planning—their performance becomes inseparable from the API infrastructure that orchestrates those capabilities. A benchmark that strips away that scaffolding may undercount a model's effective intelligence; one that permits it risks turning evaluations into system-integration contests.

ARC-AGI-3 was designed specifically to resist the kind of benchmark gaming that plagued earlier evaluations. It tests abstract reasoning on novel visual puzzles, and its difficulty is calibrated so that current AI systems score well below humans. GPT-5.6 Sol scored 92.5% on the older ARC-AGI-2 but collapsed to single digits on ARC-AGI-3 under official conditions, illustrating how much harder the new benchmark is [4].

For investors and enterprise buyers comparing OpenAI and Anthropic, the episode is a reminder that headline benchmark numbers require scrutiny of testing methodology. A 38.3% score and a 7.8% score describe the same model on the same test—the only difference is how the API manages the model's memory.

What's Next #

The ball is now in ARC Prize's court to decide whether future evaluations should accommodate provider-specific API features like retained reasoning. Chollet's comments suggest the organization is open to reporting such results alongside official scores, even if the standardized harness remains the primary measurement [3].

OpenAI's blog post also raises the question of whether Anthropic's Opus 5 score would improve with analogous optimizations on its own API. Until both models are tested under equivalent conditions—whether that means stripping features from both or enabling them for both—the ARC-AGI-3 leaderboard will remain contested.

Companies mentioned #

Further sources #

[[1] OpenAI blog post: 'How enabling two settings tripled our scores on the ARC-AGI-… ↗](https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/)

[[2] ARC Prize official results page for GPT-5.6 Sol ↗](https://arcprize.org/results/openai-gpt-5-6-sol)

[3] The Decoder: 'OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its late… ↗

[[4] ARC Prize official results page for GPT-5.6 Sol (ARC-AGI-2 scores) ↗](https://arcprize.org/results/openai-gpt-5-6-sol)

[[5] Rohan Paul on X, analysis of ARC-AGI-3 methodology and reasoning erasure ↗](https://x.com/rohanpaul_ai/status/2082660329872646465)

The stories that matter, in one email. Free — unsubscribe anytime.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/openai-claims-gpt-5-…] indexed:0 read:5min 2026-07-30 ·