- Claude Opus 5 scored 30.16% on ARC-AGI-3 at high reasoning effort, versus the previous official record of 7.78% for GPT-5.6 Sol at maximum effort. [1] - Anthropic released Opus 5 publicly on July 24 across its products and API, priced at $5 per million input tokens and $25 per million output tokens. [2] - The score measures action efficiency against humans across unfamiliar interactive games; it is not the percentage of games solved or proof of general intelligence.
[3] Anthropic’s Claude Opus 5 has set a verified record on ARC-AGI-3, scoring 30.16% at high reasoning effort. The previous official high was GPT-5.6 Sol’s 7.78% at maximum effort, while Claude Opus 4.8 scored 1.52% under the same high-effort setting used for Opus 5.[1]
The result is supported by ARC Prize and Anthropic’s system card. Opus 5 is also publicly available: Anthropic released it July 24 on all its platforms and through the Claude API. Its standard API price is $5 per million input tokens and $25 per million output tokens.[2]
What the score does — and does not — show #
ARC-AGI-3 puts models into unfamiliar turn-based games without rules or stated goals. They must explore, infer the objective and complete progressively harder levels. Its Relative Human Action Efficiency score rewards both completion and economical action use, with runs capped at five times the median human action count for each level.[3]
Anthropic’s system card says the reported Opus 5 result used ARC Prize’s 55-environment semi-private set, high reasoning effort and the benchmark’s standard setup. Anthropic’s general evaluation note says results are averaged across five trials unless stated otherwise; the ARC-AGI-3 section gives no exception. Its chart places total evaluation cost slightly above $20,000 but does not publish an exact dollar amount.[1][4]
ARC Prize said Opus 5 completed five public demonstration environments no earlier model had beaten. In one replay, it converted a reflection puzzle into algebra and completed eight levels in 294 actions. François Chollet, who created ARC-AGI, called the 30% result an “impressive jump.” Still, Opus 5 scored below one-third of the human-normalized maximum, and the verified documentation does not provide a directly comparable Fable 5 score. The evidence therefore supports a benchmark record, not the broader claim that Opus 5 has surpassed Fable 5 or demonstrated “real intelligence.”[1][4][5]
Companies mentioned #
Further sources #
[[1] ARC Prize, “Claude Opus 5 — ARC-AGI Results,” July 24, 2026. ↗](https://arcprize.org/results/anthropic-claude-opus-5)
[[2] Anthropic, “Introducing Claude Opus 5,” July 24, 2026. ↗](https://www.anthropic.com/news/claude-opus-5)
[[3] ARC Prize Foundation, “ARC-AGI-3: A New Challenge for Frontier Agentic Intellig… ↗](https://arcprize.org/media/ARC_AGI_3_Technical_Report.pdf)
[4] Anthropic, “Claude Opus 5 System Card,” July 24, 2026. ↗
[5] François Chollet, comment on Claude Opus 5’s ARC-AGI-3 result, July 24, 2026. ↗ The stories that matter, in one email. Free — unsubscribe anytime.