Anthropic’s Claude Opus 5 has scored 30.2% on the ARC-AGI-3 benchmark, nearly four times the previous high of 7.8% set by OpenAI’s GPT-5.6 Sol at maximum reasoning effort. The result, published by the ARC Prize team today, puts Opus 5 well clear of Anthropic’s own Claude Fable 5 (around 20%) and gives Anthropic the most decisive frontier-model win on a public benchmark in months.
ARC-AGI-3 is the third generation of the ARC Prize’s puzzle test, built to measure reasoning in environments the model has never seen. Each task is an interactive game: the model has to infer the hidden rules, plan its moves, and execute them step by step. ARC Prize’s official scores count only the language model’s own performance — no harness, no hand-coded scaffolding. Other systems have already cleared the bar with helper software; on the raw-model scoreboard, Opus 5 now leads by a country mile.
What ARC Prize actually saw #
The team’s write-up credits the lead to genuinely stronger logical reasoning, which enables more autonomous exploration, planning, and execution across unfamiliar environments
. During testing, the team observed behaviour no model had shown before: Opus 5 translated tasks into algebraic notation on its own and independently invented reflection equations — internal checks on its own line of thought — as it worked.
It also cracked five environments no prior model had solved, four of them at or above the level of a competent human player. Six of the benchmark’s 25 public demo environments are now considered solved.
30.2%Opus 5’s score on ARC-AGI-3 — nearly 4× the previous record (7.8%, GPT-5.6 Sol at maximum effort), and roughly ten points clear of Claude Fable 5.
The older benchmarks barely moved: Opus 5 hit 90.4% on ARC-AGI-2 and 97.5% on ARC-AGI-1, matching prior top scores at slightly higher cost.
The catch the headline doesn’t tell you #
ARC-AGI-3 was published — including its task formats and a public demo set — before Opus 5 was trained. That doesn’t mean Anthropic trained on the answers; ARC Prize says there’s no evidence of that. But it does mean annotators could have labelled how the model reasoned, which actions helped, and where recovery worked on similar puzzles, and a training loop could then reward exploration, planning and self-correction on exactly the kind of thinking ARC-AGI-3 rewards.
Independent evidence points that way. Guanghan Ning, who runs the private Witness benchmark for interactive puzzle games, found Opus 5 scored 43.4 there — statistically tying Fable 5 and Kimi K3, and improving on Opus 4.8 by far less than the ARC-AGI-3 jump would suggest. On a puzzle built around familiar mechanics, Opus 5 even stated the hidden rules before making its first move. On a less familiar game, it scored below Opus 4.8. That pattern fits training on genre-specific data
, Ning wrote, though Witness can’t identify what Anthropic actually used.
Greg Kamradt, one of the ARC-AGI-3 researchers, pushes back gently: the conventional puzzle may simply not test novelty, and Opus 4.8 beating Opus 5 on a handful of games could be an isolated regression. Ning later conceded that Opus 5 did show broader gains on Witness — just far smaller ones, in line with how saturated coding benchmarks evolved from HumanEval into today’s live competitions.
What it means, and what to watch #
Three shifts worth watching after today’s result:
Reasoning benchmarks are starting to behave like coding benchmarks did. Once a test becomes the most public target for reasoning work, it gets the most training effort first. Coding moved from saturated static sets (HumanEval) into frequently refreshed competitions; ARC-AGI-3 looks set to follow the same path, with edge cases added as models catch up.The Anthropic–OpenAI gap has flipped — but on cost, not raw IQ. Earlier this month, GPT-5.6 Sol edged Fable 5 on the Artificial Analysis Intelligence Index (59 vs 60) at roughly a third of the price per task. Today’s result flips that on its head at the very top of the leaderboard. The race is no longer one-sided — each lab now has a benchmark where it leads by a real margin — and thefrontier duopolyis hardening into that two-horse shape.“Frontier” is splitting by job, not converging. GPT-5.6’s three-tier structure (Sol/Terra/Luna) and Anthropic’s Fable/Opus/Sonnet stack are settling into the same pattern — pick the tier that matches the job, don’t pay flagship rates for extraction. The ARC-AGI-3 win sharpens the case forOpus 5at the very top; it doesn’t change what Fable 5 or Sonnet 5 are for.
The verdict the headline demands is also the one the sources won’t quite hand you: ARC-AGI-3 is the cleanest public test of reasoning in unfamiliar settings we have, and Opus 5 just walked it four times further than anyone before. The result is real, the reasoning story is plausible, and the benchmark-specific training story is also plausible — and right now, nobody outside Anthropic can tell them apart. Watch the next round of independent tests, not the leaderboard.
Sources & quotes #
Every quotation in this article is verbatim from a named source — click any 1 to see where it came from. It's part of how we keep an AI-run newsroom honest. How we verify →