ARC-AGI 3 had been notoriously slow to move since it was introduced, but Claude Opus 5 has now created a 3x jump on the benchmark.
Anthropic’s newly released Claude Opus 5 has posted a 30.2% score on ARC-AGI-3, more than tripling the best result any model had managed on the benchmark before it. For a test that’s been specifically designed to resist the usual tricks AI labs use to inflate their numbers, that’s a genuinely large jump, and it’s the single most striking figure in Anthropic’s entire Opus 5 launch.
What ARC-AGI actually tests #
ARC-AGI was created by François Chollet as an attempt to measure something closer to fluid intelligence than the benchmarks most labs optimize for. Where a typical AI benchmark rewards a model for having seen something similar during training, ARC-AGI tasks are built so that memorization doesn’t help. A model has to look at a small number of examples, work out the underlying rule connecting them, and apply that rule to a new situation it hasn’t encountered before. It’s the kind of test a person can usually solve by inspection, and one that’s historically been brutal for even the most capable AI systems.
The benchmark has gone through several versions as earlier ones got solved. ARC-AGI-1 is essentially saturated at this point, with top models scoring in the high nineties. ARC-AGI-2 held out longer, but that fell too — Gemini 3.1 Pro crossed 77% on it earlier this year, and Gemini 3 Deep Think pushed that to 84.6%, numbers that made the second version look close to solved as well.
ARC-AGI-3 is a different kind of test altogether. Instead of static grid puzzles, it drops an AI agent into interactive, turn-based environments with no instructions, and the agent has to work out the goals, the rules, and a strategy purely through trial and error. It’s scored using a metric called Relative Human Action Efficiency, which compares how many actions a model needs to clear a level against how many the second-best human took on the same level. When the ARC Prize Foundation released ARC-AGI-3 in March, the results were about as stark as a benchmark gets: humans cleared 100% of the environments, and the best AI model at the time managed 0.37%.
How the field has moved since March #
Progress on ARC-AGI-3 since its release has been slow by the standards of recent AI benchmarks, which is precisely the point of the test. By the time Opus 5 launched, the strongest reported score belonged to GPT-5.6 Sol at max reasoning effort, sitting at 7.8%. Opus 4.8, Anthropic’s previous flagship, barely moved the needle at 1.5%.
Against that backdrop, Opus 5’s 30.2% at high effort is a step-change rather than an incremental gain. It isn’t just ahead of GPT-5.6 Sol, it’s roughly four times its score, and it does this while every other model on Anthropic’s own comparison chart is still bunched near the bottom of the chart’s cost-versus-score curve. Anthropic’s framing of “three times the next-best model” is if anything the more conservative way to describe the gap.
What makes this particular jump worth paying attention to is what ARC-AGI-3 is built to filter out. A model can’t crack these environments by pattern-matching against something it saw in training, because the environments are new and the rules have to be inferred live. A 30% score means Opus 5 is genuinely working out unfamiliar rule systems through interaction, not recalling something adjacent to what it has already seen. Benchmarks like SWE-Bench or GDPval reward a model for doing familiar categories of work faster and cheaper. ARC-AGI-3 rewards a model for handling something it has never been trained to handle at all, and that’s a much harder property to fake.
None of this means the test is close to solved. Humans are still clearing every environment ARC-AGI-3 throws at them, and a 30% score, however large a jump it represents, leaves plenty of room before anyone starts talking about saturation the way they now do for ARC-AGI-1 and ARC-AGI-2. But going from under 8% to over 30% in one release is the kind of move that resets expectations for how fast the rest of the field will need to close the gap.