{"slug": "claude-opus-scores-30-on-arc-agi-3-triples-previous-best-score-by-any-model", "title": "Claude Opus Scores 30% On ARC-AGI 3, Triples Previous Best Score By Any Model", "summary": "Anthropic's Claude Opus 5 has posted a 30.2% score on ARC-AGI-3, more than tripling the previous best result of 7.8% held by GPT-5.6 Sol. The benchmark, designed by François Chollet to measure fluid intelligence through interactive turn-based environments, had seen the best AI model score only 0.37% at its March release. The jump is significant because ARC-AGI-3 resists memorization, requiring models to infer rules through live interaction.", "body_md": "ARC-AGI 3 had been notoriously slow to move since it was introduced, but Claude Opus 5 has now created a 3x jump on the benchmark.\n\nAnthropic’s newly released Claude Opus 5 has posted a 30.2% score on ARC-AGI-3, more than tripling the best result any model had managed on the benchmark before it. For a test that’s been specifically designed to resist the usual tricks AI labs use to inflate their numbers, that’s a genuinely large jump, and it’s the single most striking figure in Anthropic’s entire Opus 5 launch.\n\n## What ARC-AGI actually tests\n\nARC-AGI was created by François Chollet as an attempt to measure something closer to fluid intelligence than the benchmarks most labs optimize for. Where a typical AI benchmark rewards a model for having seen something similar during training, ARC-AGI tasks are built so that memorization doesn’t help. A model has to look at a small number of examples, work out the underlying rule connecting them, and apply that rule to a new situation it hasn’t encountered before. It’s the kind of test a person can usually solve by inspection, and one that’s historically been brutal for even the most capable AI systems.\n\nThe benchmark has gone through several versions as earlier ones got solved. [ARC-AGI-1](https://officechai.com/ai/googles-gemini-3-model-makes-2x-jump-over-previous-state-of-the-art-on-arc-agi-prize/) is essentially saturated at this point, with top models scoring in the high nineties. ARC-AGI-2 held out longer, but that fell too — Gemini 3.1 Pro crossed 77% on it earlier this year, and Gemini 3 Deep Think pushed that to 84.6%, numbers that made the second version look close to solved as well.\n\nARC-AGI-3 is a different kind of test altogether. Instead of static grid puzzles, it drops an AI agent into [interactive, turn-based environments](https://officechai.com/ai/agi-likely-by-early-2030s-when-arc-agi-6-or-7-will-be-released-arc-prizes-francois-chollet/) with no instructions, and the agent has to work out the goals, the rules, and a strategy purely through trial and error. It’s scored using a metric called Relative Human Action Efficiency, which compares how many actions a model needs to clear a level against how many the second-best human took on the same level. When the ARC Prize Foundation [released ARC-AGI-3](https://officechai.com/ai/arc-agi-3/) in March, the results were about as stark as a benchmark gets: humans cleared 100% of the environments, and the best AI model at the time managed 0.37%.\n\n## How the field has moved since March\n\nProgress on ARC-AGI-3 since its release has been slow by the standards of recent AI benchmarks, which is precisely the point of the test. By the time Opus 5 launched, the strongest reported score belonged to GPT-5.6 Sol at max reasoning effort, sitting at 7.8%. Opus 4.8, Anthropic’s previous flagship, barely moved the needle at 1.5%.\n\nAgainst that backdrop, Opus 5’s 30.2% at high effort is a step-change rather than an incremental gain. It isn’t just ahead of GPT-5.6 Sol, it’s roughly four times its score, and it does this while every other model on Anthropic’s own comparison chart is still bunched near the bottom of the chart’s cost-versus-score curve. Anthropic’s framing of “three times the next-best model” is if anything the more conservative way to describe the gap.\n\nWhat makes this particular jump worth paying attention to is what ARC-AGI-3 is built to filter out. A model can’t crack these environments by pattern-matching against something it saw in training, because the environments are new and the rules have to be inferred live. A 30% score means Opus 5 is genuinely working out unfamiliar rule systems through interaction, not recalling something adjacent to what it has already seen. Benchmarks like SWE-Bench or GDPval reward a model for doing familiar categories of work faster and cheaper. ARC-AGI-3 rewards a model for handling something it has never been trained to handle at all, and that’s a much harder property to fake.\n\nNone of this means the test is close to solved. Humans are still clearing every environment ARC-AGI-3 throws at them, and a 30% score, however large a jump it represents, leaves plenty of room before anyone starts talking about saturation the way they now do for ARC-AGI-1 and ARC-AGI-2. But going from under 8% to over 30% in one release is the kind of move that resets expectations for how fast the rest of the field will need to close the gap.", "url": "https://wpnews.pro/news/claude-opus-scores-30-on-arc-agi-3-triples-previous-best-score-by-any-model", "canonical_source": "https://officechai.com/ai/claude-opus-5-arc-agi-3/", "published_at": "2026-07-24 17:18:55+00:00", "updated_at": "2026-07-24 17:57:39.980614+00:00", "lang": "en", "topics": ["artificial-intelligence", "ai-research", "ai-products"], "entities": ["Anthropic", "Claude Opus 5", "ARC-AGI-3", "François Chollet", "GPT-5.6 Sol", "ARC Prize Foundation", "Gemini 3.1 Pro", "Gemini 3 Deep Think"], "alternates": {"html": "https://wpnews.pro/news/claude-opus-scores-30-on-arc-agi-3-triples-previous-best-score-by-any-model", "markdown": "https://wpnews.pro/news/claude-opus-scores-30-on-arc-agi-3-triples-previous-best-score-by-any-model.md", "text": "https://wpnews.pro/news/claude-opus-scores-30-on-arc-agi-3-triples-previous-best-score-by-any-model.txt", "jsonld": "https://wpnews.pro/news/claude-opus-scores-30-on-arc-agi-3-triples-previous-best-score-by-any-model.jsonld"}}