Claude Opus 5 Benchmarks: What the Numbers Actually Show Anthropic shipped Claude Opus 5 on July 24, posting 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10-point jump with no change in per-token price. The model also shows gains on internal life sciences benchmarks, scoring 10.2 percentage points above Opus 4.8 on organic chemistry and 7.7 points higher on protein tasks. A developer noted that Anthropic published most gains as ratios rather than absolute scores, making independent verification difficult. Opus 5 posts 79.2 percent on SWE-bench Pro against Opus 4.8 at 69.2, a 10 point jump with no change in per-token price Anthropic published most gains as ratios three times ARC-AGI-3, more than double Frontier-Bench rather than absolute scores On CursorBench 3.2 at max effort it lands within 0.5 percent of Fable 5's peak at half the cost per task Public GDPval-AA figures disagree across sources by up to 117 Elo, so I left that row out entirely Anthropic shipped Claude Opus 5 on July 24, and the coverage filled up with ratios instead of scores. Three times the next-best model. More than double the previous Opus. Just over a third of the cost. I went looking for the actual numbers behind those phrases. What I found says as much about how model launches get reported as it does about the model. The cleanest row is SWE-bench Pro, which runs a model against real GitHub issues and checks whether the patch passes the repository's own tests. It is harder than the older SWE-bench Verified set and it is the row the whole industry now quotes. Model SWE-bench Pro Released Claude Fable 5 80.3 2026-06-09 Claude Opus 5 79.2 2026-07-24 Claude Opus 4.8 69.2 2026-05-29 GPT-5.6 Sol 64.6 2026-07-09 That is a 10 point jump from Opus 4.8 to Opus 5 inside two months, and the per-token price did not move both tiers run at 5 and 25 per million tokens . Fable 5 keeps a 1.1 point lead and charges double for it. Those two facts together are the actual story of this release, and neither one is a ratio. On SWE-bench Verified, the older and easier set, Opus 5 reports 96.0 percent averaged over five trials. The averaging matters. A single run on a set that saturated above 90 percent tells you very little, because the spread between runs starts to rival the gap between models. Five trials is better practice than most launch tables bother with, and it is worth noticing when a lab does it. It is worth being precise about why those two rows behave differently, because they get quoted interchangeably and they should not be. SWE-bench Verified is a human-filtered subset of issues confirmed to be solvable, which made it a fair test in 2024 and makes it close to useless as a discriminator now that four models sit above 88 percent on it. Once a benchmark saturates, the remaining points are mostly measuring flaky tests and formatting luck. SWE-bench Pro was built to restore headroom, with longer issues, larger repositories, and less forgiving grading. When a lab leads with Verified rather than Pro, that choice is itself information. The same logic applies to the older rows in my own archive. Opus 4.8 scored 88.6 on Verified and 69.2 on Pro. The Verified gap between Opus 4.8 and Opus 5 is 7.4 points; the Pro gap is 10. The harder benchmark shows the bigger difference, which is the pattern you want to see if the improvement is real rather than an artifact of a test running out of room. The life sciences numbers are the other place Anthropic gave real figures. Opus 5 scores 10.2 percentage points above Opus 4.8 on an internal organic chemistry benchmark and 7.7 points higher on protein tasks. Both are internal, both are stated as point deltas against a named baseline, and that is a perfectly honest way to publish a private benchmark. Most of the launch page is written in relative terms. Frontier-Bench v0.1: more than doubles Opus 4.8 at a lower cost per task. ARC-AGI-3: three times the next-best model. Zapier AutomationBench: roughly 1.5 times the pass rate of the next-best model at the same cost per task. OSWorld 2.0: beats Fable 5's best result at just over a third of the cost. CursorBench 3.2: within 0.5 percent of Fable 5's peak at half the cost per task. None of that is dishonest. All of it is hard to check. A ratio without its baseline cannot be falsified, and "the next-best model" is a moving target that changes meaning the week someone else ships. Third parties filled the gap with absolute numbers Frontier-Bench 43.3 for Opus 5 against 34.4 for GPT-5.6 Sol, ARC-AGI-3 30.2 against 7.8 , and those line up with the ratios, so the framing appears to be presentation rather than spin. Two of those claims are worth pulling apart because they are doing something unusual. The Zapier AutomationBench result is quoted as roughly 1.5 times the pass rate of the next-best model at the same cost per task , and the OSWorld 2.0 result as beating Fable 5 at just over a third of the cost. Both hold price fixed and let capability vary, or hold capability fixed and let price vary. That is a more useful shape than a raw score for anyone actually paying the bill, and it is also a shape that only looks good if your token efficiency improved. A lab whose model got smarter but chattier cannot publish this chart. The more interesting choice is cost per task. Anthropic anchored nearly every headline claim to it rather than to the score alone. That is a legitimate metric, and for anyone running a model inside an agent loop it is closer to the thing that actually hurts than a leaderboard position is. It is also a flattering metric for a model that got cheaper to run at the same sticker price, because it bundles token efficiency into the result. Both things are true at once. I would rather have it than not, and I would not treat it as interchangeable with capability. The gaps are the part launch coverage skips, so here they are. On general reasoning the two current frontier models are close enough to call even, with third-party aggregates putting GPT-5.6 Sol at 92.5 average against 90.4 for Opus 5. Sol also keeps DeepSWE 1.1 and HealthBench Professional. Sol's Terminal-Bench 2.1 result of 91.9 percent in its top mode still stands as the highest agentic-terminal figure anyone has published. If your workload looks like that benchmark, the leaderboard did not change for you last week. Then there is GDPval-AA, the human-graded knowledge work benchmark, where I could not get the sources to agree. One set of tables puts Fable 5 at 1,932 Elo, another at 1,815. Opus 4.8 appears at 1,890 in the same tables that put Opus 5 at 1,861, which would mean the new model went backwards. Different benchmark versions and different effort tiers are almost certainly being compared as if they were the same column. I could have picked the number that made the cleanest paragraph. Instead I am leaving the row out and saying why, because a 117 point disagreement between sources is not a rounding difference, it is a sign nobody has published a like-for-like run yet. Four checks, in order, before a number goes anywhere near a decision. Which effort level produced it. Opus 5 runs five levels from low to max, and Anthropic's own CursorBench claim is explicitly a max effort result. Most third-party tables do not state the setting at all, which makes them roughly unusable for cost comparison even when the score is real. How many trials. One run on a saturated benchmark is noise. The five-trial SWE-bench Verified figure is the exception, not the norm. Who ran it. Vendor benchmarks are marketing artifacts written by people with access to the model and an interest in the outcome. That does not make them wrong. It makes them a starting hypothesis. Whether the comparison model is current. A chart that beats a model retired two releases ago is measuring the calendar. I wrote up the previous generation of this exact question in Claude Fable 5 vs Opus 4.8: Is Double the Price Worth It https://dev.to/blogs/lab/claude-fable-5-vs-opus-4-8-is-double-the-price-worth-it , and half of that comparison is already historical. The honest summary of Opus 5's scoreboard is narrower than the headlines and better than a routine bump: a 10 point SWE-bench Pro gain over Opus 4.8 at an unchanged sticker price, near-parity with a model that costs twice as much, a genuine outlier result on ARC-AGI-3, and a handful of rows where OpenAI still leads. The ratios are probably accurate. They are just not the same kind of claim as a number. What I would not do is treat any of this as settled. Three labs shipped flagship models in the fifteen days before I wrote this, the GDPval column is incoherent across sources, and nobody has published an independent like-for-like run at a stated effort level. For the full spec breakdown of what shipped, see Claude Opus 4.8 Is Here: Everything That Changed https://dev.to/blogs/lab/claude-opus-4-8-is-here-everything-that-changed for the previous release in this line, and the rest of the model coverage is indexed on the Lab overview https://dev.to/pages/lab-overview . I will run these models against my own work for a few weeks and write up what actually held.