cd /news/large-language-models/opus-5-5-costs-37-of-opus-5-per-task · home topics large-language-models article
[ARTICLE · art-138740] src=thenewway.ai ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Opus 5.5 costs 37% of Opus 5 per task

Two independent benchmarks put Anthropic's Claude Opus 5.5 at 37 to 38% of Claude Opus 5's cost per task at each model's default effort setting, beating Anthropic's own claim of "40% less than Opus 5 on typical workloads." Artificial Analysis measured Opus 5.5 at $1.34 per task on its Intelligence Index versus $3.61 for Opus 5 (scores 51 and 48), while Paweł Huryn's Bug Hunt Bench found Opus 5.5 fixed 30.3 of 105 planted bugs for $15.68 against Opus 5's 21 bugs for $41.51. At max effort the saving disappears: Artificial Analysis has Opus 5.5 at $5.98 per task versus $5.86 for Opus 5, 2% more, because Opus 5.5 writes about 119,000 output tokens per task where Opus 5 wrote about 73,000.

read5 min views3 publishedSep 24, 2026
Opus 5.5 costs 37% of Opus 5 per task
Image: Thenewway (auto-discovered)

At each model's default setting, two independent benchmarks put Opus 5.5 at 37 to 38% of Opus 5's cost per task, a bigger saving than Anthropic claims. At max effort it costs more.

Anthropic says Claude Opus 5.5 costs 40% less to run than Opus 5. OpenAI says GPT-6 Sol and Luna cost half what GPT-5.6 did. Two independent measurements now cover every effort setting, the dial that sets how long a model thinks. Artificial Analysis, an independent benchmarking firm, runs one; Paweł Huryn runs the other, Bug Hunt Bench, and publishes its raw data. The setting decides whether either claim holds. At defaults, Anthropic's saving is bigger than it says. At max, it disappears. GPT-6 Sol costs about half at every setting, but on real code it fixes fewer bugs. This newsletter's first two reads compared the wrong settings, and this is the corrected one.

Opus 5.5 defaults to medium effort, Opus 5 to high

This newsletter's Tuesday post and Wednesday issue compared Opus 5.5 with Opus 5 at max effort, the highest setting, where Artificial Analysis found the two about level per task. Anthropic's 40% claim is about default settings, and the two models don't share one. Anthropic's docs: "Claude Opus 5.5 defaults to medium." Other Claude models, Opus 5 included, default to high. Artificial Analysis published every effort level on launch day, so the default comparison was there to make. Both earlier posts now carry corrections.

At default settings, Opus 5.5 costs 37% of Opus 5 per task

Set each model to its default and the numbers beat Anthropic's pitch. Artificial Analysis measures Opus 5.5 at $1.34 per task on its Intelligence Index, against $3.61 for Opus 5, 37% of the cost, and scores it higher, 51 to 48. Bug Hunt Bench plants 105 real bugs in two production codebases, and an independent judge grades each model's fixes without knowing which model made them. There, Opus 5.5 fixed 30.3 bugs for $15.68, averaged over three runs; Opus 5 fixed 21 for $41.51 in one run: 38% of the cost, and nine more bugs. Anthropic's own claim is "40% less than Opus 5 on typical workloads." Both independent reads land above it.

| Default vs default | Opus 5 (high) | Opus 5.5 (medium) | 
|---|---|---|
| Artificial Analysis, cost per task | $3.61 (score 48) | $1.34 (score 51) | 
| Bug Hunt Bench, run cost | $41.51 (21 fixed) | $15.68 (30.3 fixed) | 

At max effort, Opus 5.5 costs more than Opus 5

Max is the one setting where the saving disappears. Artificial Analysis has Opus 5.5 at $5.98 per task on its Intelligence Index against $5.86, 2% more, because it writes about 119,000 output tokens per task where Opus 5 wrote about 73,000. It also scores 58 to 51, the highest Artificial Analysis has measured. On Bug Hunt Bench, max-effort Opus 5.5 fixed 41.7 bugs for $58.53, and Opus 5 fixed 27 for $54.94. That run cost Opus 5.5 7% more. On Artificial Analysis's coding-agent index, run in Claude Code at max effort, the gap is wider: $13.04 a task against $10.79, 21% more. Moving Opus 5.5 from medium to max multiplies its cost per task by about 4.5 on Artificial Analysis's Intelligence Index. The price cut doesn't make that choice for you.

Artificial Analysis finds GPT-6 Sol half the cost at every setting

GPT-6 Sol and Luna default to medium, as GPT-5.6 did, so this comparison is like for like. At every effort level Artificial Analysis measures GPT-6 Sol at 47 to 55% less per task than GPT-5.6 Sol, with scores within a point. At medium it's $0.25 against $0.50. Its explanation: the saving "is driven by the price cut, as both models use slightly more output tokens per task." OpenAI measures its "50% cheaper" against what it calls GPT-5.6's "promotional pricing." Per task, on this index, it holds.

On real code, GPT-6 Sol fixes fewer bugs at every setting tested

Bug Hunt Bench disagrees about what the saving buys. At the four settings where both were tested, GPT-6 Sol runs 85 to 90% cheaper than GPT-5.6 Sol and fixes 14 or 15 fewer bugs. At medium, the default, it fixed 14 for $2.39, where GPT-5.6 Sol fixed 29 for $15.77. One author runs the bench, the rows below max are single runs, and three max runs of GPT-6 Sol scored 32, 30 and 26. The OpenAI runs used Codex, OpenAI's coding agent, on a ChatGPT account, so their costs are token counts priced at list rates, not bills. Artificial Analysis's coding index has Sol up two points. If a second tester reproduces the drop, the half-price default comes with half the fixes.

Opus 5.5 at medium matches Opus 5 at max for 77% less

The effort setting moves the bill more than the price cut does. On Artificial Analysis's index, Opus 5.5 at its default scores 51, the same as Opus 5 at max, for $1.34 a task against $5.86. Raising Opus 5.5 from medium to max adds $4.64 a task, about twice the gap between the two models at their defaults. The API uses the default when you leave effort unset: Anthropic's docs say setting it to the default "produces exactly the same behavior as omitting the effort parameter entirely." If you raised effort on Opus 5, test Opus 5.5 at medium before you keep it.

Know someone who'd want this in their inbox? Forward it — that's how this grows. And if we got something wrong, or you think we buried the real story today, hit reply. A person reads every one.

The New Way is written with AI. It gathers the day's stories, checks them against their sources and drafts every summary. A person decides what runs and reviews every issue before we hit send.

── more in #large-language-models 4 stories · sorted by recency
── more on @anthropic 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/opus-5-5-costs-37-of…] indexed:0 read:5min 2026-09-24 ·