Sonnet 5.5 looks cheaper per token than Opus 5.5, but run it at max effort and the cost-per-task math flips. Here's why.
Is Sonnet 5.5 actually cheaper than Opus 5.5? #
Not always. Sonnet 5.5 carries a lower per-token price than Opus 5.5, which is the normal pattern for Anthropic’s lineup: Sonnet positioned as the value tier, Opus as the flagship. But per-token price isn’t the same thing as per-task cost. When Sonnet 5.5 is pushed to its “max effort” setting to compete on benchmarks, it can burn enough extra tokens reasoning through a task that the total bill lands close to, or above, what Opus 5.5 would have charged to do the same job with less thinking.
TL;DR #
- Per-token pricing still favors Sonnet 5.5 over Opus 5.5, but that gap shrinks or disappears once you measure cost per completed task instead of cost per million tokens.
- Max effort mode on Sonnet 5.5 generates significantly more reasoning tokens to hit benchmark-topping scores, and those extra tokens are what get billed.
- Cost-per-task benchmarks , like the ones tracked on aggregators such as Artificial Analysis, showed both Sonnet 5.5 and Opus 5.5 landing in a similarly expensive range, roughly $6 to $8 per task, far above cheaper competitors.
- GPT 6.1 Soul , announced the same week at OpenAI’s dev day, undercuts both Anthropic models dramatically on cost per task, landing around 72 cents, while trailing them by only a few points on blended intelligence benchmarks.
- Benchmark mode and production mode are different products in practice: a model tuned to maximize benchmark score by spending more compute isn’t the same cost profile you’ll see running it at default settings in your own app.
- Buyers should budget by workload , not by the headline per-token rate, because a “cheaper” model set to its most capable mode can erase its own discount.
Why does Sonnet’s advertised price not match real-world cost? #
The sticker price for an API model is quoted per million tokens, split between input and output. That number is accurate and not misleading on its own. The problem shows up when you compare models at the settings each one needs to reach its best benchmark score.
Opus 5.5 is Anthropic’s top-end model. It doesn’t need extended reasoning effort cranked to maximum to post strong numbers; its baseline is already tuned for capability. Sonnet 5.5, to close the gap with Opus on hard benchmarks, often needs to run in its highest reasoning or “max effort” configuration. That setting makes the model think longer and generate more intermediate tokens before producing a final answer. Since output tokens (and extended thinking tokens) are billed, a longer reasoning chain adds up fast, even at a lower per-token rate.
So the math becomes: Sonnet’s rate is cheaper, but the token volume needed to hit a comparable score is higher. Multiply a lower price by a much larger token count, and the total can meet or exceed Opus’s cost for the same output quality.
What did the benchmark data actually show? #
Looking at blended intelligence-versus-cost tracking from aggregator sites like Artificial Analysis, Opus 5.5 and Sonnet 5.5 both clustered in the same expensive territory on a cost-per-task basis, somewhere around $6 to $8 per task. That’s a strange result if you only looked at headline per-token pricing and assumed Sonnet would be meaningfully cheaper to run.
For contrast, OpenAI’s GPT-6 Astra landed near $3.26 per task while scoring close behind Opus 5.5 on the same combined benchmark. And GPT 6.1 Soul, the new lower-cost model OpenAI introduced at its developer event the same week, came in around 72 cents per task, only a handful of points below Astra on the intelligence score and about six points below Opus 5.5. Soul’s list price is $2 per million input tokens and $10 per million output tokens, compared to Astra’s $10 input and $50 output. The pattern across nearly every benchmark mentioned was consistent: the more expensive model scored a bit higher, and the cheaper model landed close behind at a fraction of the cost per completed task. Sonnet 5.5 broke that pattern when pushed to max effort, costing Opus-level money without delivering an Opus-level jump in score.
How should developers think about “max effort” settings? #
Max effort, extended thinking, or high-reasoning modes exist because some tasks genuinely need more deliberation: multi-step coding problems, long chains of tool calls, or ambiguous instructions that benefit from the model checking its own work. These modes aren’t a gimmick. They’re a real lever that trades latency and cost for accuracy.
The catch is that benchmark leaderboards often report a model’s best possible score, which usually means the most expensive configuration. If you’re deciding which model to put into production, the leaderboard entry that looks cheap based on list price might only hit that score at a setting that erases the discount entirely.
Remy doesn't build the plumbing. It inherits it. #
Other agents wire up auth, databases, models, and integrations from scratch every time you ask them to build something.
Remy ships with all of it from MindStudio — so every cycle goes into the app you actually want.
The practical fix is to test models at the configuration you’ll actually ship with, not the configuration used to win a benchmark. If your use case doesn’t need max reasoning effort, run Sonnet 5.5 at a lower effort tier and measure the real cost per task for your workload. If it does need max effort to hit acceptable quality, compare that cost directly against Opus 5.5’s default cost rather than assuming Sonnet wins by default.
Is Sonnet 5.5 worth it over Opus 5.5? #
It depends entirely on the task and the effort setting required to do it well. For straightforward tasks where Sonnet 5.5’s baseline (non-max) performance is already good enough, it’s cheaper in every sense: lower per-token rate, fewer tokens needed, lower total cost. For harder tasks where Sonnet needs max effort to match Opus-level output, the cost advantage can shrink to nothing, and you’re left running a model with a lower list price but a comparable bill.
This also matters for anyone building products on top of these APIs. A pricing page that advertises Sonnet as the budget option is only true at certain usage patterns. Teams running heavy coding agents, long multi-step automations, or anything that triggers extended reasoning by default should benchmark their own real workloads rather than trusting the per-token comparison sheet.
Frequently Asked Questions #
Why would a “cheaper” model cost more to run than a flagship model?
Because the list price is per token, not per task. If the cheaper model needs far more tokens (through extended reasoning or max effort settings) to reach a comparable quality bar, the total bill per completed task can match or exceed the flagship model’s cost, even though its per-token rate is lower.
What is “max effort” mode in Claude models?
It’s a setting that lets the model spend more computation and generate more reasoning tokens before answering, generally used to improve accuracy on hard problems. It raises both latency and cost, and benchmark scores reported at max effort don’t reflect the cost of running the model at lower, more typical settings.
How does GPT 6.1 Soul compare to Opus 5.5 and Sonnet 5.5 on price?
GPT 6.1 Soul is priced at $2 per million input tokens and $10 per million output tokens, well below Opus 5.5 and Astra’s rates, and its cost per task on blended benchmarks came in far lower too, around 72 cents versus roughly $6 to $8 for Opus 5.5 and Sonnet 5.5 at max effort.
Should I pick a model based on per-token price or cost per task?
Cost per task is the more reliable signal for budgeting, since it accounts for how many tokens a model actually needs to complete a given job at an acceptable quality level, not just the rate it charges per token.
Does this mean Sonnet 5.5 is a bad value?
No. At settings below max effort, Sonnet 5.5 can still be meaningfully cheaper than Opus 5.5 for tasks that don’t require heavy reasoning. The issue only arises when a task demands Sonnet’s highest effort tier to match Opus-level quality, at which point the price advantage narrows or disappears.