cd /news/large-language-models/gpt-6-sol-vs-claude-opus-5-5-cost-pe… · home topics large-language-models article
[ARTICLE · art-137455] src=digitalapplied.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

GPT-6 Sol vs Claude Opus 5.5: Cost per Task and Benchmarks

Artificial Analysis found that OpenAI's GPT-6 Sol is the cheaper way to reach any Intelligence Index score up to about 44, but Anthropic's Claude Opus 5.5 outscores it above that threshold, according to the evaluator's Intelligence Index v4.3.2 read on September 22, 2026. Sol at its maximum setting scored 47.5 for $1.06 per task, below Opus 5.5 at its default medium setting, which scored 51.2 for $1.34 per task. The two models launched the same day, with Sol listed at $2 per million input tokens and $10 per million output tokens versus Opus 5.5's $4 and $20, though Artificial Analysis noted cache reads cost $0.20 on both and requests over 272K input tokens erase Sol's price advantage.

read10 min views1 publishedSep 22, 2026
GPT-6 Sol vs Claude Opus 5.5: Cost per Task and Benchmarks
Image: Digitalapplied (auto-discovered)

OpenAI released GPT-6 Sol and Anthropic released Claude Opus 5.5 on the same day, September 22, 2026. Sol lists at $2 per million input tokens and $10 per million output tokens, half of Opus 5.5’s $4 and $20. Neither vendor compared the two: OpenAI’s charts show Opus 5, and Anthropic’s table shows GPT-5.6 Sol.

The independent evaluator Artificial Analysis has now run both models on the same ten tests at every effort level. On its index, Sol is the cheaper way to reach any score up to about 44. Above that, Sol runs out of headroom. Its best result, 47.5 at max for $1.06 a task, is below Opus 5.5 at its default medium setting, 51.2 for $1.34.

Every Artificial Analysis figure below was read from its GPT-6 Sol and Claude Opus 5.5 model pages on September 22, Intelligence Index version 4.3.2. It measures cost per task at list prices, including cache reads and writes.

  1. 01Up to an index score of about 44, GPT-6 Sol reaches it for less.Sol at high scores 42.8 for $0.37 a task. Opus 5.5 at low scores 42.3 for $0.55, about 50% more.
  2. 02Opus 5.5 at its default outscores Sol at its maximum.Opus 5.5 at medium: 51.2 for $1.34. Sol at max: 47.5 for $1.06. That is 3.7 more points for 26% more money.
  3. 03The biggest gaps are coding and knowledge work; business workflows are level.On Terminal-Bench 4.0, Opus 5.5 at medium scores 52.5% against Sol's 43.9% at max. On AutomationBench-AA, Sol at xhigh matches Opus 5.5 at medium for 40% of the cost.
  4. 04Half the list price does not mean half the bill.Cache reads cost $0.20 on both. On a cache-heavy agent task Sol costs about 57% of Opus 5.5, and requests over 272K input tokens erase the gap.

01 — The ladderThe same tests, at every effort level #

The Intelligence Index averages ten evaluations: agentic coding, business workflows, knowledge-work documents, science, long-context reasoning and factual knowledge. Both models default to medium, and both offer low through max. Each cell shows the index score and the average cost per task.

Index score · cost per task at list prices. Source: Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026.
Effort Claude Opus 5.5 GPT-6 Sol
--- --- ---
none Not offered 28.1 · $0.33
low 42.3 · $0.55 33.9 · $0.13
medium (default) 51.2 · $1.34 39.8 · $0.25
high 53.6 · $1.82 42.8 · $0.37
xhigh 56.0 · $3.46 44.1 · $0.53
max 57.6 · $5.98 47.5 · $1.06

Intelligence Index score, sorted by cost per task

Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026 Sorted by cost, the two models barely overlap. Every Sol setting below max costs less than the cheapest Opus 5.5 setting, and Sol at xhigh (44.1 for $0.53) already beats Opus 5.5 at low (42.3 for $0.55). Everything Opus 5.5 does from medium upwards scores higher than anything Sol can reach. The choice is less “which model is better” than “which price band does this task belong in.”

One result goes against intuition. Sol with reasoning switched off scored 28.1 for $0.33 a task, below Sol at low (33.9) and more than twice as expensive. Almost all of the extra cost is input. One likely reason is that a model without reasoning takes more steps to finish, and each step re-sends the conversation. Turning reasoning off is not a reliable way to save money on agent work.

02 — The detailTest by test: where the gap opens and where it closes #

The fairest pairing is each model at the setting a team is most likely to use for demanding work: Opus 5.5 at its default medium, against Sol at max, which costs about the same. Sol at xhigh is added as the cost-conscious option.

Source: Artificial Analysis model pages, Intelligence Index v4.3.2, read September 22, 2026. Costs are averages across the whole index.
Evaluation Opus 5.5 medium Sol max Sol xhigh
--- --- --- ---
Cost per index task $1.34 $1.06 $0.53
Output tokens per task 25.7k 31.2k 16.0k
Intelligence Index v4.3.2 51.2 47.5 44.1
Terminal-Bench 4.0 (agentic coding) 52.5% 43.9% 30.3%
AutomationBench-AA (business workflows, partial credit) 61.2% 61.6% 61.7%
GDPval-AA (knowledge work, Elo) 1576 1487 1437
AA-Briefcase (work documents, Elo) 1642 1483 1364
Humanity’s Last Exam 54.7% 47.9% 46.3%
SciCode (scientific coding) 59.3% 57.6% 55.1%
CritPt (physics research) 27.7% 30.9% 28.0%
AA-LCR (long-context reasoning) 84.3% 83.7% 81.3%
AA-Omniscience accuracy 64.5% 54.5% 53.8%
AA-Omniscience hallucination rate (lower is better) 68.4% 60.1% 58.9%

Coding is the widest gap. On Terminal-Bench 4.0, Opus 5.5 at medium scores 52.5% against Sol’s 43.9% at max, and Opus 5.5 reaches 59.6% at xhigh. Sol falls away fast at lower settings: 30.3% at xhigh and 26.3% at high. For coding agents that run in a terminal, Sol’s lower price buys noticeably weaker results.

Knowledge work also favours Opus 5.5. On GDPval-AA and AA-Briefcase, which rate work documents in head-to-head comparisons scored as Elo, Opus 5.5 at medium leads Sol at max by about 90 and 160 Elo points. At max, Opus 5.5 reaches 1846 on GDPval-AA.

Business workflows are level. On AutomationBench-AA, Artificial Analysis’s partial-credit run of business tasks across apps, Sol at xhigh (61.7%) matches Opus 5.5 at medium (61.2%) at 40% of the average cost per task. Long-context reasoning is also level, and at these settings Sol leads on CritPt, a physics research test.

AA-Omniscience rewards right answers, penalises wrong ones and does not penalise declining. Opus 5.5 answers more questions correctly (64.5% against 54.5%). When it doesn’t know, it gives a wrong answer 68.4% of the time, against 60.1% for Sol at max and 50.7% for Sol at low. Sol knows less but guesses less. Neither rate is low enough to skip verification on factual work.

03 — The launchesThe vendors’ own numbers, and why they don’t line up #

Each vendor published results for the benchmarks both models share. They point the same way as the independent data, but they are not a clean comparison.

  • AutomationBench (Zapier). Anthropic reports Opus 5.5 at 40.0% atmax , from Zapier’s early-access run. OpenAI’s chart puts Sol’s best at 33.2%, atxhigh . These are Zapier’s own scores, which run far lower than Artificial Analysis’s partial-credit version in the table above.
  • FrontierCode 1.1 Main (Cognition). Anthropic reports Opus 5.5 at 54.4%; OpenAI’s chart puts Sol at 49.3%, both atmax . But the two vendors disagree about the same model: Anthropic lists Opus 5 at 48.0%, OpenAI’s chart at 53.4%. A five-point gap between vendors’ runs is larger than the gap between these two models, so treat this row as inconclusive.
  • OSWorld 2.0. Anthropic reports 81.8% on the partial-credit score; OpenAI reports 64.4% for Sol on the offline set. These are different task sets and cannot be compared.

Each launch also claims cost wins, but against other models: OpenAI against Opus 5, Anthropic against GPT-6 Astra and GPT-5.6 Sol. Our GPT-6 Sol and Luna launch analysis and Opus 5.5 launch analysis read those claims against each vendor’s own chart data.

At these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences.Anthropic, Introducing Claude Opus 5.5, September 22, 2026

04 — The invoiceThe price sheets: half the list price is not half the bill #

Per million tokens. Sources: Anthropic pricing and OpenAI pricing , September 22, 2026. Ratios are our arithmetic.
Line Opus 5.5 GPT-6 Sol Sol as % of Opus
--- --- --- ---
Input $4 $2 50%
Cache read $0.20 $0.20 100%
Cache write (Anthropic: 5-minute) $5 $2.50 50%
Output $20 $10 50%
Request over 272K input, input / output $4 / $20 $4 / $15 100% / 75%
Batch, input / output $2 / $10 $1 / $5 50%
Fast mode, input / output $8 / $40 $4 / $20 50%

Anthropic pricingand OpenAI pricing, September 22, 2026. Ratios are our arithmetic.

Two lines break the “half price” rule. Cache reads cost $0.20 per million tokens on both models, and for agents, cache reads are usually the largest share of the bill. And Anthropic charges the full 1M-token window at standard rates, while OpenAI bills any request over 272K input tokens at twice the input and cache rates and 1.5 times the output rate.

The table below reuses the illustrative agent task from our Opus 5.5 launch post. The token counts are invented for the example and kept the same for both models; the rates are the published ones.

Illustrative token counts at list prices, September 22, 2026. The last column assumes every request carries more than 272K input tokens.
Illustrative task Opus 5.5 Sol Sol, long prompts
--- --- --- ---
Cache reads · 8,000,000 $1.60 $1.60 $3.20
Uncached input · 400,000 $1.60 $0.80 $1.60
Cache writes · 600,000 $3.00 $1.50 $3.00
Output incl. reasoning · 300,000 $6.00 $3.00 $4.50
Total $12.20 $6.90 $12.30

At equal token counts, Sol costs 57% of Opus 5.5 on this cache-heavy task, not 50%. If every request goes past 272K input tokens, Sol costs slightly more than Opus 5.5. Real token counts differ between the models, and Artificial Analysis’s measured figures already include that difference. At their defaults, Sol costs $0.25 a task against $1.34 for Opus 5.5, about a fifth, but for an index score of 39.8 against 51.2.

05 — The plumbingIntegration differences that decide a migration #

Sources: Anthropic’s Opus 5.5 model and migration pages; OpenAI’s GPT-6 Sol model page and announcement, September 22, 2026.
Item Claude Opus 5.5 GPT-6 Sol
--- --- ---
Reasoning off Not possible; low is the minimum effort: none
Changing effort mid-conversation Invalidates the prompt cache; a per-message effort beta avoids it Keeps the prompt cache, per OpenAI
Tool calling tool_choice any or tool returns a 400 Responses API; Chat Completions only with effort none
Context / max output 1M at standard rates / 128K 1.05M, surcharge above 272K / 128K
Knowledge cutoff June 2026 April 20, 2026
Where to run it Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud, Microsoft Foundry OpenAI API
Subscriptions Claude apps and Claude Code ChatGPT Work and Codex on Plus, Pro, Business, Enterprise and Edu

The cache behaviour matters most for routers. A system that raises effort on hard turns and lowers it on easy ones keeps its cache on Sol. On Opus 5.5 it needs the per-message effort beta, or it pays to rebuild the cache after every change. Opus 5.5 also routes most cybersecurity tasks to Opus 4.8 behind the scenes, so security teams should test that workload directly before choosing.

06 — ConclusionSol wins the low price band; Opus 5.5 owns everything above it #

Route by price band: send budget work to Sol and demanding work to Opus 5.5, then check both on your own tasks

One independent harness is a strong start, but not a verdict on your workload. Take ten to twenty real tasks, run Sol at xhigh and Opus 5.5 at medium, and compare the cost of each completed task, not the cost per token. If you also use Anthropic’s top model, our Opus 5.5 and GPT-6 Astra comparison covers the next tier up, and the frontier model price index tracks list prices across vendors.

── more in #large-language-models 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gpt-6-sol-vs-claude-…] indexed:0 read:10min 2026-09-22 ·