Grok 4.7 lists at $2 per million input tokens and $6 per million output tokens. Claude Opus 5.5 lists at $4 and $20. On price per token, Grok 4.7 looks like the cheaper model. On Artificial Analysis’s Intelligence Index, it costs about twice as much per task at its default effort ($2.73 against $1.34 for Opus 5.5 at its default) and scores lower (46.3 against 51.2).
The difference is token use. To run the index at its default effort, Grok 4.7 generated about 65,900 output tokens per task. Opus 5.5 at its default generated about 25,700. Muse Spark 1.3, the cheapest of the three at list price, sits between them: it scores above Grok 4.7 and costs less per task, but not less than Opus 5.5 at its default.
Every figure in this post comes from Artificial Analysis, an independent benchmarking firm that runs every model through the same ten evaluations and prices each run at the vendor’s list rates. We read the data on September 22, 2026, from index version 4.3.2. Opus 5.5 was added the day it launched, so its numbers may be revised as more runs complete.
- 01At each model’s default effort, Grok 4.7 cost $2.73 per task and Opus 5.5 cost $1.34.Grok 4.7 generated about 2.6× as many output tokens per task (65.9k against 25.7k), and 85% of its cost per task was input re-read across turns. It also scored lower (46.3 against 51.2).
- 02Grok 4.7 used 1.8× to 2.1× the output tokens per task of Grok 4.6 for about 2 index points.Cost per task rose 47% at high effort and 61% at xhigh; time per task rose from 619 to 1,068 seconds at high.
- 03Muse Spark 1.3 beats Grok 4.7 on both score and cost, but not Opus 5.5 at its default.Muse Spark 1.3 at max scored 48.1 for $1.60 per task, and Artificial Analysis measured it at about 213 output tokens per second.
- 04Opus 5.5 is token-efficient at medium and high, not at max.At max it used 119.2k output tokens per task, the most of any run here, for a top score of 57.6 at $5.98 per task.
01 — The rate cardsList prices, per million tokens #
Sources: Anthropic’s pricing page, xAI’s release notes and Meta’s Model API page, all read on September 22, 2026. Meta also sells a Muse Spark 1.3 Contributor tier at $0.10 / $0.20 whose prompts and outputs Meta may use to improve its products, so we leave it out of this comparison.
| Published list prices per million tokens, September 22, 2026. Grok 4.7’s higher band applies to the whole request once the prompt passes 200,000 tokens. | |||
|---|---|---|---|
| Line | Opus 5.5 | Grok 4.7 | Muse Spark 1.3 |
| --- | --- | --- | --- |
| Input | $4 | $2 | $1.25 |
| Cached input | $0.20 | $0.50 | $0.15 |
| Output | $20 | $6 | $4.25 |
| Long prompts | Same rate to 1M | $4 / $1 / $12 above 200K | One rate listed |
| Context window | 1M | 500K | 1M |
One line in this table reverses the usual order. Opus 5.5’s cached input costs $0.20 per million tokens, less than Grok 4.7’s $0.50, even though Opus 5.5’s uncached input costs twice as much. Agent loops re-send their context on every turn, so cached input is often the largest part of an agent’s bill. That is one reason list prices predict agent costs so poorly.
02 — The measurementWhat each model cost to run the same ten evaluations #
Artificial Analysis runs each model at several effort levels and records the tokens it used and the cost at list prices. The table shows every run it has published for these three models. Cost per task is Artificial Analysis’s own figure. It tested Muse Spark 1.3 only at xhigh and max.
| Source: Artificial Analysis Intelligence Index v4.3.2 model pages for Opus 5.5, Grok 4.7 and Muse Spark 1.3, read September 22, 2026. Output tokens are Artificial Analysis’s per-task average and include reasoning tokens. | |||
|---|---|---|---|
| Model · effort | Index | Output tokens / task | Cost per task |
| --- | --- | --- | --- |
| Opus 5.5 · low | 42.3 | 10.2k | $0.55 |
| Opus 5.5 · medium (default) | 51.2 | 25.7k | $1.34 |
| Opus 5.5 · high | 53.6 | 35.6k | $1.82 |
| Opus 5.5 · xhigh | 56.0 | 65.7k | $3.46 |
| Opus 5.5 · max | 57.6 | 119.2k | $5.98 |
| Grok 4.7 · high (default) | 46.3 | 65.9k | $2.73 |
| Grok 4.7 · xhigh | 46.4 | 80.6k | $3.74 |
| Muse Spark 1.3 · xhigh | 45.1 | 54.5k | $1.37 |
| Muse Spark 1.3 · max | 48.1 | 60.2k | $1.60 |
Output tokens per Intelligence Index task (thousands)
Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026 The cost split explains most of the gap, and it isn’t where the list prices suggest. Per task, 85% of Grok 4.7’s cost at high effort ($2.33 of $2.73) was input: context sent to the model again on each turn. Its output cost per task, $0.40, was lower than Opus 5.5’s $0.51 at medium. Opus 5.5 spent 61% of its $1.34 on input. Longer reasoning and more turns both mean the same context is paid for again, and cheap output tokens don’t offset that. Artificial Analysis doesn’t publish turn counts, so the mechanism is our reading of its cost split.
OpenRouter listed Grok 4.7 at $1.60 / $4.80 on September 22, 20% below xAI’s list price on every line. Applied to Artificial Analysis’s figure, that would bring Grok 4.7 at high effort to about $2.18 per task (our arithmetic), still above Opus 5.5’s $1.34.
Opus 5.5 is not token-efficient at every setting. At max it generated about 119,200 output tokens per task, more than Grok 4.7 at xhigh (80,600), and cost $5.98 per task. It also posted the highest score here, 57.6. The step from xhigh to max added 1.6 index points and cost 73% more per task. That matches Anthropic’s own advice to reserve the top settings for work where you’ve measured a gain.
03 — The regressionGrok 4.7 against Grok 4.6: small gains, much larger bills #
xAI kept Grok 4.7 at Grok 4.6’s list price, which our Grok 4.7 launch post covered. The Artificial Analysis runs show what the same price buys per task.
| Source: Artificial Analysis Intelligence Index v4.3.2, read September 22, 2026. Changes are our arithmetic. | |||
|---|---|---|---|
| Metric · effort | Grok 4.6 | Grok 4.7 | Change |
| --- | --- | --- | --- |
| Intelligence Index · high | 44.3 | 46.3 | +2.0 points |
| Output tokens per task · high | 35.8k | 65.9k | 1.8× |
| Cost per task · high | $1.86 | $2.73 | +47% |
| Time per task · high | 619 s | 1,068 s | +73% |
| Intelligence Index · xhigh | 44.2 | 46.4 | +2.2 points |
| Output tokens per task · xhigh | 37.6k | 80.6k | 2.1× |
| Cost per task · xhigh | $2.32 | $3.74 | +61% |
For about two index points, Grok 4.7 roughly doubles its output and takes 65–73% longer per task. The gains are real on some evaluations: Terminal-Bench 4.0 rose from 21.2% to 24.7% at high effort and GDPval-AA from 1605 to 1694 Elo. For a team already on Grok 4.6, the question is whether those specific gains matter enough to pay 47–61% more per task for them.
04 — The scorecardWhere each model leads #
The index is an average, and the averages hide real differences. The table shows the evaluations inside it, with each model at its default effort. Muse Spark 1.3 is shown at max, its highest-scoring tested setting.
| Opus 5.5 at medium, Grok 4.7 at high, Muse Spark 1.3 at max. Source: Artificial Analysis v4.3.2, read September 22, 2026. | |||
|---|---|---|---|
| Evaluation | Opus 5.5 | Grok 4.7 | Muse 1.3 |
| --- | --- | --- | --- |
| Terminal-Bench 4.0 | 52.5% | 24.7% | 33.3% |
| Humanity’s Last Exam | 54.7% | 42.3% | 48.7% |
| CritPt (physics research) | 27.7% | 18.0% | 24.9% |
| AA-LCR (long-context reasoning) | 84.3% | 77.0% | 83.0% |
| GDPval-AA v2.1 (Elo) | 1576 | 1694 | 1674 |
| AutomationBench-AA (partial score) | 61.2% | 63.5% | 57.9% | | SciCode | 59.3% | 57.8% | 58.8% | | AA-Omniscience index | 40.3 | 30.9 | 25.0 |
At the default settings, Grok 4.7 leads on two evaluations: GDPval-AA, which scores real professional tasks, and AutomationBench. Opus 5.5 at medium leads everywhere else, by about 28 points on Terminal-Bench 4.0 and 12 points on Humanity’s Last Exam. The Grok leads are not fixed. Opus 5.5 at high, which still costs less per task than Grok 4.7 at high ($1.82 against $2.73), scores 1692 on GDPval-AA and 63.2% on AutomationBench, level with Grok 4.7 on both.
Muse Spark 1.3 is the middle model on most rows. Its advantage is speed. Artificial Analysis measured it at about 213 output tokens per second at max, and it averaged 278 seconds per index task, against 1,068 seconds for Grok 4.7 at high. Artificial Analysis had not yet published a time per task for Opus 5.5. Our Muse Spark 1.3 launch post covers Meta’s own token claims.
Artificial Analysis’s knowledge test, AA-Omniscience, rewards right answers and penalises wrong ones, and doesn’t penalise declining to answer. Opus 5.5 answered the most questions correctly (64.5% at medium, against 47.8% for Grok 4.7 and 43.6% for Muse Spark 1.3). But when it didn’t know, it gave a wrong answer 68.4% of the time, against about 32–33% for the other two, which more often declined. For work where a confident wrong answer is costly, that matters as much as the score.
05 — The caveatsHow to read these numbers #
Artificial Analysis’s numbers are independent, which is why this post uses them, but they measure one fixed workload. Four things to keep in mind:
- Its harness scores lower than the vendors’. On Terminal-Bench 4.0 at xhigh, Artificial Analysis measured Opus 5.5 at 59.6% where Anthropic reports 66.4%, and Grok 4.7 at 25.8% where xAI reports 38.0%. The drop is larger for Grok 4.7.
- Opus 5.5 was run with fallback on. Artificial Analysis labels its Opus 5.5 runs “Default Fallback,” meaning requests that tripped a safeguard were answered by another Claude model, as they would be in production with that option enabled.
- Scores change between index versions. Version 4.3.2 includes ten evaluations, among them AA-Briefcase, GDPval-AA, AutomationBench-AA and Terminal-Bench 4.0. Earlier versions used a different set of evaluations. Our September 2 Muse Spark 1.3 post quoted an older version, and its figures can’t be compared with the ones here.
- Cost depends on your mix. Cost per task reflects this set of evaluations at list prices. A workload with shorter contexts, more caching or fewer turns will cost differently, and the ranking can change.
For more on each model, see our Opus 5.5 launch post and Opus 5.5 vs GPT-6 Astra. The raw Artificial Analysis pages for Opus 5.5, Grok 4.7 and Muse Spark 1.3 will carry any updates.
06 — ConclusionToken use, not list price, decides which model is cheapest per task #
Price your own tasks per completed job, not per million tokens, before you pick the cheaper-looking model
On Artificial Analysis’s workload, the model with the highest list price cost the least per task at its default setting, because it used far fewer tokens to finish. That won’t hold for every workload, and it didn’t hold for Opus 5.5 at max. Run a sample of your own tasks on each model at its default and one higher effort level, record total cost per completed task, and compare those numbers. Our price index keeps the list prices current.