OpenAI’s much-awaited Astra model hasn’t brought about any gains at all in the popular Artificial Analysis Intelligence Index.
GPT-6 Astra, running at max reasoning effort, has landed on the benchmark with a score of just 61 — identical to the score its predecessor, GPT-5.6 Sol, put up when it launched. For a model that was supposed to represent OpenAI’s next big leap, a flat line on the industry’s most closely watched intelligence benchmark is not the headline the company would have wanted.
The 61 score places GPT-6 Astra in a three-way tie with GPT-5.6 Sol and xAI’s Grok 4.6 (high), five points behind Anthropic’s Claude Fable 5.1, which currently tops the index at 66. It also trails Claude Opus 5 at 63 and, more awkwardly for OpenAI, comes in behind Meta’s newly released Muse Spark 1.3, which scored 62. Claude Fable 5 also edges it out at 62. In an index where model generations typically fight for single-digit gains, GPT-6 Astra’s inability to move the needle at all over Sol is a rare outcome.
Making It Up On Price, Not Score #
The stagnant score comes alongside a steep price increase. Artificial Analysis notes that GPT-6 Astra is priced at 2.5x GPT-5.6 Sol’s current rates, jumping from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount on cache reads and a 25% premium on cache writes carried over. GPT-6 Astra does use about 10% fewer output tokens than Sol at max effort, and it technically defines a new Pareto frontier for the Intelligence Index when plotted against output tokens per task. But that token efficiency isn’t nearly enough to offset the pricing jump — Artificial Analysis calculates that Astra ends up 75% more expensive per task than Sol for essentially the same intelligence score, leaving it sitting behind its predecessor on the Intelligence Index vs Cost per Task frontier.
It’s a similar story to what played out when Claude Opus 4.8 topped the index at launch and when GPT-5.5 briefly led the pack — pricing and efficiency have become just as central to these launches as the raw score itself, and this time the efficiency gains simply don’t cover the cost hike.
A Much Better Story In Coding #
If the Intelligence Index is a wash, GPT-6 Astra’s performance on the Artificial Analysis Coding Agent Index tells a more flattering story. Running in the Codex harness, Astra scores 67 — roughly on par with Claude Opus 5 and Claude Fable 5 running in Claude Code, and with Muse Spark 1.3 in Muse Code. Claude Fable 5.1 leads that index outright at 70. The efficiency gains here are real and substantial. Artificial Analysis says GPT-6 Astra uses roughly a third of the tokens that GPT-5.6 Sol (max) needed in the Codex harness for a similar task, and about a fifth of what Claude Opus 5 (xhigh) uses. At max effort, Astra costs about the same as GPT-5.6 Sol (max) while scoring two points higher on the index, and on a per-task basis it comes in at less than half the cost of Claude Fable 5 for an equivalent score. That puts GPT-6 Astra on the Pareto frontier for the Coding Agent Index versus cost — a much stronger positioning than it manages on the Intelligence Index.
Hallucinations Cut In Half, But Other Metrics Slip #
The other notable improvement is on AA-Omniscience, Artificial Analysis’s knowledge and hallucination benchmark. GPT-6 Astra’s hallucination rate drops from 92% for GPT-5.6 Sol down to 51% — a significant reduction — and Artificial Analysis is careful to point out this doesn’t come at the cost of accuracy, since Astra’s accuracy score actually improved by 4 points over Sol at the same time.
On AA-Briefcase, Artificial Analysis’s own long-horizon knowledge-work evaluation that tests models on multi-week projects spanning thousands of source files, GPT-6 Astra gains roughly 80 Elo points over its predecessor, driven by higher rubric scores and a jump in Analytical Quality Elo. That gain comes with a trade-off, though: Presentation Quality Elo actually falls, and GPT-5.6 Sol (max) continues to lead every model on that specific sub-metric.
Elsewhere, the picture is mixed at best. GPT-6 Astra gains 6 points on Humanity’s Last Exam, but that’s offset by an 80-Elo-point drop on GDPval-AA v2, the benchmark adapted from OpenAI’s own dataset measuring economically valuable work across 44 occupations. Artificial Analysis also flags 2-3 point regressions across τ³-Banking (customer support), SciCode (scientific Python problems), and AA-LCR (long-context reasoning) — a fairly broad spread of small step-backs for a model billed as a major upgrade.
For a launch that OpenAI clearly wanted to frame as a decisive leap forward, GPT-6 Astra’s Artificial Analysis results end up reading as a coding and efficiency story wrapped around an Intelligence Index score that simply didn’t move.