# GPT 6 Sol Shows Modest Gain Over GPT 5.6 Sol On Artificial Analysis Intelligence Index, But At A Much Cheaper Price

> Source: <https://officechai.com/ai/gpt-6-sol-shows-modest-gain-over-gpt-5-6-sol-on-artificial-analysis-intelligence-index-but-at-a-much-cheaper-price/>
> Published: 2026-09-22 18:27:56+00:00

Independent benchmarking firm Artificial Analysis has published its first full evaluation of [GPT-6 Sol and GPT-6 Luna](https://officechai.com/ai/openai-releases-gpt-6-sol-and-gpt-6-luna-with-50-lower-prices-than-gpt-5-6/), OpenAI’s newly released mid-tier and budget models sitting below flagship [GPT-6 Astra](https://officechai.com/ai/openai-says-its-models-have-resolved-more-than-100-long-standing-mathematical-problems-in-addition-to-navier-stokes/). The topline finding: both models are roughly half the price of their GPT-5.6 predecessors, but their actual intelligence gains are modest at best, and in some evaluations, the models score lower than the versions they’re replacing.

## Pricing Cut In Half, Intelligence Roughly Flat

GPT-6 Sol now costs $2 per million input tokens and $10 per million output tokens, down from $4 and $20 for GPT-5.6 Sol. GPT-6 Luna drops from $0.20/$1.20 to $0.10/$0.50 per million input/output tokens. Cache reads keep their 90% discount and cache writes still carry a 25% premium over standard input pricing.

On Artificial Analysis’s own Intelligence Index, a composite benchmark spanning ten evaluations including AA-Briefcase, GDPval-AA v2, AutomationBench-AA, Terminal-Bench 4.0, SciCode, and Humanity’s Last Exam, GPT-6 Sol at maximum effort scores 48, essentially level with GPT-5.6 Sol’s 47. GPT-6 Luna at maximum effort scores 37, tied exactly with GPT-5.6 Luna. For comparison, Claude Opus 5.5 tops the current leaderboard at 58, followed by Claude Fable 5.1 at 53 and GPT-6 Astra at 53, while rival mid-tier models like [Grok 4.7](https://officechai.com/ai/grok-4-7s-score-jumps-2-points-on-artificial-analysis-intelligence-index-but-scores-below-gpt-5-6-sol-muse-spark-1-3-fable-5/) and Xiaomi’s [MiMo-V2.6-Pro](https://officechai.com/ai/xiaomi-mimo-v-2-6-pro-benchmarks/) score 46 apiece, putting both new OpenAI models in tight competition with a crowded pack of similarly priced models.

Where GPT-6 Sol and Luna do stand out is cost efficiency. Artificial Analysis calculates that running GPT-6 Sol at max effort through the full Intelligence Index costs $1.06 per task, about 50% less than GPT-5.6 Sol’s $1.99, even though the new model actually uses slightly more output tokens per task (31,000 versus 29,000). GPT-6 Luna costs just $0.07 per task to run the same suite, roughly 60% less than GPT-5.6 Luna’s $0.18, despite also using more tokens (51,000 versus 41,000). That combination of flat or slightly improved intelligence scores and sharply lower running costs pushes both models onto, or close to, the Pareto frontier of intelligence versus cost — meaning that for their price point, nothing currently scores meaningfully higher.

## Coding Results Are Mixed Between The Two Models

The picture diverges more sharply on Artificial Analysis’s Coding Agent Index, which is built from DeepSWE v1.1, Terminal-Bench 4.0, and SWE-Atlas-QnA and run through OpenAI’s own Codex harness. GPT-6 Sol at max effort scores 57, a two-point improvement over GPT-5.6 Sol’s 55, driven by gains in Terminal-Bench 4.0 (43% versus 37%) and SWE-Atlas-QnA (58% versus 54%). At $2.99 per task, that’s roughly half of what GPT-5.6 Sol cost to run the same benchmark, and it lands the model on the coding cost-efficiency frontier alongside top scorers like Codex running GPT-6 Astra and Claude Code running Claude Opus 5.

GPT-6 Luna moves the other way. It scores 41 on the Coding Agent Index, two points below GPT-5.6 Luna’s 43, with declines in both SWE-Atlas-QnA (44% versus 49%) and DeepSWE v1.1 (64% versus 66%). It still costs around 60% less per task than its predecessor, but the regression means Luna’s coding upgrade is, by Artificial Analysis’s numbers, a step backward in raw capability even as it gets cheaper.

## A Sharp Drop In Hallucinations, With A Catch

The clearest improvement for both models shows up in AA-Omniscience, Artificial Analysis’s knowledge-reliability benchmark, which scores models on correct answers while penalizing confident wrong ones and applying no penalty for declining to answer. GPT-6 Sol’s hallucination rate falls from 92% to 60%, and GPT-6 Luna’s from 93% to 77%.

The catch is that Sol achieves much of that improvement by answering fewer questions rather than getting more of them right. It now attempts only 83% of questions, down from 99% for GPT-5.6 Sol, which cuts wrong answers by roughly a quarter but also drags accuracy down five points, from 59% to 54%. Luna’s accuracy is essentially unchanged, at 44% versus 43% for its predecessor, while it too answers a smaller share of questions. Net-net, the AA-Omniscience Index score, which factors in both correctness and hallucination avoidance, still rises meaningfully for both models: Sol climbs from 22 to 27, and Luna goes from a negative score of -10 up to 1.

## Real-World Work Benchmarks Show Regressions

Not every evaluation moved in the new models’ favour. Both Sol and Luna improved on AutomationBench-AA (Sol at 62% versus 60%, Luna at 53% versus 50%) and Terminal-Bench 4.0 (Sol at 44% versus 40%, Luna at 13% versus 12%), but both regressed on two benchmarks built around longer, more realistic knowledge work.

On GDPval-AA v2.1, an Elo-based leaderboard built from OpenAI’s own dataset of economically valuable tasks spanning 44 occupations, GPT-6 Sol at max effort scores 1,487 Elo, about 100 points below GPT-5.6 Sol’s 1,588. GPT-6 Luna drops even further in relative terms, falling roughly 75 Elo points to 1,367. Luna also loses around 45 Elo points on AA-Briefcase v1.1, a private, multi-week knowledge-work benchmark involving thousands of input files, while Sol holds roughly steady there. Artificial Analysis says its team manually reviewed hundreds of outputs behind these regressions and found they were generally driven by lower presentation quality and deliverables that skipped required elements from the task rubric, rather than by outright factual or reasoning failures.

## The Bigger Picture

Taken together, the numbers suggest OpenAI’s pitch for Sol and Luna rests almost entirely on cost rather than a fresh leap in raw intelligence. Both models roughly match their predecessors on the general Intelligence Index, Sol gets modestly better at coding while Luna gets modestly worse, both hallucinate far less but also answer fewer questions, and both lose ground on tasks that resemble real, deliverable-based office work. For teams choosing between the GPT-5.6 and GPT-6 mid-tier models purely on capability, the newer models aren’t a clear across-the-board upgrade — but at roughly half the price per task, they don’t need to be, and that alone is enough to shift where they sit on the cost-efficiency frontier relative to competitors still catching up on price.
