cd /news/artificial-intelligence/gpt-5-6-sol-pro-tops-editorial-craft… · home topics artificial-intelligence article
[ARTICLE · art-105236] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score

OpenAI's GPT-5.6 Sol Pro topped the 12-task Editorial Craft benchmark with a mean score of 0.97 at an estimated cost of $0.0091 per task, according to RuntimeWire's evaluation of 12 AI models. OpenAI claimed four of the top six spots, while SpaceXAI's Grok 4.6 and Anthropic's Claude Opus 4.8 tied at 0.93 among leading non-OpenAI models.

read2 min views7 publishedAug 20, 2026
GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score
Image: Runtimewire (auto-discovered)

In a 12-task Editorial Craft evaluation of 12 AI models, OpenAI’s GPT-5.6 Sol Pro ranked first with a 0.97 score at $0.0091 per task. OpenAI claimed four of the top six spots, while Grok 4.6 and Claude Opus 4.8 tied at 0.93 among the leading non-OpenAI models.

By RuntimeWire Staff · Published

Leaderboard

Rank Model Mean score Est. cost/task
1 OpenAI: GPT-5.6 Sol Pro 0.97 $0.0091
2 OpenAI: GPT-5.6 Luna 0.94 $0.0007
3 OpenAI: GPT-5.6 Terra Pro 0.94 $0.0073
4 OpenAI: GPT-5.6 Luna Pro 0.93 $0.0007
5 SpaceXAI: Grok 4.6 0.93 $0.0038
6 OpenAI: GPT-5.6 Terra 0.93 $0.0075
7 Anthropic: Claude Opus 4.8 0.93 $0.0160
8 OpenAI: GPT-5.6 Sol 0.92 $0.0091
9 Anthropic: Claude Sonnet 5 0.88 $0.0063
10 Anthropic: Claude Fable 5 0.80 $0.0297
11 Google: Gemini 3.7 Flash 0.78 $0.0011

| 12 | Phi-4-reasoning | 0.47 | — |

How we scored it

Every model answered the same 12-task battery from Editorial Craft, one task at a time, with no tools and no retries on content.

2 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.

10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort

, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the interactive leaderboard.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @openai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gpt-5-6-sol-pro-tops…] indexed:0 read:2min 2026-08-20 ·