In a 12-task Editorial Craft evaluation of 12 AI models, OpenAI’s GPT-5.6 Sol Pro ranked first with a 0.97 score at $0.0091 per task. OpenAI claimed four of the top six spots, while Grok 4.6 and Claude Opus 4.8 tied at 0.93 among the leading non-OpenAI models.
By RuntimeWire Staff · Published
Leaderboard
| Rank | Model | Mean score | Est. cost/task |
|---|---|---|---|
| 1 | OpenAI: GPT-5.6 Sol Pro | 0.97 | $0.0091 |
| 2 | OpenAI: GPT-5.6 Luna | 0.94 | $0.0007 |
| 3 | OpenAI: GPT-5.6 Terra Pro | 0.94 | $0.0073 |
| 4 | OpenAI: GPT-5.6 Luna Pro | 0.93 | $0.0007 |
| 5 | SpaceXAI: Grok 4.6 | 0.93 | $0.0038 |
| 6 | OpenAI: GPT-5.6 Terra | 0.93 | $0.0075 |
| 7 | Anthropic: Claude Opus 4.8 | 0.93 | $0.0160 |
| 8 | OpenAI: GPT-5.6 Sol | 0.92 | $0.0091 |
| 9 | Anthropic: Claude Sonnet 5 | 0.88 | $0.0063 |
| 10 | Anthropic: Claude Fable 5 | 0.80 | $0.0297 |
| 11 | Google: Gemini 3.7 Flash | 0.78 | $0.0011 |
| 12 | Phi-4-reasoning | 0.47 | — |
How we scored it
Every model answered the same 12-task battery from Editorial Craft, one task at a time, with no tools and no retries on content.
2 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.
10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.
Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort
, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.
A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.
Explore every prompt, answer, and per-task grade in the interactive leaderboard.