GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score OpenAI's GPT-5.6 Sol Pro topped the 12-task Editorial Craft benchmark with a mean score of 0.97 at an estimated cost of $0.0091 per task, according to RuntimeWire's evaluation of 12 AI models. OpenAI claimed four of the top six spots, while SpaceXAI's Grok 4.6 and Anthropic's Claude Opus 4.8 tied at 0.93 among leading non-OpenAI models. GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score In a 12-task Editorial Craft evaluation of 12 AI models, OpenAI’s GPT-5.6 Sol Pro ranked first with a 0.97 score at $0.0091 per task. OpenAI claimed four of the top six spots, while Grok 4.6 and Claude Opus 4.8 tied at 0.93 among the leading non-OpenAI models. By RuntimeWire Staff /author/runtimewire-staff · Published Leaderboard | Rank | Model | Mean score | Est. cost/task | |---|---|---|---| | 1 | OpenAI: GPT-5.6 Sol Pro | 0.97 | $0.0091 | | 2 | OpenAI: GPT-5.6 Luna | 0.94 | $0.0007 | | 3 | OpenAI: GPT-5.6 Terra Pro | 0.94 | $0.0073 | | 4 | OpenAI: GPT-5.6 Luna Pro | 0.93 | $0.0007 | | 5 | SpaceXAI: Grok 4.6 | 0.93 | $0.0038 | | 6 | OpenAI: GPT-5.6 Terra | 0.93 | $0.0075 | | 7 | Anthropic: Claude Opus 4.8 | 0.93 | $0.0160 | | 8 | OpenAI: GPT-5.6 Sol | 0.92 | $0.0091 | | 9 | Anthropic: Claude Sonnet 5 | 0.88 | $0.0063 | | 10 | Anthropic: Claude Fable 5 | 0.80 | $0.0297 | | 11 | Google: Gemini 3.7 Flash | 0.78 | $0.0011 | | 12 | Phi-4-reasoning | 0.47 | — | How we scored it Every model answered the same 12-task battery from Editorial Craft , one task at a time, with no tools and no retries on content. 2 tasks carry ground truth exact, numeric, multiple-choice, or the benchmark's official matcher and were scored deterministically against the reference answer — correct is 1, incorrect is 0. 10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale. Reasoning effort was pinned to max for every model whose lane exposes a control OpenAI-style reasoning effort , Anthropic extended thinking, Gemini thinking config . Models marked vendor default on the interactive board expose no such control and ran as shipped. A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing "—" where pricing isn't public , so treat it as directional, not billing-exact. Explore every prompt, answer, and per-task grade in the interactive leaderboard /benchmarks/editorial-craft-leaderboard-2 .