{"slug": "gpt-5-6-sol-pro-tops-editorial-craft-benchmark-with-0-97-score", "title": "GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score", "summary": "OpenAI's GPT-5.6 Sol Pro topped the 12-task Editorial Craft benchmark with a mean score of 0.97 at an estimated cost of $0.0091 per task, according to RuntimeWire's evaluation of 12 AI models. OpenAI claimed four of the top six spots, while SpaceXAI's Grok 4.6 and Anthropic's Claude Opus 4.8 tied at 0.93 among leading non-OpenAI models.", "body_md": "# GPT-5.6 Sol Pro tops Editorial Craft benchmark with 0.97 score\n\n**In a 12-task Editorial Craft evaluation of 12 AI models, OpenAI’s GPT-5.6 Sol Pro ranked first with a 0.97 score at $0.0091 per task. OpenAI claimed four of the top six spots, while Grok 4.6 and Claude Opus 4.8 tied at 0.93 among the leading non-OpenAI models.**\n\nBy [RuntimeWire Staff](/author/runtimewire-staff)\n· Published\n\n### Leaderboard\n\n| Rank | Model | Mean score | Est. cost/task |\n|---|---|---|---|\n| 1 | OpenAI: GPT-5.6 Sol Pro | 0.97 | $0.0091 |\n| 2 | OpenAI: GPT-5.6 Luna | 0.94 | $0.0007 |\n| 3 | OpenAI: GPT-5.6 Terra Pro | 0.94 | $0.0073 |\n| 4 | OpenAI: GPT-5.6 Luna Pro | 0.93 | $0.0007 |\n| 5 | SpaceXAI: Grok 4.6 | 0.93 | $0.0038 |\n| 6 | OpenAI: GPT-5.6 Terra | 0.93 | $0.0075 |\n| 7 | Anthropic: Claude Opus 4.8 | 0.93 | $0.0160 |\n| 8 | OpenAI: GPT-5.6 Sol | 0.92 | $0.0091 |\n| 9 | Anthropic: Claude Sonnet 5 | 0.88 | $0.0063 |\n| 10 | Anthropic: Claude Fable 5 | 0.80 | $0.0297 |\n| 11 | Google: Gemini 3.7 Flash | 0.78 | $0.0011 |\n| 12 | Phi-4-reasoning | 0.47 | — |\n\n### How we scored it\n\nEvery model answered the same 12-task battery from **Editorial Craft**, one task at a time, with no tools and no retries on content.\n\n2 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.\n\n10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.\n\nReasoning effort was pinned to **max** for every model whose lane exposes a control (OpenAI-style `reasoning_effort`\n\n, Anthropic extended thinking, Gemini thinking config). Models marked *vendor default* on the interactive board expose no such control and ran as shipped.\n\nA model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing (\"—\" where pricing isn't public), so treat it as directional, not billing-exact.\n\nExplore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/editorial-craft-leaderboard-2).", "url": "https://wpnews.pro/news/gpt-5-6-sol-pro-tops-editorial-craft-benchmark-with-0-97-score", "canonical_source": "https://runtimewire.com/article/gpt-5-6-sol-pro-tops-editorial-craft-benchmark-with-0-97-score", "published_at": "2026-08-20 22:42:30+00:00", "updated_at": "2026-08-20 23:15:07.469713+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products", "ai-tools"], "entities": ["OpenAI", "GPT-5.6 Sol Pro", "SpaceXAI", "Grok 4.6", "Anthropic", "Claude Opus 4.8", "Google", "Gemini 3.7 Flash"], "alternates": {"html": "https://wpnews.pro/news/gpt-5-6-sol-pro-tops-editorial-craft-benchmark-with-0-97-score", "markdown": "https://wpnews.pro/news/gpt-5-6-sol-pro-tops-editorial-craft-benchmark-with-0-97-score.md", "text": "https://wpnews.pro/news/gpt-5-6-sol-pro-tops-editorial-craft-benchmark-with-0-97-score.txt", "jsonld": "https://wpnews.pro/news/gpt-5-6-sol-pro-tops-editorial-craft-benchmark-with-0-97-score.jsonld"}}