{"slug": "grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78", "title": "Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.78", "summary": "SpaceXAI's Grok 4.6 topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.78, outperforming OpenAI's GPT-5.6 Luna Pro (0.77) and GPT-5.6 Sol Pro (0.76), Meta's Muse Spark 1.2 (0.73), Anthropic's Claude Opus 4.8 (0.71), and Google's Gemini 3.7 Flash (0.67). The benchmark evaluated 50 open-ended tasks graded by gpt-5.4, with costs per task ranging from $0.0005 for GPT-5.6 Luna Pro to $0.0215 for GPT-5.6 Sol Pro.", "body_md": "### Leaderboard\n\n| Rank |\nModel |\nMean score |\nEst. cost/task |\n| 1 |\nSpaceXAI: Grok 4.6 |\n0.78 |\n$0.0056 |\n| 2 |\nOpenAI: GPT-5.6 Luna Pro |\n0.77 |\n$0.0005 |\n| 3 |\nOpenAI: GPT-5.6 Sol Pro |\n0.76 |\n$0.0215 |\n| 4 |\nMeta: Muse Spark 1.2 |\n0.73 |\n$0.0044 |\n| 5 |\nAnthropic: Claude Opus 4.8 |\n0.71 |\n$0.0211 |\n| 6 |\nGoogle: Gemini 3.7 Flash |\n0.67 |\n$0.0014 |\n\n### How we scored it\n\nEvery model answered the same 50-task battery from **Newsroom Reliability v0.2**, one task at a time, with no tools and no retries on content.\n\n50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.\n\nReasoning effort was pinned to **max** for every model whose lane exposes a control (OpenAI-style `reasoning_effort`\n\n, Anthropic extended thinking, Gemini thinking config). Models marked *vendor default* on the interactive board expose no such control and ran as shipped.\n\nA model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing (\"—\" where pricing isn't public), so treat it as directional, not billing-exact.\n\nExplore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/newsroom-reliability-v0-2-leaderboard-2).", "url": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78", "canonical_source": "https://runtimewire.com/article/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78", "published_at": "2026-08-15 19:21:15+00:00", "updated_at": "2026-08-15 19:40:58.175418+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence", "ai-research"], "entities": ["SpaceXAI", "OpenAI", "Meta", "Anthropic", "Google", "Grok 4.6", "GPT-5.6 Luna Pro", "Claude Opus 4.8"], "alternates": {"html": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78", "markdown": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78.md", "text": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78.txt", "jsonld": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78.jsonld"}}