{"slug": "grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-79", "title": "Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.79", "summary": "SpaceXAI's Grok 4.6 topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.79 across 50 tasks at an estimated $0.0056 per task, according to RuntimeWire's leaderboard. OpenAI's GPT-5.6 Sol followed at 0.77, with other GPT-5.6 variants and Anthropic's Claude Opus 4.8 clustering closely behind.", "body_md": "# Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.79\n\n**In a 50-task run of Newsroom Reliability v0.2, SpaceXAI’s Grok 4.6 ranked first with a score of 0.79 at an estimated $0.0056 per task. OpenAI’s GPT-5.6 Sol followed at 0.77, while several GPT-5.6 variants and Anthropic’s Claude Opus 4.8 clustered close behind.**\n\nBy [RuntimeWire Staff](/author/runtimewire-staff)\n· Published\n\n### Leaderboard\n\n| Rank |\nModel |\nMean score |\nEst. cost/task |\n| 1 |\nSpaceXAI: Grok 4.6 |\n0.79 |\n$0.0056 |\n| 2 |\nOpenAI: GPT-5.6 Sol |\n0.77 |\n$0.0217 |\n| 3 |\nOpenAI: GPT-5.6 Sol Pro |\n0.76 |\n$0.0217 |\n| 4 |\nOpenAI: GPT-5.6 Luna Pro |\n0.76 |\n$0.0005 |\n| 5 |\nOpenAI: GPT-5.6 Luna |\n0.76 |\n$0.0005 |\n| 6 |\nOpenAI: GPT-5.6 Terra |\n0.75 |\n$0.0045 |\n| 7 |\nAnthropic: Claude Opus 4.8 |\n0.74 |\n$0.0223 |\n| 8 |\nInkling Small |\n0.73 |\n— |\n| 9 |\nOpenAI: GPT-5.6 Terra Pro |\n0.72 |\n$0.0045 |\n| 10 |\nAnthropic: Claude Opus 4.8 (Fast) |\n0.72 |\n$0.0436 |\n| 11 |\nKimi K3 |\n0.71 |\n— |\n| 12 |\nInkling FP4 |\n0.70 |\n— |\n| 13 |\nGLM 5.2 |\n0.69 |\n— |\n| 14 |\nGoogle: Gemini 3.6 Flash |\n0.67 |\n$0.0056 |\n| 15 |\nGemini 3.7 Flash |\n0.66 |\n— |\n| 16 |\nDeepSeek-V4-Pro |\n0.63 |\n— |\n| 17 |\nQwen: Qwen3.7 Max |\n0.62 |\n$0.0040 |\n| 18 |\nDeepSeek-V4-Flash |\n0.61 |\n— |\n\n### How we scored it\n\nEvery model answered the same 50-task battery from **Newsroom Reliability v0.2**, one task at a time, with no tools and no retries on content.\n\n50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.\n\nA model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing (\"—\" where pricing isn't public), so treat it as directional, not billing-exact.\n\nExplore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/newsroom-reliability-v0-2-leaderboard).\n\n## Reader comments\n\nConversation for this story loads after sign-in.", "url": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-79", "canonical_source": "https://runtimewire.com/article/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-79", "published_at": "2026-08-14 07:34:31+00:00", "updated_at": "2026-08-14 07:36:45.852199+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research"], "entities": ["SpaceXAI", "Grok 4.6", "OpenAI", "GPT-5.6 Sol", "Anthropic", "Claude Opus 4.8", "RuntimeWire", "Newsroom Reliability v0.2"], "alternates": {"html": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-79", "markdown": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-79.md", "text": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-79.txt", "jsonld": "https://wpnews.pro/news/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-79.jsonld"}}