cd /news/artificial-intelligence/grok-4-6-tops-newsroom-reliability-v… · home topics artificial-intelligence article
[ARTICLE · art-96469] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.79

SpaceXAI's Grok 4.6 topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.79 across 50 tasks at an estimated $0.0056 per task, according to RuntimeWire's leaderboard. OpenAI's GPT-5.6 Sol followed at 0.77, with other GPT-5.6 variants and Anthropic's Claude Opus 4.8 clustering closely behind.

read2 min views1 publishedAug 14, 2026
Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.79
Image: Runtimewire (auto-discovered)

In a 50-task run of Newsroom Reliability v0.2, SpaceXAI’s Grok 4.6 ranked first with a score of 0.79 at an estimated $0.0056 per task. OpenAI’s GPT-5.6 Sol followed at 0.77, while several GPT-5.6 variants and Anthropic’s Claude Opus 4.8 clustered close behind.

By RuntimeWire Staff · Published

Leaderboard

| Rank | Model | Mean score | Est. cost/task | | 1 | SpaceXAI: Grok 4.6 | 0.79 | $0.0056 | | 2 | OpenAI: GPT-5.6 Sol | 0.77 | $0.0217 | | 3 | OpenAI: GPT-5.6 Sol Pro | 0.76 | $0.0217 | | 4 | OpenAI: GPT-5.6 Luna Pro | 0.76 | $0.0005 | | 5 | OpenAI: GPT-5.6 Luna | 0.76 | $0.0005 | | 6 | OpenAI: GPT-5.6 Terra | 0.75 | $0.0045 | | 7 | Anthropic: Claude Opus 4.8 | 0.74 | $0.0223 | | 8 | Inkling Small | 0.73 | — | | 9 | OpenAI: GPT-5.6 Terra Pro | 0.72 | $0.0045 | | 10 | Anthropic: Claude Opus 4.8 (Fast) | 0.72 | $0.0436 | | 11 | Kimi K3 | 0.71 | — | | 12 | Inkling FP4 | 0.70 | — | | 13 | GLM 5.2 | 0.69 | — | | 14 | Google: Gemini 3.6 Flash | 0.67 | $0.0056 | | 15 | Gemini 3.7 Flash | 0.66 | — | | 16 | DeepSeek-V4-Pro | 0.63 | — | | 17 | Qwen: Qwen3.7 Max | 0.62 | $0.0040 | | 18 | DeepSeek-V4-Flash | 0.61 | — |

How we scored it

Every model answered the same 50-task battery from Newsroom Reliability v0.2, one task at a time, with no tools and no retries on content.

50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the interactive leaderboard.

Reader comments #

Conversation for this story loads after sign-in.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @spacexai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/grok-4-6-tops-newsro…] indexed:0 read:2min 2026-08-14 ·