# Grok 4.6 tops Newsroom Reliability v0.2 benchmark at 0.78

> Source: <https://runtimewire.com/article/grok-4-6-tops-newsroom-reliability-v0-2-benchmark-at-0-78>
> Published: 2026-08-15 19:21:15+00:00

### Leaderboard

| Rank |
Model |
Mean score |
Est. cost/task |
| 1 |
SpaceXAI: Grok 4.6 |
0.78 |
$0.0056 |
| 2 |
OpenAI: GPT-5.6 Luna Pro |
0.77 |
$0.0005 |
| 3 |
OpenAI: GPT-5.6 Sol Pro |
0.76 |
$0.0215 |
| 4 |
Meta: Muse Spark 1.2 |
0.73 |
$0.0044 |
| 5 |
Anthropic: Claude Opus 4.8 |
0.71 |
$0.0211 |
| 6 |
Google: Gemini 3.7 Flash |
0.67 |
$0.0014 |

### How we scored it

Every model answered the same 50-task battery from **Newsroom Reliability v0.2**, one task at a time, with no tools and no retries on content.

50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

Reasoning effort was pinned to **max** for every model whose lane exposes a control (OpenAI-style `reasoning_effort`

, Anthropic extended thinking, Gemini thinking config). Models marked *vendor default* on the interactive board expose no such control and ran as shipped.

A model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/newsroom-reliability-v0-2-leaderboard-2).
