# Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark

> Source: <https://runtimewire.com/article/claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark>
> Published: 2026-08-27 01:47:32+00:00

# Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark

**In a 50-task run of Newsroom Reliability v0.2, Claude Opus 5 (Fast) led the field with a 0.73 score at $0.0460 per task. GLM 5.3 Flash, Gemini 3.7 Flash, and Qwen3.8 Flash followed at 0.69, each at substantially lower cost.**

By [RuntimeWire Staff](/author/runtimewire-staff)
· Published

### Leaderboard

| Rank | Model | Mean score | Est. cost/task |
|---|---|---|---|
| 1 | Claude Opus 5 (Fast) | 0.73 | $0.0460 |
| 2 | Z.ai: GLM 5.3 Flash | 0.69 | $0.0003 |
| 3 | Google: Gemini 3.7 Flash | 0.69 | $0.0014 |
| 4 | Qwen: Qwen3.8 Flash | 0.69 | $0.0005 |
| 5 | step-3.7-flash | 0.59 | — |

### How we scored it

Every model answered the same 50-task battery from **Newsroom Reliability v0.2**, one task at a time, with no tools and no retries on content.

50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

Reasoning effort was pinned to **max** for every model whose lane exposes a control (OpenAI-style `reasoning_effort`

, Anthropic extended thinking, Gemini thinking config). Models marked *vendor default* on the interactive board expose no such control and ran as shipped.

A model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/newsroom-reliability-v0-2-leaderboard-4).
