cd /news/large-language-models/claude-opus-5-fast-tops-newsroom-rel… · home topics large-language-models article
[ARTICLE · art-112553] src=runtimewire.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark

Claude Opus 5 (Fast) topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.73 at an estimated $0.0460 per task, according to a 50-task run by RuntimeWire. Z.ai's GLM 5.3 Flash, Google's Gemini 3.7 Flash, and Qwen's Qwen3.8 Flash tied for second at 0.69, with GLM 5.3 Flash costing just $0.0003 per task and Gemini 3.7 Flash at $0.0014.

read1 min views1 publishedAug 27, 2026
Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark
Image: Runtimewire (auto-discovered)

In a 50-task run of Newsroom Reliability v0.2, Claude Opus 5 (Fast) led the field with a 0.73 score at $0.0460 per task. GLM 5.3 Flash, Gemini 3.7 Flash, and Qwen3.8 Flash followed at 0.69, each at substantially lower cost.

By RuntimeWire Staff · Published

Leaderboard

Rank Model Mean score Est. cost/task
1 Claude Opus 5 (Fast) 0.73 $0.0460
2 Z.ai: GLM 5.3 Flash 0.69 $0.0003
3 Google: Gemini 3.7 Flash 0.69 $0.0014
4 Qwen: Qwen3.8 Flash 0.69 $0.0005

| 5 | step-3.7-flash | 0.59 | — |

How we scored it

Every model answered the same 50-task battery from Newsroom Reliability v0.2, one task at a time, with no tools and no retries on content.

50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort

, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the interactive leaderboard.

── more in #large-language-models 4 stories · sorted by recency
── more on @claude opus 5 (fast) 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-opus-5-fast-t…] indexed:0 read:1min 2026-08-27 ·