cd /news/large-language-models/claude-opus-5-fast-tops-newsroom-rel… Β· home β€Ί topics β€Ί large-language-models β€Ί article
[ARTICLE Β· art-112553] src=runtimewire.com β†— pub= topic=large-language-models verified=true sentiment=Β· neutral

Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark

Claude Opus 5 (Fast) topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.73 at an estimated $0.0460 per task, according to a 50-task run by RuntimeWire. Z.ai's GLM 5.3 Flash, Google's Gemini 3.7 Flash, and Qwen's Qwen3.8 Flash tied for second at 0.69, with GLM 5.3 Flash costing just $0.0003 per task and Gemini 3.7 Flash at $0.0014.

read1 min views29 publishedAug 27, 2026
Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark
Image: Runtimewire (auto-discovered)

In a 50-task run of Newsroom Reliability v0.2, Claude Opus 5 (Fast) led the field with a 0.73 score at $0.0460 per task. GLM 5.3 Flash, Gemini 3.7 Flash, and Qwen3.8 Flash followed at 0.69, each at substantially lower cost.

By RuntimeWire Staff Β· Published

Leaderboard

Rank Model Mean score Est. cost/task
1 Claude Opus 5 (Fast) 0.73 $0.0460
2 Z.ai: GLM 5.3 Flash 0.69 $0.0003
3 Google: Gemini 3.7 Flash 0.69 $0.0014
4 Qwen: Qwen3.8 Flash 0.69 $0.0005

| 5 | step-3.7-flash | 0.59 | β€” |

How we scored it

Every model answered the same 50-task battery from Newsroom Reliability v0.2, one task at a time, with no tools and no retries on content.

50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric β€” blind to which model wrote the answer β€” and normalized to the same 0–1 scale.

Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort

, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("β€”" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the interactive leaderboard.

── more in #large-language-models 4 stories Β· sorted by recency
── more on @claude opus 5 (fast) 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain β€” perfect for shipping the agent you just read about.

$git push zahid main
β†’ Live at https://your-agent.zahid.host βœ“
Get free account β†’ Pricing
from €0/mo Β· no card required
LIVE [news/claude-opus-5-fast-t…] indexed:0 read:1min 2026-08-27 Β· β€”