{"slug": "claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark", "title": "Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark", "summary": "Claude Opus 5 (Fast) topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.73 at an estimated $0.0460 per task, according to a 50-task run by RuntimeWire. Z.ai's GLM 5.3 Flash, Google's Gemini 3.7 Flash, and Qwen's Qwen3.8 Flash tied for second at 0.69, with GLM 5.3 Flash costing just $0.0003 per task and Gemini 3.7 Flash at $0.0014.", "body_md": "# Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark\n\n**In a 50-task run of Newsroom Reliability v0.2, Claude Opus 5 (Fast) led the field with a 0.73 score at $0.0460 per task. GLM 5.3 Flash, Gemini 3.7 Flash, and Qwen3.8 Flash followed at 0.69, each at substantially lower cost.**\n\nBy [RuntimeWire Staff](/author/runtimewire-staff)\n· Published\n\n### Leaderboard\n\n| Rank | Model | Mean score | Est. cost/task |\n|---|---|---|---|\n| 1 | Claude Opus 5 (Fast) | 0.73 | $0.0460 |\n| 2 | Z.ai: GLM 5.3 Flash | 0.69 | $0.0003 |\n| 3 | Google: Gemini 3.7 Flash | 0.69 | $0.0014 |\n| 4 | Qwen: Qwen3.8 Flash | 0.69 | $0.0005 |\n| 5 | step-3.7-flash | 0.59 | — |\n\n### How we scored it\n\nEvery model answered the same 50-task battery from **Newsroom Reliability v0.2**, one task at a time, with no tools and no retries on content.\n\n50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.\n\nReasoning effort was pinned to **max** for every model whose lane exposes a control (OpenAI-style `reasoning_effort`\n\n, Anthropic extended thinking, Gemini thinking config). Models marked *vendor default* on the interactive board expose no such control and ran as shipped.\n\nA model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing (\"—\" where pricing isn't public), so treat it as directional, not billing-exact.\n\nExplore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/newsroom-reliability-v0-2-leaderboard-4).", "url": "https://wpnews.pro/news/claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark", "canonical_source": "https://runtimewire.com/article/claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark", "published_at": "2026-08-27 01:47:32+00:00", "updated_at": "2026-08-27 02:18:51.400145+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-products"], "entities": ["Claude Opus 5 (Fast)", "RuntimeWire", "Z.ai", "GLM 5.3 Flash", "Google", "Gemini 3.7 Flash", "Qwen", "Qwen3.8 Flash"], "alternates": {"html": "https://wpnews.pro/news/claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark", "markdown": "https://wpnews.pro/news/claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark.md", "text": "https://wpnews.pro/news/claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark.txt", "jsonld": "https://wpnews.pro/news/claude-opus-5-fast-tops-newsroom-reliability-v0-2-benchmark.jsonld"}}