Claude Opus 5 (Fast) tops Newsroom Reliability v0.2 benchmark Claude Opus 5 (Fast) topped the Newsroom Reliability v0.2 benchmark with a mean score of 0.73 at an estimated $0.0460 per task, according to a 50-task run by RuntimeWire. Z.ai's GLM 5.3 Flash, Google's Gemini 3.7 Flash, and Qwen's Qwen3.8 Flash tied for second at 0.69, with GLM 5.3 Flash costing just $0.0003 per task and Gemini 3.7 Flash at $0.0014. Claude Opus 5 Fast tops Newsroom Reliability v0.2 benchmark In a 50-task run of Newsroom Reliability v0.2, Claude Opus 5 Fast led the field with a 0.73 score at $0.0460 per task. GLM 5.3 Flash, Gemini 3.7 Flash, and Qwen3.8 Flash followed at 0.69, each at substantially lower cost. By RuntimeWire Staff /author/runtimewire-staff · Published Leaderboard | Rank | Model | Mean score | Est. cost/task | |---|---|---|---| | 1 | Claude Opus 5 Fast | 0.73 | $0.0460 | | 2 | Z.ai: GLM 5.3 Flash | 0.69 | $0.0003 | | 3 | Google: Gemini 3.7 Flash | 0.69 | $0.0014 | | 4 | Qwen: Qwen3.8 Flash | 0.69 | $0.0005 | | 5 | step-3.7-flash | 0.59 | — | How we scored it Every model answered the same 50-task battery from Newsroom Reliability v0.2 , one task at a time, with no tools and no retries on content. 50 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale. Reasoning effort was pinned to max for every model whose lane exposes a control OpenAI-style reasoning effort , Anthropic extended thinking, Gemini thinking config . Models marked vendor default on the interactive board expose no such control and ran as shipped. A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing "—" where pricing isn't public , so treat it as directional, not billing-exact. Explore every prompt, answer, and per-task grade in the interactive leaderboard /benchmarks/newsroom-reliability-v0-2-leaderboard-4 .