Z.ai GLM 5.3 Flash tops Editorial Craft benchmark at 0.94 Z.ai's GLM 5.3 Flash ranked first in a 12-task Editorial Craft benchmark with a mean score of 0.94 and an estimated cost of $0.0002 per task, according to RuntimeWire. Step-3.7-flash followed at 0.92, while Qwen3.8 Flash, Claude Opus 5 (Fast), Gemini 3.7 Flash, and DeepSeek-V4-Flash-0731 rounded out the leaderboard. The evaluation scored 2 tasks deterministically and 10 open-ended tasks via gpt-5.4 against a fixed rubric, with reasoning effort set to max where available. Z.ai GLM 5.3 Flash tops Editorial Craft benchmark at 0.94 In a 12-task Editorial Craft evaluation, Z.ai’s GLM 5.3 Flash ranked first with a 0.94 score and an estimated cost of $0.0002 per task. Step-3.7-flash followed at 0.92, while Qwen3.8 Flash, Claude Opus 5 Fast , Gemini 3.7 Flash, and DeepSeek-V4-Flash-0731 rounded out the leaderboard. By RuntimeWire Staff /author/runtimewire-staff · Published Leaderboard | Rank | Model | Mean score | Est. cost/task | |---|---|---|---| | 1 | Z.ai: GLM 5.3 Flash | 0.94 | $0.0002 | | 2 | step-3.7-flash | 0.92 | — | | 3 | Qwen: Qwen3.8 Flash | 0.90 | $0.0003 | | 4 | Claude Opus 5 Fast | 0.88 | $0.0334 | | 5 | Google: Gemini 3.7 Flash | 0.87 | $0.0011 | | 6 | DeepSeek-V4-Flash-0731 | 0.72 | — | How we scored it Every model answered the same 12-task battery from Editorial Craft , one task at a time, with no tools and no retries on content. 2 tasks carry ground truth exact, numeric, multiple-choice, or the benchmark's official matcher and were scored deterministically against the reference answer — correct is 1, incorrect is 0. 10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale. Reasoning effort was pinned to max for every model whose lane exposes a control OpenAI-style reasoning effort , Anthropic extended thinking, Gemini thinking config . Models marked vendor default on the interactive board expose no such control and ran as shipped. A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing "—" where pricing isn't public , so treat it as directional, not billing-exact. Explore every prompt, answer, and per-task grade in the interactive leaderboard /benchmarks/editorial-craft-leaderboard .