{"slug": "z-ai-glm-5-3-flash-tops-editorial-craft-benchmark-at-0-94", "title": "Z.ai GLM 5.3 Flash tops Editorial Craft benchmark at 0.94", "summary": "Z.ai's GLM 5.3 Flash ranked first in a 12-task Editorial Craft benchmark with a mean score of 0.94 and an estimated cost of $0.0002 per task, according to RuntimeWire. Step-3.7-flash followed at 0.92, while Qwen3.8 Flash, Claude Opus 5 (Fast), Gemini 3.7 Flash, and DeepSeek-V4-Flash-0731 rounded out the leaderboard. The evaluation scored 2 tasks deterministically and 10 open-ended tasks via gpt-5.4 against a fixed rubric, with reasoning effort set to max where available.", "body_md": "# Z.ai GLM 5.3 Flash tops Editorial Craft benchmark at 0.94\n\n**In a 12-task Editorial Craft evaluation, Z.ai’s GLM 5.3 Flash ranked first with a 0.94 score and an estimated cost of $0.0002 per task. Step-3.7-flash followed at 0.92, while Qwen3.8 Flash, Claude Opus 5 (Fast), Gemini 3.7 Flash, and DeepSeek-V4-Flash-0731 rounded out the leaderboard.**\n\nBy [RuntimeWire Staff](/author/runtimewire-staff)\n· Published\n\n### Leaderboard\n\n| Rank | Model | Mean score | Est. cost/task |\n|---|---|---|---|\n| 1 | Z.ai: GLM 5.3 Flash | 0.94 | $0.0002 |\n| 2 | step-3.7-flash | 0.92 | — |\n| 3 | Qwen: Qwen3.8 Flash | 0.90 | $0.0003 |\n| 4 | Claude Opus 5 (Fast) | 0.88 | $0.0334 |\n| 5 | Google: Gemini 3.7 Flash | 0.87 | $0.0011 |\n| 6 | DeepSeek-V4-Flash-0731 | 0.72 | — |\n\n### How we scored it\n\nEvery model answered the same 12-task battery from **Editorial Craft**, one task at a time, with no tools and no retries on content.\n\n2 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.\n\n10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.\n\nReasoning effort was pinned to **max** for every model whose lane exposes a control (OpenAI-style `reasoning_effort`\n\n, Anthropic extended thinking, Gemini thinking config). Models marked *vendor default* on the interactive board expose no such control and ran as shipped.\n\nA model's **mean score** averages its graded tasks; a generation failure counts as 0. **Cost per task** is estimated from each model's published per-token pricing (\"—\" where pricing isn't public), so treat it as directional, not billing-exact.\n\nExplore every prompt, answer, and per-task grade in the [interactive leaderboard](/benchmarks/editorial-craft-leaderboard).", "url": "https://wpnews.pro/news/z-ai-glm-5-3-flash-tops-editorial-craft-benchmark-at-0-94", "canonical_source": "https://runtimewire.com/article/z-ai-glm-5-3-flash-tops-editorial-craft-benchmark-at-0-94", "published_at": "2026-08-27 00:45:16+00:00", "updated_at": "2026-08-27 00:48:40.622714+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-research", "ai-products"], "entities": ["Z.ai", "GLM 5.3 Flash", "Step-3.7-flash", "Qwen3.8 Flash", "Claude Opus 5 (Fast)", "Gemini 3.7 Flash", "DeepSeek-V4-Flash-0731", "RuntimeWire"], "alternates": {"html": "https://wpnews.pro/news/z-ai-glm-5-3-flash-tops-editorial-craft-benchmark-at-0-94", "markdown": "https://wpnews.pro/news/z-ai-glm-5-3-flash-tops-editorial-craft-benchmark-at-0-94.md", "text": "https://wpnews.pro/news/z-ai-glm-5-3-flash-tops-editorial-craft-benchmark-at-0-94.txt", "jsonld": "https://wpnews.pro/news/z-ai-glm-5-3-flash-tops-editorial-craft-benchmark-at-0-94.jsonld"}}