cd /news/artificial-intelligence/z-ai-glm-5-3-flash-tops-editorial-cr… · home topics artificial-intelligence article
[ARTICLE · art-112510] src=runtimewire.com ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Z.ai GLM 5.3 Flash tops Editorial Craft benchmark at 0.94

Z.ai's GLM 5.3 Flash ranked first in a 12-task Editorial Craft benchmark with a mean score of 0.94 and an estimated cost of $0.0002 per task, according to RuntimeWire. Step-3.7-flash followed at 0.92, while Qwen3.8 Flash, Claude Opus 5 (Fast), Gemini 3.7 Flash, and DeepSeek-V4-Flash-0731 rounded out the leaderboard. The evaluation scored 2 tasks deterministically and 10 open-ended tasks via gpt-5.4 against a fixed rubric, with reasoning effort set to max where available.

read2 min views1 publishedAug 27, 2026
Z.ai GLM 5.3 Flash tops Editorial Craft benchmark at 0.94
Image: Runtimewire (auto-discovered)

In a 12-task Editorial Craft evaluation, Z.ai’s GLM 5.3 Flash ranked first with a 0.94 score and an estimated cost of $0.0002 per task. Step-3.7-flash followed at 0.92, while Qwen3.8 Flash, Claude Opus 5 (Fast), Gemini 3.7 Flash, and DeepSeek-V4-Flash-0731 rounded out the leaderboard.

By RuntimeWire Staff · Published

Leaderboard

Rank Model Mean score Est. cost/task
1 Z.ai: GLM 5.3 Flash 0.94 $0.0002
2 step-3.7-flash 0.92
3 Qwen: Qwen3.8 Flash 0.90 $0.0003
4 Claude Opus 5 (Fast) 0.88 $0.0334
5 Google: Gemini 3.7 Flash 0.87 $0.0011

| 6 | DeepSeek-V4-Flash-0731 | 0.72 | — |

How we scored it

Every model answered the same 12-task battery from Editorial Craft, one task at a time, with no tools and no retries on content.

2 tasks carry ground truth (exact, numeric, multiple-choice, or the benchmark's official matcher) and were scored deterministically against the reference answer — correct is 1, incorrect is 0.

10 open-ended tasks were graded 0–10 by gpt-5.4 against a fixed rubric — blind to which model wrote the answer — and normalized to the same 0–1 scale.

Reasoning effort was pinned to max for every model whose lane exposes a control (OpenAI-style reasoning_effort

, Anthropic extended thinking, Gemini thinking config). Models marked vendor default on the interactive board expose no such control and ran as shipped.

A model's mean score averages its graded tasks; a generation failure counts as 0. Cost per task is estimated from each model's published per-token pricing ("—" where pricing isn't public), so treat it as directional, not billing-exact.

Explore every prompt, answer, and per-task grade in the interactive leaderboard.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @z.ai 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/z-ai-glm-5-3-flash-t…] indexed:0 read:2min 2026-08-27 ·