Thibaud Colas of the Wagtail core team spent September 2026 coding on GLM 5.3 Flash, Z.ai's MIT-licensed model with 320B parameters and 18B active, and published the numbers on 2 October: about 1B tokens on Flash for $68 and about 4 kWh of GPU energy. Flash kept the default slot rather than a planner or executor role. Wagtail now recommends it for 95% of tasks and GLM 5.3 for the hardest 5%. On Wagtail's own 20-task benchmark it passed 11, at a median $0.16 a task.
The run is his, not mine, and it did not use Claude Code. He drove Flash from omp, billed per token through two European providers. The Claude Code and OpenCode setup at the end comes from Z.ai's and Anthropic's docs.
What is GLM 5.3 Flash, and what does a task cost? #
GLM 5.3 Flash is Z.ai's coding and vision model, released on 26 August 2026: 320B total parameters, 18B active, MIT licence, a 1M-token context and thinking that cannot be switched off. Z.ai and OpenRouter list it at $0.15 per million input tokens, $0.03 cached and $0.50 output. On Wagtail's benchmark a task cost a median $0.16; Z.ai quotes $0.045 on Artificial Analysis's index tasks, a different task set.
| Per million tokens, read 7 October 2026 | Input | Output |
|---|---|---|
| GLM 5.3 Flash, Z.ai API | $0.15 | $0.50 |
| GLM 5.3 Flash, TensorX (the author's provider) | $0.20 | $0.50 |
| GLM 5.3, Z.ai API | $1.40 | $4.40 |
| Claude Sonnet 5.5 | $2 | $10 |
| Claude Opus 5.5 | $4 | $20 |
Z.ai's GLM Coding Plan costs $18, $80 or $168 a month, and Flash spends a third of the plan credits GLM 5.3 does per token, which is where the "3x quota" in Z.ai's marketing comes from. The plan works only inside listed tools: Claude Code, OpenCode, Pi, Codex and about a dozen others, not omp. Z.ai's own launch table puts Flash at 84.3 on Terminal-Bench 2.1, run in Claude Code, against 85.0 for Opus 4.8. Z.ai publishes no SWE-bench Verified score for it.
The stack: omp, two EU providers and per-token billing #
Thibaud Colas ran GLM 5.3 Flash in the omp coding agent, on two inference providers serving from European data centres, and paid per token with no subscription. The month was a challenge he set in a 3 September post: one efficient open-weight model for all of September, and "ditch your AI coding subscription".
- Harness: omp (oh-my-pi), confirmed by the authoron Hacker News .
- Model: GLM 5.3 Flash, with GLM 5.3 for harder work, andDeepSeek V4.1 Flash andQwen 3.8 Flash as fallbacks.
- Inference:TensorX andNeuralwatt , usage-based billing only.
- Metering:AgentsView for tokens and cost, whose costs Wagtail calls "indicative only". Energy and carbon come from Neuralwatt and count GPU energy only.
Neuralwatt supplied the energy figures, and its CTO said in the Hacker News thread that his company was likely the provider used. The post does not say whether the $68 is a provider bill or an AgentsView estimate.
The first half of September ran on Flash alone #
GLM 5.3 Flash carried all of Colas's work for the first two weeks: "Wagtail itself, sites built with it. UI tasks, AI R&D, also a bit of documentation writing. Lots of evals." He picked it for three reasons: a 1M context for long tasks, vision for building from screenshots, and many providers competing on price.
His verdict: "GLM 5.3 Flash itself was excellent!" For day-to-day work, he wrote, "it's totally viable to focus on one or two flash-tier cheap models."
One overnight build on GLM 5.3 took 150% of the budget #
The month's largest bill came from a model mix-up, not from Flash. Colas vibe-coded a Wagtail MCP server prototype overnight in omp and left it on the full GLM 5.3, which lists at about nine times Flash's input price. That session used 450M tokens, $150 and about 5 kWh.
He explained it on Hacker News: "I didn't realize that one session was a quarter of the month's spend and 150% of the budget". He estimates Flash would have done similar work "for most likely 5x less cost".
Provider capacity moved the second half to other models #
In the second half of September, Colas saw GLM 5.3 Flash's performance degrade at the European providers he used, which he put down to demand: popular open models "do not have the same capacity as the big labs who hoard all the GPUs." He moved work to DeepSeek V4.1 Flash and Qwen 3.8 Flash, and R&D needed other models, so Flash ended the month at 1B of 2B tokens.
What September cost, by model, from the post and its AgentsView screenshot:
| Line | Tokens | Cost | Energy |
|---|---|---|---|
| GLM 5.3 Flash, all month | 1.04B | $68 | about 4 kWh |
| GLM 5.3, the overnight MCP build | 451M | $150 | about 5 kWh |
| GPT-6 Astra, DeepSeek V4.1 Flash, Claude models and others | about 570M | Not published | Not published |
| September total | 2B | Not published | about 35 kWh, or "about 30kWh" in a later comment |
AgentsView shows an 84.1% cache hit rate across the month, which is how 1.04B Flash tokens came to $68: about $0.065 per million, blended, by my arithmetic. Colas says he is moving to DeepSeek V4.1 Flash in October, which "feels slightly better from a few days of use".
Wagtail's benchmark passed GLM 5.3 Flash on 11 of 20 tasks #
Wagtail's work-in-progress benchmark of 20 Wagtail tasks is the number the vendor tables lack: GLM 5.3 Flash passed 11 at a median $0.16 and 26.2 Wh a task, while DeepSeek V4.1 Flash, Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6 Sol each passed 19. The table is an image in the post, and its harness and effort level are not given.
| Model | Passed | Median cost per task | Energy per task |
|---|---|---|---|
| DeepSeek V4.1 Flash | 19 of 20 | $0.09 | 14.9 Wh |
| Claude Sonnet 5.5 | 19 of 20 | $0.14 | Not measured |
| GPT-6 Sol | 19 of 20 | $0.15 | Not measured |
| Claude Opus 5.5 | 19 of 20 | $0.55 | Not measured |
| GLM 5.3 | 14 of 20 | $0.87 | 56.9 Wh |
| GLM 5.3 Flash | 11 of 20 | $0.16 | 26.2 Wh |
Divided by the pass rate, Flash costs about $0.29 per passing task and Sonnet 5.5 about $0.15, by my arithmetic. Twenty tasks is a small sample. Effort level moves the cost too: HN user lhl, on about 300 tasks, found medium effort the best dollars per pass, with max scoring about 5% higher at more than twice the cost.
GLM 5.3 Flash kept the default slot, and commenters made it the executor #
For its author, GLM 5.3 Flash is neither planner nor executor: it is the default for routine work, with GLM 5.3 kept for the hardest 5%, and he added that he has "a hard time justifying GLM 5.3 these days". Commenters who split the roles mostly put Flash in the executor slot behind a stronger planner.
From the thread, 180 comments by 7 October:
- RussianCow: GLM 5.3 writes "a detailed plan with little to no ambiguity", and Flash "executes it just fine for a fraction of the price".
- surgical_fire: Flash is the implementation model after GLM 5.3 plans, in separate sessions that pass markdown files.
- throw930rmdkdk: Flash codes, with one model planning and another reviewing.
- esafak , on Z.ai's plan: the reverse, Flash for planning and review "because it's too slow for execution".
Colas's October plan points the same way: "Orchestrator vs. scout vs. implementer vs. reviewer agents." The model routing post has more September numbers on who sends which step to which model.
On Stackness, as of 7 October 2026, 7 real profiles list Claude Code and 2 list OpenCode, both alongside Claude Code, and none lists a GLM model (data sources). The numbers are small; the language models developers list will show when GLM appears.
Where Claude Code breaks on GLM 5.3 Flash #
Claude Code is built for Claude, and Anthropic's gateway docs say it does not support routing to non-Claude models. Pointed at GLM 5.3 Flash, Claude Code hits four known problems: requests that switch thinking off, a 200K context assumption, web search that only runs on Anthropic's backend, and tool calls that come back empty.
| Symptom | Cause | Fix |
|---|---|---|
| HTTP 400, code 1210, on background requests | Flash cannot switch thinking off, and Claude Code turns it off for some requests ( cc-switch #6737 , late August) | Z.ai's docs now say such requests run at low effort; Claude Code also drops thinking for the rest of a conversation after a rejection |
| Compaction at 200K on a 1M model | Claude Code assumes 200K for model IDs it does not know | [1m] on the model ID andCLAUDE_CODE_AUTO_COMPACT_WINDOW |
| Web search fails | WebSearch runs on Anthropic's backend | Z.ai's web search MCP, 1.2 plan credits a call |
| Empty tool calls, then repeated 400s | Reported through a gateway with Pi ( llmgateway #4366 ) and inKilo Code (#14475 ) | Open, no fix published |
| The model stops thinking after about 50 steps and loops | Ollama Cloud did not send earlier thinking back ( masc PR ) | Z.ai recommends clear_thinking: false |
| Screenshots fail in the Sonnet and Opus slots | Z.ai's manual config maps them to GLM 5.3, which is text-only | Map those slots to Flash |
| A "Claude" run that ran on GLM | A silent step-down to GLM when no Claude seat was free ( ticfac #248 ) | Log the model behind every run |
Remote Control and the Advisor tool need Anthropic's API. The DeepSeek Harness post covers what else stays behind when you switch harness or model.
Run it yourself #
Claude Code runs GLM 5.3 Flash with four settings: Z.ai's Anthropic-compatible base URL, a Z.ai API key, the Flash model ID with [1m] in all three model slots, and a 1M-token compaction window, plus the longer timeout Z.ai sets. Z.ai documents the endpoint, key and timeout on its Claude Code page and the model IDs and window on its latest-model page. Add them to ~/.claude/settings.json rather than replacing the file:
{
"env": {
"ANTHROPIC_AUTH_TOKEN": "your_zai_api_key",
"ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
"ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
"ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3-flash[1m]",
"ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3-flash[1m]",
"CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
"API_TIMEOUT_MS": "3000000"
}
}
/status should then show glm-5.3-flash[1m] as the model.
To let GLM 5.3 plan and Flash execute, set ANTHROPIC_DEFAULT_OPUS_MODEL to glm-5.3[1m] and switch to /model opusplan, which Anthropic's model docs describe as the Opus model in plan mode and the Sonnet model otherwise. I worked this out from the docs and have not run it.
Keeping Claude as the planner takes two sessions with a handoff file, the setup surgical_fire describes, or a gateway that routes by model name, such as claude-code-router, which Anthropic does not support. MoAI-ADK retired its Claude-leader-with-GLM-teammates mode and now runs one lane per model, moai cc -l and moai glm -l.
OpenCode and Ollama, from Z.ai's OpenCode page and Ollama's library page:
curl -fsSL https://opencode.ai/install | bash
opencode auth login # pick Z.AI Coding Plan, then /models in a session
ollama launch claude --model glm-5.3-flash:cloud
Ollama has only a cloud tag for Flash. Colibri serves it locally by streaming experts from disk, at about 20 seconds per token in its README. Wagtail's TensorX configs for omp and OpenCode are in its wiki, with one tip: on a "quota reached" error, set max output tokens to 16384 or 32768.
Tools in this post #
Use any of these tools? #
Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.