cd /news/ai-tools/glm-5-3-flash-as-a-coding-agent-a-mo… · home › topics › ai-tools › article
[ARTICLE · art-147123] src=stackness.dev ↗ pub= topic=ai-tools verified=true sentiment=↑ positive

GLM 5.3 Flash as a coding agent: a month of daily use, $68, and the slot it kept

Thibaud Colas of the Wagtail core team spent September 2026 coding daily on Z.ai's MIT-licensed GLM 5.3 Flash, consuming about 1B tokens for $68 and roughly 4 kWh of GPU energy, according to numbers he published on 2 October. Wagtail now recommends GLM 5.3 Flash for 95% of tasks and GLM 5.3 for the hardest 5%, and the model passed 11 of 20 tasks on Wagtail's own benchmark at a median cost of $0.16 per task. GLM 5.3 Flash, released 26 August 2026, has 320B total parameters, 18B active, a 1M-token context, and is priced at $0.15 per million input tokens and $0.50 per million output tokens on the Z.ai API.

by read10 min views2 publishedOct 7, 2026
GLM 5.3 Flash as a coding agent: a month of daily use, $68, and the slot it kept
Image: Stackness (auto-discovered)

Thibaud Colas of the Wagtail core team spent September 2026 coding on GLM 5.3 Flash, Z.ai's MIT-licensed model with 320B parameters and 18B active, and published the numbers on 2 October: about 1B tokens on Flash for $68 and about 4 kWh of GPU energy. Flash kept the default slot rather than a planner or executor role. Wagtail now recommends it for 95% of tasks and GLM 5.3 for the hardest 5%. On Wagtail's own 20-task benchmark it passed 11, at a median $0.16 a task.

The run is his, not mine, and it did not use Claude Code. He drove Flash from omp, billed per token through two European providers. The Claude Code and OpenCode setup at the end comes from Z.ai's and Anthropic's docs.

What is GLM 5.3 Flash, and what does a task cost? #

GLM 5.3 Flash is Z.ai's coding and vision model, released on 26 August 2026: 320B total parameters, 18B active, MIT licence, a 1M-token context and thinking that cannot be switched off. Z.ai and OpenRouter list it at $0.15 per million input tokens, $0.03 cached and $0.50 output. On Wagtail's benchmark a task cost a median $0.16; Z.ai quotes $0.045 on Artificial Analysis's index tasks, a different task set.

Per million tokens, read 7 October 2026 Input Output
GLM 5.3 Flash, Z.ai API $0.15 $0.50
GLM 5.3 Flash, TensorX (the author's provider) $0.20 $0.50
GLM 5.3, Z.ai API $1.40 $4.40
Claude Sonnet 5.5 $2 $10
Claude Opus 5.5 $4 $20

Z.ai's GLM Coding Plan costs $18, $80 or $168 a month, and Flash spends a third of the plan credits GLM 5.3 does per token, which is where the "3x quota" in Z.ai's marketing comes from. The plan works only inside listed tools: Claude Code, OpenCode, Pi, Codex and about a dozen others, not omp. Z.ai's own launch table puts Flash at 84.3 on Terminal-Bench 2.1, run in Claude Code, against 85.0 for Opus 4.8. Z.ai publishes no SWE-bench Verified score for it.

The stack: omp, two EU providers and per-token billing #

Thibaud Colas ran GLM 5.3 Flash in the omp coding agent, on two inference providers serving from European data centres, and paid per token with no subscription. The month was a challenge he set in a 3 September post: one efficient open-weight model for all of September, and "ditch your AI coding subscription".

  • Harness: omp (oh-my-pi), confirmed by the authoron Hacker News .
  • Model: GLM 5.3 Flash, with GLM 5.3 for harder work, andDeepSeek V4.1 Flash andQwen 3.8 Flash as fallbacks.
  • Inference:TensorX andNeuralwatt , usage-based billing only.
  • Metering:AgentsView for tokens and cost, whose costs Wagtail calls "indicative only". Energy and carbon come from Neuralwatt and count GPU energy only.

Neuralwatt supplied the energy figures, and its CTO said in the Hacker News thread that his company was likely the provider used. The post does not say whether the $68 is a provider bill or an AgentsView estimate.

The first half of September ran on Flash alone #

GLM 5.3 Flash carried all of Colas's work for the first two weeks: "Wagtail itself, sites built with it. UI tasks, AI R&D, also a bit of documentation writing. Lots of evals." He picked it for three reasons: a 1M context for long tasks, vision for building from screenshots, and many providers competing on price.

His verdict: "GLM 5.3 Flash itself was excellent!" For day-to-day work, he wrote, "it's totally viable to focus on one or two flash-tier cheap models."

One overnight build on GLM 5.3 took 150% of the budget #

The month's largest bill came from a model mix-up, not from Flash. Colas vibe-coded a Wagtail MCP server prototype overnight in omp and left it on the full GLM 5.3, which lists at about nine times Flash's input price. That session used 450M tokens, $150 and about 5 kWh.

He explained it on Hacker News: "I didn't realize that one session was a quarter of the month's spend and 150% of the budget". He estimates Flash would have done similar work "for most likely 5x less cost".

Provider capacity moved the second half to other models #

In the second half of September, Colas saw GLM 5.3 Flash's performance degrade at the European providers he used, which he put down to demand: popular open models "do not have the same capacity as the big labs who hoard all the GPUs." He moved work to DeepSeek V4.1 Flash and Qwen 3.8 Flash, and R&D needed other models, so Flash ended the month at 1B of 2B tokens.

What September cost, by model, from the post and its AgentsView screenshot:

Line Tokens Cost Energy
GLM 5.3 Flash, all month 1.04B $68 about 4 kWh
GLM 5.3, the overnight MCP build 451M $150 about 5 kWh
GPT-6 Astra, DeepSeek V4.1 Flash, Claude models and others about 570M Not published Not published
September total 2B Not published about 35 kWh, or "about 30kWh" in a later comment

AgentsView shows an 84.1% cache hit rate across the month, which is how 1.04B Flash tokens came to $68: about $0.065 per million, blended, by my arithmetic. Colas says he is moving to DeepSeek V4.1 Flash in October, which "feels slightly better from a few days of use".

Wagtail's benchmark passed GLM 5.3 Flash on 11 of 20 tasks #

Wagtail's work-in-progress benchmark of 20 Wagtail tasks is the number the vendor tables lack: GLM 5.3 Flash passed 11 at a median $0.16 and 26.2 Wh a task, while DeepSeek V4.1 Flash, Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6 Sol each passed 19. The table is an image in the post, and its harness and effort level are not given.

Model Passed Median cost per task Energy per task
DeepSeek V4.1 Flash 19 of 20 $0.09 14.9 Wh
Claude Sonnet 5.5 19 of 20 $0.14 Not measured
GPT-6 Sol 19 of 20 $0.15 Not measured
Claude Opus 5.5 19 of 20 $0.55 Not measured
GLM 5.3 14 of 20 $0.87 56.9 Wh
GLM 5.3 Flash 11 of 20 $0.16 26.2 Wh

Divided by the pass rate, Flash costs about $0.29 per passing task and Sonnet 5.5 about $0.15, by my arithmetic. Twenty tasks is a small sample. Effort level moves the cost too: HN user lhl, on about 300 tasks, found medium effort the best dollars per pass, with max scoring about 5% higher at more than twice the cost.

GLM 5.3 Flash kept the default slot, and commenters made it the executor #

For its author, GLM 5.3 Flash is neither planner nor executor: it is the default for routine work, with GLM 5.3 kept for the hardest 5%, and he added that he has "a hard time justifying GLM 5.3 these days". Commenters who split the roles mostly put Flash in the executor slot behind a stronger planner.

From the thread, 180 comments by 7 October:

  • RussianCow: GLM 5.3 writes "a detailed plan with little to no ambiguity", and Flash "executes it just fine for a fraction of the price".
  • surgical_fire: Flash is the implementation model after GLM 5.3 plans, in separate sessions that pass markdown files.
  • throw930rmdkdk: Flash codes, with one model planning and another reviewing.
  • esafak , on Z.ai's plan: the reverse, Flash for planning and review "because it's too slow for execution".

Colas's October plan points the same way: "Orchestrator vs. scout vs. implementer vs. reviewer agents." The model routing post has more September numbers on who sends which step to which model.

On Stackness, as of 7 October 2026, 7 real profiles list Claude Code and 2 list OpenCode, both alongside Claude Code, and none lists a GLM model (data sources). The numbers are small; the language models developers list will show when GLM appears.

Where Claude Code breaks on GLM 5.3 Flash #

Claude Code is built for Claude, and Anthropic's gateway docs say it does not support routing to non-Claude models. Pointed at GLM 5.3 Flash, Claude Code hits four known problems: requests that switch thinking off, a 200K context assumption, web search that only runs on Anthropic's backend, and tool calls that come back empty.

Symptom Cause Fix
HTTP 400, code 1210, on background requests Flash cannot switch thinking off, and Claude Code turns it off for some requests ( cc-switch #6737 , late August) Z.ai's docs now say such requests run at low effort; Claude Code also drops thinking for the rest of a conversation after a rejection
Compaction at 200K on a 1M model Claude Code assumes 200K for model IDs it does not know [1m] on the model ID andCLAUDE_CODE_AUTO_COMPACT_WINDOW
Web search fails WebSearch runs on Anthropic's backend Z.ai's web search MCP, 1.2 plan credits a call
Empty tool calls, then repeated 400s Reported through a gateway with Pi ( llmgateway #4366 ) and inKilo Code (#14475 ) Open, no fix published
The model stops thinking after about 50 steps and loops Ollama Cloud did not send earlier thinking back ( masc PR ) Z.ai recommends clear_thinking: false
Screenshots fail in the Sonnet and Opus slots Z.ai's manual config maps them to GLM 5.3, which is text-only Map those slots to Flash
A "Claude" run that ran on GLM A silent step-down to GLM when no Claude seat was free ( ticfac #248 ) Log the model behind every run

Remote Control and the Advisor tool need Anthropic's API. The DeepSeek Harness post covers what else stays behind when you switch harness or model.

Run it yourself #

Claude Code runs GLM 5.3 Flash with four settings: Z.ai's Anthropic-compatible base URL, a Z.ai API key, the Flash model ID with [1m] in all three model slots, and a 1M-token compaction window, plus the longer timeout Z.ai sets. Z.ai documents the endpoint, key and timeout on its Claude Code page and the model IDs and window on its latest-model page. Add them to ~/.claude/settings.json rather than replacing the file:

{
  "env": {
    "ANTHROPIC_AUTH_TOKEN": "your_zai_api_key",
    "ANTHROPIC_BASE_URL": "https://api.z.ai/api/anthropic",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "glm-5.3-flash[1m]",
    "ANTHROPIC_DEFAULT_SONNET_MODEL": "glm-5.3-flash[1m]",
    "ANTHROPIC_DEFAULT_OPUS_MODEL": "glm-5.3-flash[1m]",
    "CLAUDE_CODE_AUTO_COMPACT_WINDOW": "1000000",
    "API_TIMEOUT_MS": "3000000"
  }
}

/status should then show glm-5.3-flash[1m] as the model.

To let GLM 5.3 plan and Flash execute, set ANTHROPIC_DEFAULT_OPUS_MODEL to glm-5.3[1m] and switch to /model opusplan, which Anthropic's model docs describe as the Opus model in plan mode and the Sonnet model otherwise. I worked this out from the docs and have not run it.

Keeping Claude as the planner takes two sessions with a handoff file, the setup surgical_fire describes, or a gateway that routes by model name, such as claude-code-router, which Anthropic does not support. MoAI-ADK retired its Claude-leader-with-GLM-teammates mode and now runs one lane per model, moai cc -l and moai glm -l.

OpenCode and Ollama, from Z.ai's OpenCode page and Ollama's library page:

curl -fsSL https://opencode.ai/install | bash
opencode auth login          # pick Z.AI Coding Plan, then /models in a session
ollama launch claude --model glm-5.3-flash:cloud

Ollama has only a cloud tag for Flash. Colibri serves it locally by streaming experts from disk, at about 20 seconds per token in its README. Wagtail's TensorX configs for omp and OpenCode are in its wiki, with one tip: on a "quota reached" error, set max output tokens to 16384 or 32768.

Tools in this post #

Use any of these tools? #

Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes.

Show my stack

── more in #ai-tools 4 stories · sorted by recency
── more on @thibaud colas 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/glm-5-3-flash-as-a-c…] indexed:0 read:10min 2026-10-07 · —