GLM 5.3 Flash as a coding agent: a month of daily use, $68, and the slot it kept Thibaud Colas of the Wagtail core team spent September 2026 coding daily on Z.ai's MIT-licensed GLM 5.3 Flash, consuming about 1B tokens for $68 and roughly 4 kWh of GPU energy, according to numbers he published on 2 October. Wagtail now recommends GLM 5.3 Flash for 95% of tasks and GLM 5.3 for the hardest 5%, and the model passed 11 of 20 tasks on Wagtail's own benchmark at a median cost of $0.16 per task. GLM 5.3 Flash, released 26 August 2026, has 320B total parameters, 18B active, a 1M-token context, and is priced at $0.15 per million input tokens and $0.50 per million output tokens on the Z.ai API. GLM 5.3 Flash as a coding agent: a month of daily use, $68, and the slot it kept Thibaud Colas of the Wagtail https://stackness.dev/tools/wagtail core team spent September 2026 coding on GLM 5.3 Flash https://stackness.dev/tools/glm , Z.ai's MIT-licensed model with 320B parameters and 18B active, and published the numbers https://wagtail.org/blog/one-month-on-glm-53-flash/ on 2 October: about 1B tokens on Flash for $68 and about 4 kWh of GPU energy. Flash kept the default slot rather than a planner or executor role. Wagtail now recommends https://wagtail.org/ai/agentic-engineering-recommendations/ it for 95% of tasks and GLM 5.3 for the hardest 5%. On Wagtail's own 20-task benchmark it passed 11, at a median $0.16 a task. The run is his, not mine, and it did not use Claude Code. He drove Flash from omp https://stackness.dev/tools/omp , billed per token through two European providers. The Claude Code https://stackness.dev/tools/claude-code and OpenCode https://stackness.dev/tools/opencode setup at the end comes from Z.ai's and Anthropic's docs. What is GLM 5.3 Flash, and what does a task cost? GLM 5.3 Flash is Z.ai's coding and vision model, released on 26 August 2026: 320B total parameters, 18B active, MIT licence, a 1M-token context and thinking that cannot be switched off. Z.ai and OpenRouter list it at $0.15 per million input tokens, $0.03 cached and $0.50 output. On Wagtail's benchmark a task cost a median $0.16; Z.ai quotes $0.045 on Artificial Analysis's index tasks, a different task set. | Per million tokens, read 7 October 2026 | Input | Output | |---|---|---| | GLM 5.3 Flash, Z.ai API https://docs.z.ai/guides/overview/pricing | $0.15 | $0.50 | | GLM 5.3 Flash, TensorX the author's provider | $0.20 | $0.50 | | GLM 5.3, Z.ai API | $1.40 | $4.40 | | Claude Sonnet https://stackness.dev/tools/claude-sonnet 5.5 | $2 | $10 | | Claude Opus https://stackness.dev/tools/claude-opus 5.5 | $4 | $20 | Z.ai's GLM Coding Plan https://z.ai/subscribe costs $18, $80 or $168 a month, and Flash spends a third of the plan credits GLM 5.3 does per token, which is where the "3x quota" in Z.ai's marketing comes from. The plan works only inside listed tools: Claude Code, OpenCode, Pi https://stackness.dev/tools/pi , Codex https://stackness.dev/tools/codex-cli and about a dozen others, not omp. Z.ai's own launch table https://z.ai/blog/glm-5.3-flash puts Flash at 84.3 on Terminal-Bench 2.1, run in Claude Code, against 85.0 for Opus 4.8. Z.ai publishes no SWE-bench Verified score for it. The stack: omp, two EU providers and per-token billing Thibaud Colas ran GLM 5.3 Flash in the omp coding agent, on two inference providers serving from European data centres, and paid per token with no subscription. The month was a challenge he set in a 3 September post https://wagtail.org/blog/open-models-only-adoption-challenge/ : one efficient open-weight model for all of September, and "ditch your AI coding subscription". - Harness: omp oh-my-pi , confirmed by the author on Hacker News https://news.ycombinator.com/item?id=49938853 . - Model: GLM 5.3 Flash, with GLM 5.3 for harder work, and DeepSeek https://stackness.dev/tools/deepseek V4.1 Flash and Qwen https://stackness.dev/tools/qwen 3.8 Flash as fallbacks. - Inference: TensorX https://stackness.dev/tools/tensorx and Neuralwatt https://stackness.dev/tools/neuralwatt , usage-based billing only. - Metering: AgentsView https://stackness.dev/tools/agentsview for tokens and cost, whose costs Wagtail calls "indicative only". Energy and carbon come from Neuralwatt and count GPU energy only. Neuralwatt supplied the energy figures, and its CTO said in the Hacker News thread that his company was likely the provider used. The post does not say whether the $68 is a provider bill or an AgentsView estimate. The first half of September ran on Flash alone GLM 5.3 Flash carried all of Colas's work for the first two weeks: "Wagtail itself, sites built with it. UI tasks, AI R&D, also a bit of documentation writing. Lots of evals." He picked it for three reasons: a 1M context for long tasks, vision for building from screenshots, and many providers competing on price. His verdict: "GLM 5.3 Flash itself was excellent " For day-to-day work, he wrote, "it's totally viable to focus on one or two flash-tier cheap models." One overnight build on GLM 5.3 took 150% of the budget The month's largest bill came from a model mix-up, not from Flash. Colas vibe-coded a Wagtail MCP server prototype overnight in omp and left it on the full GLM 5.3, which lists at about nine times Flash's input price. That session used 450M tokens, $150 and about 5 kWh. He explained it on Hacker News https://news.ycombinator.com/item?id=49937588 : "I didn't realize that one session was a quarter of the month's spend and 150% of the budget". He estimates Flash would have done similar work "for most likely 5x less cost". Provider capacity moved the second half to other models In the second half of September, Colas saw GLM 5.3 Flash's performance degrade at the European providers he used, which he put down to demand: popular open models "do not have the same capacity as the big labs who hoard all the GPUs." He moved work to DeepSeek V4.1 Flash and Qwen 3.8 Flash, and R&D needed other models, so Flash ended the month at 1B of 2B tokens. What September cost, by model, from the post and its AgentsView screenshot https://media.wagtail.org/images/agentsview septembermodels split.width-800.png : | Line | Tokens | Cost | Energy | |---|---|---|---| | GLM 5.3 Flash, all month | 1.04B | $68 | about 4 kWh | | GLM 5.3, the overnight MCP build | 451M | $150 | about 5 kWh | | GPT-6 Astra, DeepSeek V4.1 Flash, Claude models and others | about 570M | Not published | Not published | | September total | 2B | Not published | about 35 kWh, or "about 30kWh" in a later comment | AgentsView shows an 84.1% cache hit rate across the month, which is how 1.04B Flash tokens came to $68: about $0.065 per million, blended, by my arithmetic. Colas says he is moving to DeepSeek V4.1 Flash in October, which "feels slightly better from a few days of use". Wagtail's benchmark passed GLM 5.3 Flash on 11 of 20 tasks Wagtail's work-in-progress benchmark of 20 Wagtail tasks is the number the vendor tables lack: GLM 5.3 Flash passed 11 at a median $0.16 and 26.2 Wh a task, while DeepSeek V4.1 Flash, Claude Sonnet 5.5, Claude Opus 5.5 and GPT-6 Sol each passed 19. The table is an image in the post, and its harness and effort level are not given. | Model | Passed | Median cost per task | Energy per task | |---|---|---|---| | DeepSeek V4.1 Flash | 19 of 20 | $0.09 | 14.9 Wh | | Claude Sonnet 5.5 | 19 of 20 | $0.14 | Not measured | | GPT-6 Sol | 19 of 20 | $0.15 | Not measured | | Claude Opus 5.5 | 19 of 20 | $0.55 | Not measured | | GLM 5.3 | 14 of 20 | $0.87 | 56.9 Wh | | GLM 5.3 Flash | 11 of 20 | $0.16 | 26.2 Wh | Divided by the pass rate, Flash costs about $0.29 per passing task and Sonnet 5.5 about $0.15, by my arithmetic. Twenty tasks is a small sample. Effort level moves the cost too: HN user lhl https://news.ycombinator.com/item?id=49947203 , on about 300 tasks, found medium effort the best dollars per pass, with max scoring about 5% higher at more than twice the cost. GLM 5.3 Flash kept the default slot, and commenters made it the executor For its author, GLM 5.3 Flash is neither planner nor executor: it is the default for routine work, with GLM 5.3 kept for the hardest 5%, and he added that he has "a hard time justifying GLM 5.3 these days". Commenters who split the roles mostly put Flash in the executor slot behind a stronger planner. From the thread https://news.ycombinator.com/item?id=49934620 , 180 comments by 7 October: - RussianCow: GLM 5.3 writes "a detailed plan with little to no ambiguity", and Flash "executes it just fine for a fraction of the price". - surgical fire: Flash is the implementation model after GLM 5.3 plans, in separate sessions that pass markdown files. - throw930rmdkdk: Flash codes, with one model planning and another reviewing. - esafak , on Z.ai's plan: the reverse, Flash for planning and review "because it's too slow for execution". Colas's October plan points the same way: "Orchestrator vs. scout vs. implementer vs. reviewer agents." The model routing post https://stackness.dev/blog/model-routing-for-coding-agents-the-plan-and-price-numbers-developers-measured-in-september-2026 has more September numbers on who sends which step to which model. On Stackness, as of 7 October 2026, 7 real profiles list Claude Code and 2 list OpenCode, both alongside Claude Code, and none lists a GLM model data sources https://stackness.dev/about/data-sources . The numbers are small; the language models developers list https://stackness.dev/categories/llms will show when GLM appears. Where Claude Code breaks on GLM 5.3 Flash Claude Code is built for Claude, and Anthropic's gateway docs https://code.claude.com/docs/en/llm-gateway say it does not support routing to non-Claude models. Pointed at GLM 5.3 Flash, Claude Code hits four known problems: requests that switch thinking off, a 200K context assumption, web search that only runs on Anthropic's backend, and tool calls that come back empty. | Symptom | Cause | Fix | |---|---|---| | HTTP 400, code 1210, on background requests | Flash cannot switch thinking off, and Claude Code turns it off for some requests cc-switch 6737 https://github.com/farion1231/cc-switch/issues/6737 , late August | Z.ai's docs now say such requests run at low effort; Claude Code also drops thinking for the rest of a conversation after a rejection | | Compaction at 200K on a 1M model | Claude Code assumes 200K for model IDs it does not know | 1m on the model ID and CLAUDE CODE AUTO COMPACT WINDOW | | Web search fails | WebSearch runs on Anthropic's backend | Z.ai's web search MCP, 1.2 plan credits a call | | Empty tool calls, then repeated 400s | Reported through a gateway with Pi llmgateway 4366 https://github.com/theopenco/llmgateway/issues/4366 and in Kilo Code https://stackness.dev/tools/kilo-code 14475 https://github.com/Kilo-Org/kilocode/issues/14475 | Open, no fix published | | The model stops thinking after about 50 steps and loops | Ollama Cloud did not send earlier thinking back masc PR https://github.com/jeong-sik/masc/pull/41404 | Z.ai recommends clear thinking: false | | Screenshots fail in the Sonnet and Opus slots | Z.ai's manual config maps them to GLM 5.3, which is text-only | Map those slots to Flash | | A "Claude" run that ran on GLM | A silent step-down to GLM when no Claude seat was free ticfac 248 https://github.com/pengelbrecht/ticfac/pull/248 | Log the model behind every run | Remote Control and the Advisor tool need Anthropic's API. The DeepSeek Harness post https://stackness.dev/blog/deepseek-harness-vs-claude-code-which-harnesses-lock-you-to-one-model-and-what-breaks-when-you-switc covers what else stays behind when you switch harness or model. Run it yourself Claude Code runs GLM 5.3 Flash with four settings: Z.ai's Anthropic-compatible base URL, a Z.ai API key, the Flash model ID with 1m in all three model slots, and a 1M-token compaction window, plus the longer timeout Z.ai sets. Z.ai documents the endpoint, key and timeout on its Claude Code page https://docs.z.ai/devpack/tool/claude and the model IDs and window on its latest-model page https://docs.z.ai/devpack/latest-model . Add them to ~/.claude/settings.json rather than replacing the file: { "env": { "ANTHROPIC AUTH TOKEN": "your zai api key", "ANTHROPIC BASE URL": "https://api.z.ai/api/anthropic", "ANTHROPIC DEFAULT HAIKU MODEL": "glm-5.3-flash 1m ", "ANTHROPIC DEFAULT SONNET MODEL": "glm-5.3-flash 1m ", "ANTHROPIC DEFAULT OPUS MODEL": "glm-5.3-flash 1m ", "CLAUDE CODE AUTO COMPACT WINDOW": "1000000", "API TIMEOUT MS": "3000000" } } /status should then show glm-5.3-flash 1m as the model. To let GLM 5.3 plan and Flash execute, set ANTHROPIC DEFAULT OPUS MODEL to glm-5.3 1m and switch to /model opusplan , which Anthropic's model docs https://code.claude.com/docs/en/model-config describe as the Opus model in plan mode and the Sonnet model otherwise. I worked this out from the docs and have not run it. Keeping Claude as the planner takes two sessions with a handoff file, the setup surgical fire describes, or a gateway that routes by model name, such as claude-code-router https://stackness.dev/tools/claude-code-router , which Anthropic does not support. MoAI-ADK https://stackness.dev/tools/moai-adk retired its Claude-leader-with-GLM-teammates mode and now runs one lane per model, moai cc -l and moai glm -l . OpenCode and Ollama https://stackness.dev/tools/ollama , from Z.ai's OpenCode page https://docs.z.ai/devpack/tool/opencode and Ollama's library page https://ollama.com/library/glm-5.3-flash : curl -fsSL https://opencode.ai/install | bash opencode auth login pick Z.AI Coding Plan, then /models in a session ollama launch claude --model glm-5.3-flash:cloud Ollama has only a cloud tag for Flash. Colibri https://stackness.dev/tools/colibri serves it locally by streaming experts from disk, at about 20 seconds per token in its README. Wagtail's TensorX configs for omp and OpenCode are in its wiki https://github.com/wagtail/wagtail/wiki/Agentic-engineering-setup , with one tip: on a "quota reached" error, set max output tokens to 16384 or 32768. Tools in this post Use any of these tools? Put them on a Stackness profile, say how you use each one and see who pairs them the same way. It takes a couple of minutes. Show my stack https://stackness.dev/register