cd /news/ai-tools/one-month-coding-with-glm-5-3-flash · home › topics › ai-tools › article
[ARTICLE · art-143958] src=wagtail.org ↗ pub= topic=ai-tools verified=true sentiment=· neutral

One month coding with GLM 5.3 Flash

A one-month experiment to run all of September on the open model GLM 5.3 Flash ended with only 50% of the 2 billion tokens generated going to the target model, according to the developer's AgentsView usage data. GLM 5.3 Flash usage cost $68 and consumed about 4kWh of energy (365 grams of carbon emissions), while a vibe-coded Wagtail MCP server prototype burned 450 million tokens, $150 and 5kWh almost overnight after the wrong model was selected, and infrastructure capacity limits forced switches to DeepSeek V4.1 Flash and Qwen 3.8 Flash. Total energy use reached about 35 kWh instead of the planned 10 kWh, and the team's October plan calls for constant measurement of tokens, energy and spend, budgeting for experimentation, and orchestrator/scout/implementer/reviewer multi-agent techniques.

read4 min views3 publishedOct 2, 2026
One month coding with GLM 5.3 Flash
Image: Wagtail (auto-discovered)

2B tokens later, 🤖 task failed successfully #

Setting a challenge to spend the whole of September on only one efficient open model felt like a great idea at the time. Turns out not so much in practice. 2B tokens later, here’s how it went.

Where tokens went this month #

Here’s the tokens distribution according to AgentsView, one of our Agentic engineering recommendations to keep tabs on AI usage:

Zooming in on the models split specifically:

The goal was to spend the whole month on GLM 5.3 Flash pictured in teal. Here’s what went well:

  • Successfully spent the first half of the month on just that model.
  • That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).

The second half of the month didn’t go so well, with 1B tokens going to other models.

Unexpected hurdles #

The cost of vibe coding

We’re pretty transparent that our experimental Wagtail MCP server is a vibe-coded prototype. Vibe coding isn’t quite what we normally aspire to, but for a prototype it’s spot on. Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:

Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.

Infrastructure woes

Another unexpected hurdle was infrastructure availability issues. We’ve written extensively about comparing inference providers. Our choices work really most of the times, but it turns out they’re very popular, and do not have the same capacity as the big labs who hoard all the GPUs. We noted degradation with the performance of GLM 5.3 Flash in particular, most likely because of it being so high up the Pareto frontier of relevant models for our work.

This meant having to switch to other similar models (DeepSeek V4.1 Flash, Qwen 3.8 Flash). Which is very simple to do, but nonetheless unexpected!

The cost of experimentation and R&D

Last but not least, beyond using one model for day-to-day engineering, it felt essential to keep experimenting with a wide range of models, keeping up with what providers are releasing. This is particularly essential as we start to benchmark models’ performance on Wagtail tasks, where we need data across a wide range of models. Sneak peek of our benchmark:

It’s much easier to guide people towards leaner options with this kind of concrete data. And for us to make those options even more viable with agent skills, or our new CLI prototype, which is intended to work well with agents.

Takeways and what to do next #

So technically this challenge was a failure. Only 50% usage on the target model, 1B out of 2B tokens. About 35 kWh of energy use instead of 10. But we did learn a lot, which is crucial for the current moment. Reflecting on this for October, here’s what will make it work:

  1. Constant measurement. Looking not just at tokens but also energy use and spend, and ideally how well this all leads to concrete positive outcomes.
  2. Budgeting for experimentation, not just day-to-day tasks. Making more concerted decisions about which prototypes are worth building, and how.
  3. Better prompt selection and multi-agent techniques. Orchestrator vs. scout vs. implementer vs. reviewer agents.
  4. Keep pushing for more efficient techniques and models. The Jev-style decision diffusion models look very promising if they can run so efficiently.
For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models. A viable target is probably that the *majority* of AI inference work should be done with such efficient models, measured in cost or energy use rather than meaningless tokens. That’s the goal for October! You should try it too, you’ll learn a lot in the process.

And come say hi at [Wagtail Space 2026](https://wagtail.org/wagtail-space-2026/) in November to hear how that all pans out!
── more in #ai-tools 4 stories · sorted by recency
── more on @glm 5.3 flash 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/one-month-coding-wit…] indexed:0 read:4min 2026-10-02 · —