Do read Z.ai’s recent blog. Z.ai says an agent running on GLM-5.3 built the inference service that now serves GLM-5.3-Flash, on a cluster of more than 100,000 Chinese-made accelerators, reaching 3.22x baseline throughput in thirteen days. This is a quick review of the architecture, and what it does to the cost, price and buildout arithmetic in Parts I to VI of my series, which starts with The House of Cards Didn’t Fall….
I like their sharing. And that of DeepSeek AI’s v4.1 Flash Technical report. If anything, this not only shows how the Chinese Labs are “doing more, with less (especially hardware), but will challenge your thoughts on how you may see the trillion-dollar AI capex-heavy-build-out supporting the Frontier Labs, and Nvidia’s role in it. And then on inferencing, how wide will these margins stay!
Start here, then read the DeepSeek piece(s), to be followed by this article:
Why this is worth reading
• A second lever on the same arithmetic. SemiAnalysis’s inference work models rising profit per gigawatt from better silicon. This run moves the identical number with scheduling and kernels, on chips Nvidia did not make.
• It is cheap, fast and portable. Thirteen days and a team of engineers, against quarters of lead time and billions of financed capex.
• It lands on the substitutable tier. Open weights, cheap verification, aggregator distribution — the exact quadrant where my Part III work says margin does not hold.
• It gives the bear case a mechanism. “Prices will fall” is a forecast. A published, reproducible way to triple throughput in a fortnight is a reason.
What’s actually stated in their blog:
- All production traffic for GLM-5.3-Flash runs on a cluster of 100,000+ domestic accelerators. Z.ai never says who owns it, who bought the chips, or on what terms.
- No chipmaker is named. LatePost reporting names Huawei, Moore Threads and Hygon as possible suppliers, with Z.ai declining to comment. QuartzIDC Atlas
- Weights shipped MIT on HuggingFace. Cambricon and Moore Threads both announced Day-0 support for serving the model. Zero Hedge
So there are two different “other providers” at work
- Compute. Almost certainly not Z.ai’s own balance sheet at 100k chips — but that’s inference on my part. I accept there is obvious hardware constraints that apply. The standard Chinese structure is vendor-supplied silicon in a state-backed or vendor-affiliated cloud, rented by the lab, often with compute vouchers or concessional terms. A flagship model proving Ascend or Cambricon at production scale is worth more to the vendor as a reference deployment than the margin on the chips.
- Serving capacity. Open weights mean third-party hosts add capacity at zero capex to Z.ai and zero revenue share back. Z.ai captures distribution and price leadership; whoever runs cheapest captures the margin. That’s the scaling channel, and it’s free to them.
Why this matters for my thesis
The “per-token cost comparable to Nvidia GPUs” claim is engineering efficiency. It is not an arm’s-length cost of capital. If someone else carries the chips, the depreciation and the power contract, then the comparison to a hyperscaler’s fully-loaded inference cost is not like-for-like — which is exactly the objection I’d raise against a neocloud quoting $/hr off a vendor-backstopped facility. The vendor-financed reference cluster is the Chinese edition of the Nvidia backstop from Part I: the same circularity, a different flag.
⚙️ What they built #
• Encode–Prefill–Decode, disaggregated. Three stages run on separate pools, so a long prompt being read never stalls tokens being written for another user.
• Everything at low precision. W8A8 weights and activations; the KV cache quantised per layer across INT8, FP8 and BF16.
• Parallelism did the heavy lifting. Layer Split and context parallel together moved the system from 1.41x to 2.49x in two days — more than half the total gain in a tenth of the calendar.
• Kernel work filled the rest. Chunked MQA, a prefill dequant kernel, fused activation plus quantisation, ReplaySSM, intra-node tensor parallelism for linear attention and the LM head.
• On hostile hardware. Thin on-chip memory capacity and bandwidth, a new model architecture, a 1M-token window, multimodal traffic, incomplete kernel support and behaviour the engineers had to guess at.
The claim, in their words
Z.ai frames the run as “The model optimizes the system; the system runs the model” — while stating plainly that it has not reached recursive self-improvement, and that objectives, boundaries and risk stay with humans.
In plain terms — Nothing about the hardware changed. They rearranged how work moves through it and got three times the output in under a fortnight.
Figure 1 · The serving path, the agent loop and the thirteen-day result. Source: Z.ai, 17 September 2026.
🔁 The loop is the product #
The interesting engineering (no pun intended), is not the stack, so much. It is the admission that an agent given only end-to-end metrics is blind: throughput fell, and nothing in the repository says which of five layers caused it. So they built the answer into the environment.
• Dense feedback = local, cheap, verifiable. Tie every signal to one kernel, one request, one thread. Let a hypothesis be tested by a microbenchmark rather than a cluster deploy. Settle causes with controlled experiments, never correlation.
• Three named wins. A TF32 precision bug in the KDA kernel under context parallelism, fixed and merged upstream as Flash Linear Attention PR #1180. A Python GIL in DeepEP blocking KV-transfer overlap: a 20% penalty cut to under 1%. A decode kernel made 1.71x faster by merging tiles and giving up some parallelism.
• Numeric acceptance criteria. An engineer-set rule that prefill plus KV transfer must land within 5% of prefill alone is what turned a vague slowdown into a bounded search.
• The compounding asset. Validated techniques are distilled into reusable optimisation skeletons, so each deployment lowers the engineering cost of the next.
Why this matters to my harness work
This is the same claim my harness experiments keep landing on, now at cluster scale: the ceiling is set by the feedback environment rather than the model. Z.ai just priced that claim in production hardware.
In plain terms — The win came from building a place where the agent could check its own guesses cheaply. That is copyable by anyone, and it gets cheaper every time it is used.
📈 Read it against the SemiAnalysis inference case #
SemiAnalysis published its Vera Rubin agentic-inference review three days before the Z.ai post. Read together they describe the same denominator from opposite ends.
• The bull case, fairly stated. Rubin NVL72 delivers up to 67x the throughput per dollar of TCO against GB300 at a 170 TPS target, and 1.4x to 3x at the 60 to 100 TPS band where most providers actually serve. SemiAnalysis models roughly 39% more annual revenue and 42% more profit per gigawatt than the best GB300 configuration. The earlier value-capture work has Anthropic’s ARR at $44bn with inference gross margin moving from 38% to over 70%.
• Why that supports the buildout. If tokens per megawatt keep climbing and realised price holds, every financed gigawatt earns more each year than the one before it. That is the assumption under the capex schedules, the SPV structures and the vendor backstops I walked through in Parts I and II.
• Where Z.ai cuts across it. The same numerator moved 3.22x in thirteen days with no new silicon, on domestic accelerators, and the fix was published. An efficiency gain that everyone can obtain reaches the customer as a lower price. It becomes margin only for whoever holds it exclusively.
• Both can be true. Rubin can be the best token factory ever built and still sell into a market where the marginal seller keeps cutting. Profit per gigawatt is tokens per megawatt times realised price times utilisation, and the bull case is strongest on the first term.
• Watch the benchmark politics too. AgentX is an open benchmark run by a firm that also publishes indices and models the hardware it measures, and the headline numbers are early-silicon results on coding-agent traffic. Governance neutrality is an open question, and the 67x is a TCO ratio at one interactivity point.
Figure 3 · Two routes to the same denominator. Sources: SemiAnalysis, 14 Sep 2026; Z.ai, 17 Sep 2026.
In plain terms — The earnings case for the buildout rests on tokens per megawatt going up. Z.ai just showed that number can also be moved by software, by anyone, on cheaper hardware — and when it moves that way, it arrives as a price cut instead of a margin.
📉 How it lands on the margin thesis #
Four of my six published parts move on this evidence.
In plain terms — The cheap tier just got cheaper to produce, on hardware nobody in the West/Frontier Labs is counting or considering (who knows, I may be wrong on this), by a method that anyone can copy.
Figure 2 · Cost, price, residual value and buildout — four legs of the underwriting question, and where the GLM evidence pushes each.
🧯 What I discount #
• Self-reported, unreproduced. Throughput, cluster size and cost parity come from the vendor. Only the Flash Linear Attention PR is independently checkable.
• The baseline flatters the multiple. A first successful run on unfamiliar silicon with incomplete kernel support is a low bar to triple. 3.22x measures the distance travelled, not the altitude reached.
• “Comparable to Nvidia” is unaudited. No power draw, no depreciation basis, no utilisation series. Per-token cost parity without those three numbers is a marketing sentence.
• One round only. The recursive claim needs a second pass, where the system optimises what it just changed. Z.ai has not published one.
• It is also an ad. The post shipped alongside the coding plan the model sells into.
In plain terms — Treat 3.22x as a direction of travel with the vendor’s thumb on the scale. The direction is still the part that matters.
🎯 Pre-registered #
• G1 — No independent third-party reproduction of a 3x-class end-to-end gain on domestic Chinese accelerators is published.
• G2 — Z.ai publishes no second-round optimisation curve (agent improving its own prior output) before 31 March 2027.
• G3 — GLM-5.3-Flash blended list price falls below $0.40 per million tokens by 30 June 2027, from $0.65 today.
• G4 — Frontier-lab blended realised revenue per million tokens falls year-on-year through FY2027, even while reported inference gross margin holds above 60%. Volume, mix and hardware carry the margin while price keeps sliding.
• G5 · expected failure — At least one US hyperscaler publicly attributes a 2x-or-better inference throughput gain to an internal agent loop by Q2 2027. I expect this one to break, because the incentive to disclose runs the wrong way.
Bottom line #
• The cost curve now has a leg that nobody finances, nobody depreciates and nobody counts in a gigawatt model. Thirteen days of agent-assisted engineering did the work of a hardware generation.
• That leg is available to the sellers with the least pricing power, on the hardware with the worst economics. It compresses the spread from below.
• The floors still hold. Memory remains an oligopoly, delivery certainty still carries a premium, and the frontier is still protected wherever substitution is hard.
• Worth reading, in one line: the buildout is underwritten on tokens per megawatt rising, and this is one of the first production-scale demonstration that the same number can be raised on the software side by a competitor with worse chips and no capex. DeepSeek is doing similar in different ways with optimizations on software
• For the buildout question, the uncomfortable version is this: if throughput per chip is a software variable, then so is the revenue per megawatt that the debt was underwritten against.
References
1. Z.ai, “Toward Recursive Self-Improvement: How GLM Built Its Own Inference Infrastructure”, 17 Sep 2026. https://z.ai/blog/glm-built-its-inference-infrastructure
2. explainx.ai, “GLM-5.3 Built Its Own Inference Stack. The Real Lesson Is Dense Feedback”, 17 Sep 2026 (day-by-day curve, case detail). https://www.explainx.ai/blog/glm-5-3-infra-agent-dense-feedback-inference-2026
3. Unite.AI, “Z.ai Details GLM-5.3-Flash Inference Build on 100,000 Chinese Chips”, 17 Sep 2026. https://www.unite.ai/z-ai-details-glm-5-3-flash-inference-build-on-100-000-chinese-chips/
4. ai-tldr.dev, “Z.ai says GLM built its own inference stack”, 17 Sep 2026 (named fixes, upstream PR). https://ai-tldr.dev/releases/zai-glm-infra-agent/
5. Trending Topics, “Forget AGI, Here Comes RSI”, 17 Sep 2026 (verifiability caveats). https://www.trendingtopics.eu/forget-agi-here-comes-rsi-z-ai-says-its-glm-model-built-its-own-inference-infra/
6. SemiAnalysis, “Vera Rubin NVL72 Agentic Inference: 67x better Performance per Dollar”, 14 Sep 2026.
7. SemiAnalysis, “AI Value Capture — The Shift To Model Labs” (Anthropic margin path, compute constraint).
8. NVIDIA, “Up to 30x More Work Per Watt: Vera Rubin NVL72”, updated 15 Sep 2026 (AgentX results). https://blogs.nvidia.com/blog/vera-rubin-nvl72-efficiency-ai-agents/
9. TechTimes, “AgentX Benchmark: Vera Rubin NVL72 Achieves 30x Efficiency Gain”, 16 Sep 2026 (benchmark governance caveat). https://www.techtimes.com/articles/327595/20260916/agentx-benchmark-vera-rubin-nvl72-achieves-30x-efficiency-gain-over-gb300-ai-agents.htm
10. OpenRouter, GLM-5.3 and GLM-5.3-Flash pricing comparison (accessed 19 Sep 2026). https://openrouter.ai/compare/z-ai/glm-5.3/z-ai/glm-5.3-flash
11. Z.ai API pricing documentation. https://docs.z.ai/guides/overview/pricing
12. computeprices.com, OpenRouter GLM-5.3-Flash rate tracking. https://computeprices.com/providers/openrouter/models/glm-5-3-flash