# China's GLM-5.3-Flash Just Beat Claude Opus 4.8 on Agent Benchmarks

> Source: <https://startupfortune.com/chinas-glm-53-flash-just-beat-claude-opus-48-on-agent-benchmarks/>
> Published: 2026-09-04 06:44:13+00:00

*Zhipu AI's new open-weight model just beat Claude Opus 4.8 on two major agent benchmarks, and it costs roughly a fortieth of the price. The full weights are free to download.*

For six days in August, developers on OpenRouter kept testing an unbranded endpoint called "Ox Alpha." They swapped notes about how well it chained together coding tasks, guessing at who had built it. Nobody knew. On August 26, 2026, Zhipu AI, the Beijing lab known as Z.ai, ended the guessing game. Ox Alpha was GLM-5.3-Flash, and the company put the full weights on Hugging Face under an MIT license that same day.

The reveal came with numbers attached, and the numbers are the story. On a benchmark for multi-step automation, GLM-5.3-Flash scored 48.8 against Claude Opus 4.8's 41.0. On a software-engineering agent benchmark tracked by LLM Stats, it hit 63.4 versus Opus's 58.0. Anthropic's flagship still edges ahead on Terminal-Bench 2.1, 85.0 to 84.3. That's close enough that it barely counts as a win.

Then there's the price. Zhipu lists GLM-5.3-Flash at $0.15 per million input tokens and $0.50 per million output, and it's running a launch promotion through September 9 that cuts that to $0.075 and $0.25, matching what OpenRouter already charges. Claude Opus 4.8 runs several dollars per million tokens. On a blended basis, LLM Stats puts GLM-5.3-Flash at roughly 42 times cheaper than Opus. You read that right. Forty-two times.

GLM-5.3-Flash is a mixture-of-experts model with 320 billion total parameters but only 18 billion active on any given pass, paired with a hybrid sparse and linear attention setup that keeps its 1,048,576-token context window from becoming prohibitively expensive to run. Zhipu trained it on a 30-trillion-token multimodal corpus, and it's the first model in the GLM-5 line that natively handles text, image and video without bolted-on adapters. Artificial Analysis scores it 57 on its Intelligence Index, tying Claude Opus 4.8 and placing it fourth out of 111 models the firm tracks, against a median of just 29 for open-weight models its size.

[Zhipu's GLM-5.3-Flash Undercuts Claude and GPT on Price, Not on Hardware](https://startupfortune.com/zhipus-glm-53-flash-undercuts-claude-and-gpt-on-price-not-on-hardware/)

Zhipu AI revealed that its viral anonymous model Ox Alpha was actually GLM-5.3-Flash, an open-weight, MIT-licensed system that undercuts Claude and GPT on API price while matching them on coding benchmarks. Running it locally, though, still requires a serious multi-GPU setup, not a single consumer card. - [GLM-5.3-Flash model pricing compared to Claude Opus](https://startupfortune.com/zhipus-glm-53-flash-undercuts-claude-and-gpt-on-price-not-on-hardware/) - [open weight model performance on coding benchmarks](https://startupfortune.com/zhipus-glm-53-flash-undercuts-claude-and-gpt-on-price-not-on-hardware/)

None of that would matter much if you couldn't run it. You can. The weights are MIT-licensed, so anyone with the hardware can self-host GLM-5.3-Flash for the cost of electricity, no per-token fee to Z.ai at all. On Z.ai's own Code Bench at high effort, the model scored 31.4% against Opus's 29.5%, and it did it using around 50,000 output tokens per task instead of Opus's roughly 120,000. That's not just a cheaper model. It's a model that finishes the job faster, too.

## Why US labs should be nervous

Frankly, this is the pattern the industry has been bracing for since DeepSeek first rattled markets in early 2025, except now it's showing up in agent workflows, which is where the real money sits. Startups running always-on coding agents or automation pipelines pay by the token, and every token routed through Opus instead of a model that's within a few points of it, at a fortieth the cost, is margin left on the table. A founder running a customer-support agent that fires off a few hundred thousand tokens a day doesn't need a lecture on open weights versus closed ones. They need a bill they can afford, and GLM-5.3-Flash just handed them one.

Anthropic hasn't cut Opus 4.8's pricing in response, at least not yet. It doesn't have to immediately. Enterprises with compliance requirements, existing contracts and a preference for a vendor they can call a support line about will keep paying the premium for a while. But the gap between good enough and nearly free versus best and expensive keeps narrowing every few months, and GLM-5.3-Flash ranks below the flagship in Zhipu's own lineup. Zhipu still has the full GLM-5.3 sitting above it in the lineup.

Six days ago, most developers testing Ox Alpha didn't know they were talking to a Chinese open-weight model. Now everyone does, and the number that will stick with people is 42.

**Also read:** [SoftBank pays double the market rate to raise $6.3 billion for its AI bets](https://startupfortune.com/softbank-pays-double-the-market-rate-to-raise-63-billion-for-its-ai-bets/) • [Novo Nordisk's Foundation Is Building a Quantum Chip Factory in Copenhagen](https://startupfortune.com/novo-nordisks-foundation-is-building-a-quantum-chip-factory-in-copenhagen/) • [Tesla launches its steering-wheel-free Cybercab robotaxi in Austin](https://startupfortune.com/tesla-launches-its-steering-wheel-free-cybercab-robotaxi-in-austin/)

## Founder discussion

[Open in the community →](/community/)

Almost there. Sign in and your reply posts straight away.
