# GLM 5.3

> Source: <https://tokenstead.ai/models/glm-5-3>
> Published: 2026-08-14 19:45:15+00:00

# GLM 5.3

MoE workstation**743B total, ~40B active per token (MoE: 256 routed experts, 8 active + 1 shared).** Same base model as GLM 5.2 - every gain comes from scaled-up post-training, not architecture change. Built on the same stack: IndexShare (long-context), SAO (RL for long-horizon tasks), and slime (open-source async RL). Training environments simulate real professional work, some spanning days of engineer effort.

-
**Context:** 1M native, same as 5.2. -
**API change:** thinking is now mandatory with three effort levels (low / high / max) - a breaking change for apps that ran with thinking off.

**Coding (vendor-reported).** Terminal-Bench 3.0 28.3 (vs 4.6 on 5.2), DeepSWE v1.1 66.9 (vs 46.2), Agents’ Last Exam (CLI) 28.5 (vs 23.8). On Z.ai’s private Code Bench it scores 31.4% at ~50K output tokens, beating Claude Opus 4.8’s 29.5% at ~120K, but trails Claude Fable 5 (39.5% at max effort). It still trails GPT-5.6 Sol and Fable 5 on the harder public suites.

**Cybersecurity - the emergent capability.** Z.ai added vulnerability-discovery data expecting incremental single-bug gains; the model instead developed multi-step exploit-chain reasoning. CyberGym 84.5% (up from 77.2%, ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6%), ExploitBench 54.4% (up from 24.4%), ExploitGym 105 tasks in 2 hours / 130 in 6 (vs 29/39 for 5.2). In real-world testing it found 2,436 vulnerabilities across 269 open-source projects - 1,097 critical/high severity, the oldest dating to 1981. 53 CVEs are publicly disclosed; 2,383 sit under embargo via the Security Disclosure Ledger (cvd.z.ai).

**Availability.** Live now via the GLM Coding Plan and ZCode; per-token API access is rolling out in stages. Open weights are held back ~2 weeks for safety review (expected ~Aug 28, 2026) - the first GLM-series release delayed this way. Self-hosting needs ~1.5TB GPU memory at full precision or ~239GB quantized.

**Honest framing.** All benchmark figures are Z.ai’s own harness configurations. Independent verification awaits the open-weight release at the end of August.

- 743.0B
- 1000k
- mit
- Aug 2026

## What people are building with GLM 5.3

Real demos from X

## Scores

## Or run it in the cloud

Live per-provider pricing, throughput and uptime. Click a column to sort.

| Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value |
|---|---|---|---|---|---|---|---|---|
| Sub | - | - | - | - | - | - | $10.00/mo Coding Plan Lite | |
| Sub | - | - | - | - | - | - | $30.00/mo Coding Plan Pro | |
| Sub | - | - | - | - | - | - | $80.00/mo Coding Plan Max |

Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing.

[Detailed API pricing page + JSON endpoint →](/models/glm-5-3/pricing)

[See who runs Zhipu AI in production →](/adoption/zhipu)

## Inference cost over time

Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.
