GLM 5.3 Zhipu AI released GLM 5.3, a 743B-parameter mixture-of-experts model with ~40B active per token, achieving a Terminal-Bench 3.0 score of 28.3 (up from 4.6 on GLM 5.2) and a DeepSWE v1.1 score of 66.9 (up from 46.2), while introducing mandatory thinking with three effort levels and a 1M native context. The model also demonstrated emergent cybersecurity capabilities, scoring 84.5% on CyberGym and finding 2,436 vulnerabilities across 269 open-source projects, with open weights delayed about two weeks for safety review, expected around August 28, 2026. GLM 5.3 MoE workstation 743B total, ~40B active per token MoE: 256 routed experts, 8 active + 1 shared . Same base model as GLM 5.2 - every gain comes from scaled-up post-training, not architecture change. Built on the same stack: IndexShare long-context , SAO RL for long-horizon tasks , and slime open-source async RL . Training environments simulate real professional work, some spanning days of engineer effort. - Context: 1M native, same as 5.2. - API change: thinking is now mandatory with three effort levels low / high / max - a breaking change for apps that ran with thinking off. Coding vendor-reported . Terminal-Bench 3.0 28.3 vs 4.6 on 5.2 , DeepSWE v1.1 66.9 vs 46.2 , Agents’ Last Exam CLI 28.5 vs 23.8 . On Z.ai’s private Code Bench it scores 31.4% at ~50K output tokens, beating Claude Opus 4.8’s 29.5% at ~120K, but trails Claude Fable 5 39.5% at max effort . It still trails GPT-5.6 Sol and Fable 5 on the harder public suites. Cybersecurity - the emergent capability. Z.ai added vulnerability-discovery data expecting incremental single-bug gains; the model instead developed multi-step exploit-chain reasoning. CyberGym 84.5% up from 77.2%, ahead of Claude Mythos 5 at 83.8% and GPT-5.6 Sol at 83.6% , ExploitBench 54.4% up from 24.4% , ExploitGym 105 tasks in 2 hours / 130 in 6 vs 29/39 for 5.2 . In real-world testing it found 2,436 vulnerabilities across 269 open-source projects - 1,097 critical/high severity, the oldest dating to 1981. 53 CVEs are publicly disclosed; 2,383 sit under embargo via the Security Disclosure Ledger cvd.z.ai . Availability. Live now via the GLM Coding Plan and ZCode; per-token API access is rolling out in stages. Open weights are held back ~2 weeks for safety review expected ~Aug 28, 2026 - the first GLM-series release delayed this way. Self-hosting needs ~1.5TB GPU memory at full precision or ~239GB quantized. Honest framing. All benchmark figures are Z.ai’s own harness configurations. Independent verification awaits the open-weight release at the end of August. - 743.0B - 1000k - mit - Aug 2026 What people are building with GLM 5.3 Real demos from X Scores Or run it in the cloud Live per-provider pricing, throughput and uptime. Click a column to sort. | Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value | |---|---|---|---|---|---|---|---|---| | Sub | - | - | - | - | - | - | $10.00/mo Coding Plan Lite | | | Sub | - | - | - | - | - | - | $30.00/mo Coding Plan Pro | | | Sub | - | - | - | - | - | - | $80.00/mo Coding Plan Max | Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing. Detailed API pricing page + JSON endpoint → /models/glm-5-3/pricing See who runs Zhipu AI in production → /adoption/zhipu Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.