GLM-5.3-Flash Z.ai released GLM-5.3-Flash on 2026-08-26, a 320B-parameter mixture-of-experts model with 18B active parameters, the first natively multimodal model in the GLM-5 series and the first open-source frontier model to pair sparse attention with linear attention, cutting attention compute by 3.01x and KV cache by 4.44x versus GLM-5.3 while supporting a 1M context. The model, available under MIT license on HuggingFace, was revealed to be the final version of the anonymous 'ox-alpha' stealth model previewed on OpenRouter/OpenCode from Aug 20, whose entire traffic load was served on Chinese AI chips, with OpenCode reporting 42T tokens served in 6 days. Z.ai self-reported benchmarks show GLM-5.3-Flash leading open-source multimodal coding results, beating GLM-5.2 on every coding and agentic benchmark, including Terminal Bench 2.1 (84.3), DeepSWE v1.1 (63.4), and Toolathlon Verified (78.4). GLM-5.3-Flash MoE enthusiast 320B total MoE, only 18B active per token - the first natively multimodal model in the GLM-5 series, released 2026-08-26 by Z.ai. The first open-source frontier model to pair sparse attention with linear attention in one hybrid architecture, cutting attention compute and KV cache by 3.01x and 4.44x vs GLM-5.3 while keeping precise long-context ability. Also adopts Manifold-Constrained Hyper-Connections mHC . 1M context, MIT license on HuggingFace at zai-org/GLM-5.3-Flash , trained on a 30T-token multimodal corpus. This is the ox-alpha reveal. Z.ai confirmed that the anonymous “ox-alpha” stealth model previewed free on OpenRouter/OpenCode from Aug 20 was an early version of GLM-5.3-Flash - and that the preview’s entire traffic load was served on Chinese AI chips . Per Z.ai Zixuan Li , the official release is stronger and significantly more stable than the ox-alpha preview. OpenCode reported 42T tokens served in 6 days , making it the most-used model after DeepSeek Flash’s 56-day run. The ox-alpha row in this catalog is superseded by this model. Native multimodal coding. Visual capabilities are built into the coding loop - the model observes interfaces, rendered results, and interaction feedback, then tests and improves its work. Coordinates across code, browsers, and GUIs BUA/CUA for frontend dev, game creation, and Blender 3D scenes. Beyond coding it handles Office, financial research, and document workflows, producing finished PPTX/PDF/DOCX/XLSX. Benchmarks Z.ai self-reported . GLM-5.3-Flash vs GLM-5.2 and the frontier field - strongest open-source multimodal-coder result across the board, beating GLM-5.2 on every coding and agentic row and leading open-source vision OfficeQA Pro, CharXiv w/ tools, Chartography w/ tools : | Benchmark | GLM-5.3-Flash | GLM-5.2 | DeepSeek-V4-Vision-Exp | Opus 4.8 | GPT-5.6 Terra | Gemini 3.7 Flash | |---|---|---|---|---|---|---| | Terminal Bench 2.1 | 84.3 | 81.0 | 83.9 | 85.0 | 87.4 | 85.8 | | DeepSWE v1.1 | 63.4 | 46.2 | 59.3 | 58.0 | 69.6 | 65.3 | | NL2Repo | 56.3 | 48.9 | 57.7 | 69.7 | - | - | | Toolathlon Verified | 78.4 | 59.9 | 75.9 | 76.2 | 74.9 | - | | AutomationBench v1.0.6 | 48.8 | 26.2 | 38.8 | 41.0 | 37.2 | 52.3 | | Agents’ Last Exam | 26.3 | 20.4 | 27.3 | 27.0 | 28.0 | - | | HLE w/ Tools | 55.3 | 54.7 | 55.1 | 57.9 | - | - | | GDPval-AA v2 | 1773 | 1504 | 1675 | 1582 | 1571 | 1527 | | OfficeQA Pro | 62.4 | - | 57.9 | 48.9 | - | - | | CharXiv Reasoning w/ Tools | 89.4 | - | 80.4 | 89.9 | 88.0 | 88.7 | | Chartography w/ Tools | 78.0 | - | 64.3 | 75.0 | 68.0 | 65.0 | | BabyVision | 53.4 | - | 35.1 | 46.8 | 61.6 | 70.9 | | MVbench | 77.8 | - | 69.4 | 67.1 | 75.0 | 82.2 | | MMVU | 80.5 | - | 72.7 | 67.4 | 75.8 | 82.3 | Access: model code glm-5.3-flash ; OpenAI- and Anthropic-compatible APIs; GLM Coding Plan 3x GLM-5.3 quota, off-peak 50% points . Recommended sampling: temperature 1, top p 0.95, reasoning effort max, thinking always on. Local-run status: weights landed on launch day under MIT, but no community quant sizes are published yet, so no model variants are seeded same posture as GLM 5.3 pre-quants . Treat launch benchmarks as Z.ai self-reported until independently replicated. - 320.0B - 1000k - mit - 🇨🇳 China - Aug 2026 Scores Or run it in the cloud Live per-provider pricing, throughput and uptime. Click a column to sort. | Provider | Type | Input $/M | Output $/M | Cache $/M | Tok/s | Latency | Uptime | Value | |---|---|---|---|---|---|---|---|---| | Sub | - | - | - | - | - | - | $10.00/mo Coding Plan Lite | | | Sub | - | - | - | - | - | - | $10.00/mo Go $5 first month | | | Sub | - | - | - | - | - | - | $30.00/mo Coding Plan Pro | | | Sub | - | - | - | - | - | - | $80.00/mo Coding Plan Max | Default order: throughput among 95%+ uptime providers, then latency; subscriptions last. Sort by any column. Subscription rows show $/mo in the Value column - per-token columns are "-". Affiliate links are marked sponsored / nofollow. Confirm current pricing on the provider's site before committing. Detailed API pricing page + JSON endpoint → /models/glm-5-3-flash/pricing See who runs Zhipu AI in production → /adoption/zhipu Inference cost over time Data accumulates from the first daily sync - longer ranges populate over time. Prices come from OpenRouter snapshots, not a historical API.