New Model Available: GLM 5.3 Prime Z.ai released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5–2× higher output throughput through inference acceleration while inheriting the base model's full capabilities. The model supports text input and output with a 1M-token context window and up to 128K output tokens, and is optimized for coding and agentic workloads including long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation. Pricing listed on the model page starts at $2.8 per million tokens, with read pricing at $0.56 per million tokens. GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× higher output throughput through inference acceleration. It supports text input and output with a 1M-token context window and up to 128K output tokens, and is optimized for coding and agentic workloads, including long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation. Back to Models https://zenmux.ai/models Providers Route requests across multiple providers. Copy a provider slug to set your preference. $2.8 / M tokens $8.8 / M tokens Read: 0.56 / M tokens Write: - / M tokens1M-- Uptime 24hours Direct request success rate on AI Gateway and per-provider. Throughput 24hours P50 throughput on live AI Gateway traffic, in tokens per second TPS . Latency 24hours P50 time to first token TTFT on live AI Gateway traffic, in milliseconds. Activity Token volume and request traffic to this model over time. Benchmarks Scores on standardized evaluations. Higher percentages are better — and rank percentile shows Metrics sourced from Artificial Analysis https://artificialanalysis.ai/ Apps Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. View All https://zenmux.ai/analytics/apps Related Models More models from Z.ai https://zenmux.ai/z-ai