{"slug": "new-model-available-glm-5-3-prime", "title": "New Model Available: GLM 5.3 Prime", "summary": "Z.ai released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5–2× higher output throughput through inference acceleration while inheriting the base model's full capabilities. The model supports text input and output with a 1M-token context window and up to 128K output tokens, and is optimized for coding and agentic workloads including long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation. Pricing listed on the model page starts at $2.8 per million tokens, with read pricing at $0.56 per million tokens.", "body_md": "GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× higher output throughput through inference acceleration. It supports text input and output with a 1M-token context window and up to 128K output tokens, and is optimized for coding and agentic workloads, including long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation.\n\n[Back to Models](https://zenmux.ai/models)\n\n## Providers\n\nRoute requests across multiple providers. Copy a provider slug to set your preference.\n\n**$2.8**\n\n*/ M tokens*\n\n**$8.8**\n\n*/ M tokens*\n\n*Read:*\n\n**0.56**/ M tokens\n\n*Write:*\n\n**-**/ M tokens1M--\n\n## Uptime\n\n24hours\nDirect request success rate on AI Gateway and per-provider.\n\n## Throughput\n\n24hours\nP50 throughput on live AI Gateway traffic, in tokens per second (TPS).\n\n## Latency\n\n24hours\nP50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.\n\n## Activity\n\nToken volume and request traffic to this model over time.\n\n## Benchmarks\n\nScores on standardized evaluations. Higher percentages are better — and rank percentile shows\n\nMetrics sourced from[Artificial Analysis](https://artificialanalysis.ai/)\n\n## Apps\n\nPublic apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. [View All](https://zenmux.ai/analytics/apps)\n\n## Related Models\n\nMore models from [Z.ai](https://zenmux.ai/z-ai)", "url": "https://wpnews.pro/news/new-model-available-glm-5-3-prime", "canonical_source": "https://zenmux.ai/z-ai/glm-5.3-prime", "published_at": "2026-10-08 12:25:46+00:00", "updated_at": "2026-10-08 12:50:31.620334+00:00", "lang": "en", "topics": ["large-language-models", "ai-agents", "ai-products", "generative-ai"], "entities": ["Z.ai", "GLM-5.3-Prime", "GLM-5.3", "Artificial Analysis"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/new-model-available-glm-5-3-prime", "markdown": "https://wpnews.pro/news/new-model-available-glm-5-3-prime.md", "text": "https://wpnews.pro/news/new-model-available-glm-5-3-prime.txt", "jsonld": "https://wpnews.pro/news/new-model-available-glm-5-3-prime.jsonld"}}