GLM 5.3 Available in OpenRouter Z.ai released GLM-5.3, a large-scale reasoning model for complex software engineering and long-horizon agent tasks, now available on OpenRouter with a 1M-token context window and pricing of $1.40 per 1M input tokens and $4.40 per 1M output tokens. The model improves on GLM-5.2 in coding and token efficiency, supports reasoning efforts low, high, and max (default), and achieves 99.91% uptime with a P50 latency of 5.71 seconds and throughput of 27 tokens per second. GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency. Reasoning is always on and cannot be disabled. Reasoning efforts low , high , and max are supported; max is the default. Modalities In / Out Price $1.40 / $4.40per 1M Context 1M Released Aug 18, 2026 | $1.40 | $4.40 | $0.26 | 5.71s | 27 tps | Throughput 27tok/s P50, best across providers Latency 5.71s P50, best provider 99.91% 99.65% When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API /docs/api/api-reference/endpoints/list-endpoints . Learn more /docs/provider-routing about our load balancing and customization options.