GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves on GLM-5.2 in coding and in the balance between performance and token efficiency.
Reasoning is always on and cannot be disabled. Reasoning efforts low
, high
, and max
are supported; max
is the default.
Modalities
In / Out Price
$1.40 / $4.40per 1M
Context
1M
Released
Aug 18, 2026
| $1.40 | $4.40 | $0.26 | 5.71s | 27 tps |
Throughput
27tok/s
P50, best across providers
Latency
5.71s
P50, best provider
99.91%
99.65%
When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.