# New Model Available: GLM 5.3 Prime

> Source: <https://zenmux.ai/z-ai/glm-5.3-prime>
> Published: 2026-10-08 12:25:46+00:00

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× higher output throughput through inference acceleration. It supports text input and output with a 1M-token context window and up to 128K output tokens, and is optimized for coding and agentic workloads, including long-horizon multi-turn agent orchestration, real-time conversation, and streaming code generation.

[Back to Models](https://zenmux.ai/models)

## Providers

Route requests across multiple providers. Copy a provider slug to set your preference.

**$2.8**

*/ M tokens*

**$8.8**

*/ M tokens*

*Read:*

**0.56**/ M tokens

*Write:*

**-**/ M tokens1M--

## Uptime

24hours
Direct request success rate on AI Gateway and per-provider.

## Throughput

24hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS).

## Latency

24hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.

## Activity

Token volume and request traffic to this model over time.

## Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows

Metrics sourced from[Artificial Analysis](https://artificialanalysis.ai/)

## Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. [View All](https://zenmux.ai/analytics/apps)

## Related Models

More models from [Z.ai](https://zenmux.ai/z-ai)
