# New Model Available: Z.AI: GLM 5.3 Flash

> Source: <https://zenmux.ai/z-ai/glm-5.3-flash>
> Published: 2026-08-26 15:12:37+00:00

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead.

[Back to Models](https://zenmux.ai/models)

## Providers

Route requests across multiple providers. Copy a provider slug to set your preference.

**$0.15**

*/ M tokens*

**$0.5**

*/ M tokens*

*Read:*

**0.03**/ M tokens

*Write:*

**-**/ M tokens1M6.3s37.6tps

**$0.15**

*/ M tokens*

**$0.5**

*/ M tokens*

*Read:*

**0.03**/ M tokens

*Write:*

**-**/ M tokens1M1.96s69.9tps

~~$0.15~~

**$0.075**

*/ M tokens*

~~$0.5~~

**$0.25**

*/ M tokens*

*Read:*

~~0.03~~

**0.015**/ M tokens

*Write:*

**-**/ M tokens1M5.64s23.6tps

## Uptime

24hours
Direct request success rate on AI Gateway and per-provider.

## Throughput

24hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS).

## Latency

24hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds.

## Activity

Token volume and request traffic to this model over time.

## Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. [View All](https://zenmux.ai/analytics/apps)

## Related Models

More models from [Z.ai](https://zenmux.ai/z-ai)
