New Model Available: Z.AI: GLM 5.3 Flash Z.ai released GLM-5.3-Flash, a native multimodal model designed for efficient coding and long-horizon agent tasks, featuring a hybrid sparse and linear attention architecture that reduces compute overhead while maintaining long-context accuracy. Pricing starts at $0.15 per million input tokens and $0.5 per million output tokens, with some providers offering discounted rates. GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while reducing compute overhead. Back to Models https://zenmux.ai/models Providers Route requests across multiple providers. Copy a provider slug to set your preference. $0.15 / M tokens $0.5 / M tokens Read: 0.03 / M tokens Write: - / M tokens1M6.3s37.6tps $0.15 / M tokens $0.5 / M tokens Read: 0.03 / M tokens Write: - / M tokens1M1.96s69.9tps ~~$0.15~~ $0.075 / M tokens ~~$0.5~~ $0.25 / M tokens Read: ~~0.03~~ 0.015 / M tokens Write: - / M tokens1M5.64s23.6tps Uptime 24hours Direct request success rate on AI Gateway and per-provider. Throughput 24hours P50 throughput on live AI Gateway traffic, in tokens per second TPS . Latency 24hours P50 time to first token TTFT on live AI Gateway traffic, in milliseconds. Activity Token volume and request traffic to this model over time. Apps Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for. View All https://zenmux.ai/analytics/apps Related Models More models from Z.ai https://zenmux.ai/z-ai