# GLM 5.3 flash on AMD GPUs 670 tok/s

> Source: <https://twitter.com/runinfrai/status/2105186912634057086>
> Published: 2026-09-30 06:57:30+00:00

we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today
670 tok/s on Vercel AI Gateway
$0.11 per 1M input, $0.45 per 1M output, $0.03 per 1M cached. 1M token context. FP8, vendor-native release
it now runs on AMD. same model, same API, more capacity behind it
99.7% cache hit rate over the last 24 hours. every hit and miss shows up in your dashboard, so you can track the cache logs yourself
OpenAI-compatible chat completions and Anthropic-compatible /v1/messages. text and image input, tool calling, JSON mode, streaming
zero data retention. never used for training
