00:00
2026-09-25
rocm.blogs.amd.com
ai-infrastructure
UltraQuant on AMD Instinct: More Efficient Agentic Serving for Qwen3.8-MXFP4
AMD reported that UltraQuant, its 4-bit KV-cache method for the Qwen3.8-2.4T MXFP4 model on AMD Instinct MI355X GPUs, serves 29% more tokens per second at 24% lower per-token latency than the fastest …