00:00
2026-08-24
rocm.blogs.amd.com
artificial-intelligence
Serving 64Mi-Token Contexts on One AMD Instinctβ’ MI355X Node
On a single 8-GPU AMD Instinct MI355X node, AMD served Kimi Linear 48B-A3B with context lengths from 1024 tokens to 64Mi (67,108,864), using vLLM with FP8 KV cache and tensor parallelism 8, recording β¦