00:00
2026-09-16
rocm.blogs.amd.com
artificial-intelligence
DFlash Speculative Decoding on AMD Instinct MI355X: Up to 5× Faster Qwen3.5 Inference
DFlash, a block-diffusion speculative decoding drafter, delivered up to 5.02× single-request throughput on Qwen3.5-27B when run on AMD Instinct MI355X through vLLM on ROCm, according to the benchmark.…