09:26
2026-09-07
vllm.ai
large-language-models
Speculative Decoding in vLLM on AMD GPUs
AMD's experiments with speculative decoding in vLLM on AMD Instinct MI300X and MI355X GPUs using the ROCm platform show that output-token throughput gains vary by drafting method, proposal length, mod…