Standard LLM serving remains bound by memory bandwidth, requiring a full forward pass for every output token. Speculative decoding in vLLM attempts to break this bottleneck on AMD Instinct MI300X and MI355X GPUs by using fast draft methods like MTP, EAGLE-3, DFlash, and DSpark to propose candidates verified in a single pass.
However, speculative decoding is not an automatic win. On structured outputs like code or JSON, poor draft acceptance rates introduce significant compute overhead. Discarded candidate tokens and KV cache rollbacks can make speculative serving slower than standard autoregressive baselines.
Maximizing inference throughput on non-Nvidia hardware requires monitoring token acceptance histograms and dynamic proposal length tuning rather than relying on aggregate benchmark speedups.
Read the full article: Evaluating Speculative Decoding in vLLM on AMD MI300X GPUs