17:52
2026-09-15
gist.github.com
large-language-models
FP8 MTP draft stack for vLLM ā +6.6% decode on modelopt NVFP4 Qwen3.5-family checkpoints
A vLLM contributor published an FP8 multi-token-prediction (MTP) draft stack for vLLM that yields a 6.6% decode throughput gain on modelopt NVFP4 Qwen3.5-family checkpoints. The change quantizes the pā¦