18:28
2026-09-03
dev.to
large-language-models
LLMs Don't Have to Generate One Token at a Time: How Medusa and Multi-Token Prediction Cheat Autoregression
Shrijith Venkatramana, developer of LiveReview, explains how autoregressive LLM decoding can be accelerated using speculative decoding and Medusa-style multi-token prediction, which break the sequentiβ¦