LLMs Don't Have to Generate One Token at a Time: How Medusa and Multi-Token Prediction Cheat Autoregression
Shrijith Venkatramana, developer of LiveReview, explains how autoregressive LLM decoding can be accelerated using speculative decoding and Medusa-style multi-token prediction, which break the sequenti…