Speculative Decoding, Explained: The Free Speed Toggle Your Local LLM Is Probably Not Using
Speculative decoding can speed up local large language model inference by 1.5 to 2.5 times without changing output quality, according to research from Google and DeepMind. The technique uses a small draft model to guess …