09:00
2026-09-07
arxiv.org
artificial-intelligence
Fast Inference from Transformers via Speculative Decoding
Researchers introduced speculative decoding, an algorithm that accelerates inference from large autoregressive models like Transformers by computing several tokens in parallel without changing outputs…