The economics of speculative decoding
Speculative decoding, a lossless inference optimisation that predicts future tokens to reduce latency, faces new economic constraints as modern mixture-of-experts (MoE) architectures replace dense transformers. MoE layer…