What Actually Happens During Speculative Decoding in LLMs
A developer's technical breakdown of speculative decoding explains that single-batch LLM generation is memory-bandwidth bound rather than compute bound: on an Nvidia H100, a 70B FP16 model spends roug…