16:32
2026-08-01
pub.towardsai.net
machine-learning
Speculative Decoding Collapses at Batch 224 on Llama-3-70B and Never on gpt-oss-120B
A 180-line roofline model built by an independent developer shows speculative decoding on Llama-3-70B collapses at batch 224, turning a 1.92x speedup into a 0.74x slowdown by batch 512, while the sameβ¦