Speculative Decoding Collapses at Batch 224 on Llama-3-70B and Never on gpt-oss-120B
A 180-line roofline model built by an independent developer shows speculative decoding on Llama-3-70B collapses at batch 224, turning a 1.92x speedup into a 0.74x slowdown by batch 512, while the same…