SpecLA: Efficient Speculative Decoding for Linear-Attention Models
SpecLA, a speculative decoding runtime for stateful linear-attention models, achieves up to 1.70x end-to-end speedup over autoregressive decoding on an NVIDIA H100 with a GDN-1.3B target, according to…