19:00
2026-09-21
systems.seas.harvard.edu
artificial-intelligence
H-Spec: Parallel Speculative Decoding Without A Drafter-Side KV Cache
Researchers Weifan Jiang and colleagues at Harvard University proposed H-Spec, a hybrid Mamba-attention parallel speculative decoding drafter that eliminates the drafter-side KV cache by reusing targeβ¦