23:55
2026-08-27
supercomputing-system-ai-lab.github.io
artificial-intelligence
XPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding
XPress, a lightweight causal refiner for diffusion drafters in speculative decoding, raises acceptance length by ~30% on average (up to +56%) and decoding throughput by ~1.3Ć on average (up to 1.7Ć) cā¦