04:00
2026-07-28
machinebrief.com
artificial-intelligence
DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference
Researchers propose DraftExpert, an expansion-aware self-speculative decoding framework for expert-offloaded Mixture-of-Experts (MoE) inference on end devices, achieving 1.45x average decode throughpuβ¦