04:00
2026-07-29
arxiv.org
artificial-intelligence
SpecPrefetch: Parameter-Efficient Expert Prefetching for Sparse MoE Foundation Models
Researchers propose SpecPrefetch, a parameter-efficient prefetching framework for offloaded sparse Mixture-of-Experts (MoE) inference that uses a shared lightweight adapter to predict next-layer experβ¦