04:00
2026-07-13
arxiv.org
machine-learning
Sticky Routing: Training MoE Models for Memory-Efficient Inference
Researchers propose StickyMoE, a differentiable routing consistency loss that reduces expert switch rates by up to 60% in Mixture-of-Experts models during pretraining, with less than 4% perplexity degβ¦