05:00
2026-08-23
dev.to
machine-learning
Expert-locality-aware decode routing reduces MoE serving latency
Researchers introduced ELDR, an expert-locality-aware decode routing method for mixture-of-experts (MoE) models, which reduces median time-per-output-token by 5.9โ13.9% across three MoE models and twoโฆ