{"slug": "routepack-expert-placement-and-attention-aware-data-packing-for-moe-learning", "title": "RoutePack: Expert Placement and Attention-Aware Data Packing for MoE Reinforcement Learning", "summary": "A new system called RoutePack, developed by researchers, improves training throughput for Mixture-of-Experts (MoE) reinforcement learning models by coordinating expert placement and data packing. On Ling-3.0-Tiny and Ling-3.0-Flash, RoutePack boosts mean trainer-measured token throughput by 8.85% and 14.89% over baseline, respectively, combining expert rerouting gains of 3.80% and 10.50% with routing-aware packing gains of 4.86% and 3.98%.", "body_md": "arXiv:2608.12146v1 Announce Type: cross\nAbstract: Training Mixture-of-Experts (MoE) models for reinforcement learning (RL) couples two load-balancing problems: sequence composition determines dense attention work in each data-parallel microbatch, while token routing determines sparse expert work on expert-parallel ranks. Optimizing either alone can shift the bottleneck to the other. In MoE RL, rollout-time routing replay exposes every sample's sequence length and layer-wise expert demand before its training step. We present RoutePack, a hierarchical planner that coordinates state-consistent, layer-wise expert rerouting with joint attention- and expert-aware data packing over an optimizer-step window. RoutePack first places experts independently at each MoE layer using aggregate routing demand. It then packs samples into the smallest certified, or best-known feasible, number of token-capped execution rows and optimizes their DP layout with a projected EDP-shard-aware objective. The objective combines a window-normalized linear-quadratic attention proxy with per-layer physical EP-rank peaks and minimizes the accumulated cost of the slowest EDP shard. Parallel population annealing searches fixed-row feasible layouts while preserving sample coverage, capacity, nonempty cells, equal microbatch counts, and communicator topology. State-consistent materialization preserves logical top-k routing and existing MoE kernels without microbatch-level expert replication. Across Ling-3.0-Tiny and Ling-3.0-Flash, expert rerouting improves mean trainer-measured token throughput by 3.80% and 10.50%, while routing-aware packing adds another 4.86% and 3.98%, respectively. Overall, RoutePack improves throughput by 8.85% and 14.89% over the baseline.", "url": "https://wpnews.pro/news/routepack-expert-placement-and-attention-aware-data-packing-for-moe-learning", "canonical_source": "https://www.machinebrief.com/news/routepack-expert-placement-and-attention-aware-data-packing-74bo", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 05:11:18.464885+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["RoutePack", "Ling-3.0-Tiny", "Ling-3.0-Flash"], "alternates": {"html": "https://wpnews.pro/news/routepack-expert-placement-and-attention-aware-data-packing-for-moe-learning", "markdown": "https://wpnews.pro/news/routepack-expert-placement-and-attention-aware-data-packing-for-moe-learning.md", "text": "https://wpnews.pro/news/routepack-expert-placement-and-attention-aware-data-packing-for-moe-learning.txt", "jsonld": "https://wpnews.pro/news/routepack-expert-placement-and-attention-aware-data-packing-for-moe-learning.jsonld"}}