Boosting MoE Training Throughput with Advanced Fusion Kernels
NVIDIA introduced advanced fused MLP kernels for mixture-of-experts (MoE) models, built with the CuTe DSL, delivering 1.3x–2x kernel-level speedups and enabling sync-free MoE execution. The optimization contributed an 8%…