MoE Memory Breakthrough: Training Trillion-Parameter Models with Ease
Researchers have developed a memory-efficient training stack for Mixture-of-Experts models that achieves 4.7x to 8.2x higher per-GPU throughput compared to a finely-tuned FSDP2 baseline, using fewer than 12 8x H200 GPU n…