Efficient MoE Training for Biological Foundation Models
NVIDIA published a tutorial showing how its Transformer Engine (TE) and BioNeMo MoE recipe cut the overhead of training mixture-of-experts biological foundation models, using GroupedLinear to submit expert work as one gr…