Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine
NVIDIA reported that combining its Transformer Engine library with the JAX Python library raised DeepSeek-V3 MoE training throughput on NVIDIA GB200 GPUs from an unoptimized baseline of 103 TFLOPS/GPU…