From 36 Minutes to 2.6: Making a Latent-GRPO Training Step on Qwen 3.8 27B 14× Faster
A latent-GRPO training step for the Qwen 3.8 27B model dropped from a median of 36 minutes on a single H100 to 2.6 minutes on one 8× NVIDIA B300 node, a 14× speedup, according to the training team's a…