FP8 Reinforcement Learning in SkyRL: Preserving Policy Consistency Across Training and Rollout
SkyRL now supports FP8-accelerated training and rollout with on-policy weight sync, reducing end-to-end step time by up to 23% while closely tracking BF16 convergence in long-run RL experiments. The s…