# Towards Full Pipeline FP8 Reinforcement Learning for LLMs

> Source: <https://aiflash.com/news/124643/>
> Published: 2026-09-22 21:01:17+00:00

Reinforcement learning (RL) has become a key technique for improving the reasoning and agentic abilities of large language models (LLMs). Although FP8 quantization can accelerate RL training, maintaining stability throughout an FP8 RL pipeline remains challenging. While previous works have focused o
