TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models A paper titled TRACE introduces rollout-guided quantization-aware training for FP4 reinforcement learning of mixture-of-experts (MoE) language models, targeting the computation and memory overhead of rollout generation in RL post-training. The work addresses a limitation in existing FP4 RL methods, which the paper says primarily optimize other aspects of the pipeline rather than the rollout. No specific benchmark figures, dates, or author names were provided in the available text. Reinforcement learning RL for post-training large language models LLMs incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily opti