RL Systems Mind the Gap: Matching Trainer and Generator Throughput
Anthropic CEO Dario Amodei said reinforcement learning shows the same log-linear scaling as pre-training, but RL system efficiency is critical to afford enough training. Experiments on open models show that matching trai…