Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It A study of training-inference mismatch in reinforcement learning with verifiable rewards (RLVR) for large language models finds that rollouts sampled by an inference engine and gradients computed by a training engine assign different probabilities to the same tokens. The research examines where the mismatch arises and how to correct it. We study training-inference mismatch in reinforcement learning with verifiable rewards RLVR for large language models, where rollouts are sampled by an inference engine while gradients are computed by a training engine, and the two engines assign different probabilities to the same tokens. To acco