Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR A new reinforcement learning technique called Off-Policy-Aware Cross-Model Trajectory Exchange targets the all-fail group problem in Reinforcement Learning with Verifiable Rewards (RLVR) methods such as GRPO, where finite rollout budgets yield no reward-based policy-gradient signal. The approach exchanges successful trajectories across models rather than spending more rollouts to raise the chance of success at higher cost. Reinforcement Learning with Verifiable Rewards RLVR methods such as GRPO rely on successful self-generated trajectories, but finite rollout budgets can produce all-fail groups with no reward-based policy-gradient signal. While additional rollouts improve the chance of success at higher cost, succe