When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
A new arXiv study (2607.20543v1) finds that reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while causing pass@k inversion, where the trained policy solves fewer distinct problems th…