04:00
2026-07-24
arxiv.org
artificial-intelligence
When RLVR Shrinks the Reasoning Boundary: Diagnosing Pass@k Inversion
A new arXiv study (2607.20543v1) finds that reinforcement learning with verifiable rewards (RLVR) can improve one-sample accuracy while causing pass@k inversion, where the trained policy solves fewer โฆ