Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning
A new arXiv preprint (2608.00301v1) proves that error-penalized scoring rules, which reward correct answers with +1, penalize wrong ones with -λ, and give 0 for abstention, can cause KL-anchored gradient learners to drif…