{"slug": "abstention-as-an-action-can-kill-both-the-reward-gradient-and-the-kl-anchor-law", "title": "Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning", "summary": "A new arXiv preprint (2608.00301v1) proves that error-penalized scoring rules, which reward correct answers with +1, penalize wrong ones with -λ, and give 0 for abstention, can cause KL-anchored gradient learners to drift toward refusing all questions, with mean training reward rising to zero like 1/t while coverage collapses. The authors show that when abstention is a discrete action, the reward gradient and the anchor's restoring force are throttled by the same gate-saturation factor, and that the advantage estimator compounds the failure by replacing designed penalties with an effective penalty of one, shifting the learned threshold from t* to 1/2. They propose a structural repair: train a mandatory confidence report with a strictly proper score plus a correctness reward, abstaining only at deployment by thresholding the report, which simulations and language model experiments confirm improves coverage, accuracy, and calibration.", "body_md": "arXiv:2608.00301v1 Announce Type: new\nAbstract: Error-penalized scoring rules ($+1$ for a correct answer, $-\\lambda$ for a wrong one, $0$ for abstaining) are increasingly prescribed against hallucination: a rational agent facing such a rule answers exactly when its correctness probability exceeds Chow's threshold $t^\\ast=\\lambda/(1+\\lambda)$. We prove that a KL-anchored gradient learner can do the opposite. When abstention is a discrete action, the reward gradient and the anchor's restoring force are throttled by the same gate-saturation factor and die together: under explicit conditions (among them, blanket answering loses score in expectation and prompts share a bounded readout) the model drifts toward refusing everything, its mean training reward rising to zero like $1/t$ in training time $t$, so the curve reads as improvement while coverage collapses. The advantage estimator compounds the failure: in its sparse-answer regime, group normalization silently replaces every designed penalty with an effective penalty of one, moving the learned threshold from $t^\\ast$ to $1/2$. The repair is structural: train a mandatory confidence report with a strictly proper score plus a correctness reward, and abstain only at deployment by thresholding the report. The always-emitted report has no gate to saturate, so no shared factor can kill its reward gradient and its anchor together, and its calibrated optimum is attracting. Simulations confirm every prediction, and experiments on language models at two scales confirm the mechanism live: the rule silences questions the models demonstrably still solve within ten optimizer steps, an ablation isolates the cause, and report-level training raises coverage, accuracy, and calibration together.", "url": "https://wpnews.pro/news/abstention-as-an-action-can-kill-both-the-reward-gradient-and-the-kl-anchor-law", "canonical_source": "https://arxiv.org/abs/2608.00301", "published_at": "2026-08-04 04:00:00+00:00", "updated_at": "2026-08-04 04:34:25.756944+00:00", "lang": "en", "topics": ["machine-learning", "ai-safety", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/abstention-as-an-action-can-kill-both-the-reward-gradient-and-the-kl-anchor-law", "markdown": "https://wpnews.pro/news/abstention-as-an-action-can-kill-both-the-reward-gradient-and-the-kl-anchor-law.md", "text": "https://wpnews.pro/news/abstention-as-an-action-can-kill-both-the-reward-gradient-and-the-kl-anchor-law.txt", "jsonld": "https://wpnews.pro/news/abstention-as-an-action-can-kill-both-the-reward-gradient-and-the-kl-anchor-law.jsonld"}}