cd /news/machine-learning/abstention-as-an-action-can-kill-bot… · home topics machine-learning article
[ARTICLE · art-85548] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

Abstention as an Action Can Kill Both the Reward Gradient and the KL Anchor: Collapse Law and Repair for Error-Penalized Reinforcement Learning

A new arXiv preprint (2608.00301v1) proves that error-penalized scoring rules, which reward correct answers with +1, penalize wrong ones with -λ, and give 0 for abstention, can cause KL-anchored gradient learners to drift toward refusing all questions, with mean training reward rising to zero like 1/t while coverage collapses. The authors show that when abstention is a discrete action, the reward gradient and the anchor's restoring force are throttled by the same gate-saturation factor, and that the advantage estimator compounds the failure by replacing designed penalties with an effective penalty of one, shifting the learned threshold from t* to 1/2. They propose a structural repair: train a mandatory confidence report with a strictly proper score plus a correctness reward, abstaining only at deployment by thresholding the report, which simulations and language model experiments confirm improves coverage, accuracy, and calibration.

read1 min views1 publishedAug 4, 2026

arXiv:2608.00301v1 Announce Type: new Abstract: Error-penalized scoring rules ($+1$ for a correct answer, $-\lambda$ for a wrong one, $0$ for abstaining) are increasingly prescribed against hallucination: a rational agent facing such a rule answers exactly when its correctness probability exceeds Chow's threshold $t^\ast=\lambda/(1+\lambda)$. We prove that a KL-anchored gradient learner can do the opposite. When abstention is a discrete action, the reward gradient and the anchor's restoring force are throttled by the same gate-saturation factor and die together: under explicit conditions (among them, blanket answering loses score in expectation and prompts share a bounded readout) the model drifts toward refusing everything, its mean training reward rising to zero like $1/t$ in training time $t$, so the curve reads as improvement while coverage collapses. The advantage estimator compounds the failure: in its sparse-answer regime, group normalization silently replaces every designed penalty with an effective penalty of one, moving the learned threshold from $t^\ast$ to $1/2$. The repair is structural: train a mandatory confidence report with a strictly proper score plus a correctness reward, and abstain only at deployment by thresholding the report. The always-emitted report has no gate to saturate, so no shared factor can kill its reward gradient and its anchor together, and its calibrated optimum is attracting. Simulations confirm every prediction, and experiments on language models at two scales confirm the mechanism live: the rule silences questions the models demonstrably still solve within ten optimizer steps, an ablation isolates the cause, and report-level training raises coverage, accuracy, and calibration together.

── more in #machine-learning 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/abstention-as-an-act…] indexed:0 read:1min 2026-08-04 ·