cd /news/large-language-models/rewarding-efficient-reasoning-improv… · home topics large-language-models article
[ARTICLE · art-135523] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=↑ positive

Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models

A new arXiv paper (2609.20846v1) reports that fine-tuning several 4B large reasoning models with a novel GRPO reward that encourages efficient reasoning about whether a task contains enough information improves abstention performance by 12.8% on average while shortening chains of thought by 44% on average. The authors compared large reasoning model behavior against a human study, finding that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas the models generate longer chains of thought on unanswerable than on answerable prompts. The reward is inspired by a resource-rational perspective on human cognition and retains the models' answering capabilities.

by read1 min views2 publishedSep 21, 2026

arXiv:2609.20846v1 Announce Type: new Abstract: While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models' efficiency (44% shorter CoTs on average).

── more in #large-language-models 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/rewarding-efficient-…] indexed:0 read:1min 2026-09-21 ·