{"slug": "from-refuse-to-richness-rubric-rewards-for-long-form-hallucination-reinforcement", "title": "From Refuse to Richness: Rubric Rewards for Long-Form Hallucination Reinforcement Learning", "summary": "A new arXiv preprint (2608.12337v1) from researchers studying long-form hallucination reinforcement learning finds that strict grounding rewards improve factual support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. The authors propose a soft combination of grounding, rubric coverage, and relevance rewards, which achieves the best balance in experiments, improving in-distribution support and transferring better to out-of-distribution checklist tasks than either reward alone.", "body_md": "arXiv:2608.12337v1 Announce Type: new\nAbstract: Rewards that penalize unsupported claims can improve grounding in long-form generation, but they can also teach models to answer less. We study this refusal-to-richness trade-off in long-form hallucination RL. Instead of using global richness proxies such as length, claim count, detail, or pairwise relevance, we represent each question with a key-point rubric that specifies the required and optional information a useful answer should cover. These rubrics define coverage directly and are used both for evaluation and as reward signals. Across grounding-only, proxy-based, rubric-only, and combined rewards, we find a stable trade-off: strict grounding rewards improve support but suppress coverage, while unconstrained rubric rewards improve coverage but weaken grounding. A soft combination of grounding, rubric coverage, and relevance gives the best balance in our experiments, improving in-distribution support while transferring better to out-of-distribution checklist tasks than either grounding-only or rubric-only rewards.", "url": "https://wpnews.pro/news/from-refuse-to-richness-rubric-rewards-for-long-form-hallucination-reinforcement", "canonical_source": "https://arxiv.org/abs/2608.12337", "published_at": "2026-08-14 04:00:00+00:00", "updated_at": "2026-08-14 04:12:57.071711+00:00", "lang": "en", "topics": ["large-language-models", "machine-learning", "ai-research"], "entities": ["arXiv"], "alternates": {"html": "https://wpnews.pro/news/from-refuse-to-richness-rubric-rewards-for-long-form-hallucination-reinforcement", "markdown": "https://wpnews.pro/news/from-refuse-to-richness-rubric-rewards-for-long-form-hallucination-reinforcement.md", "text": "https://wpnews.pro/news/from-refuse-to-richness-rubric-rewards-for-long-form-hallucination-reinforcement.txt", "jsonld": "https://wpnews.pro/news/from-refuse-to-richness-rubric-rewards-for-long-form-hallucination-reinforcement.jsonld"}}