cd /news/large-language-models/reinforcement-learning-for-large-lan… · home topics large-language-models article
[ARTICLE · art-69683] src=machinebrief.com ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results

A new benchmark and training set called SelectBench, introduced by researchers in arXiv:2607.20090v1, shows that reinforcement learning via DAPO can modestly improve selective evidence adoption in retrieval-augmented large language models. Post-training Qwen3.5-4B with DAPO-Rule and DAPO-DeepSeek raised strict success on SelectBench-v2 from 22.46% to 25.54% and 26.46%, respectively, while reducing forbidden-content adoption and preserving general capabilities on MMLU and HotpotQA. However, the gains were not statistically significant after Holm correction, and prompt-injection resistance did not improve, highlighting remaining challenges.

read1 min views1 publishedJul 23, 2026

arXiv:2607.20090v1 Announce Type: new Abstract: Retrieval-augmented large language models frequently face contexts that interleave useful evidence with misleading statements or instruction-like content. Blanket refusal discards valid evidence, whereas uncritical adoption yields incorrect or unsafe answers. The ability to selectively adopt relevant information while rejecting deceptive or harmful content is therefore critical for reliable deployment in real-world retrieval settings. We introduce SelectBench, a controlled benchmark and training set for selective evidence adoption, and post-train Qwen3.5-4B directly with DAPO using either deterministic rule rewards or a frozen semantic judge. On the corrected 325-example SelectBench-v2 test set, strict success rises from 22.46% for the original checkpoint to 25.54% with DAPO-Rule and 26.46% with DAPO-DeepSeek. Both trained policies reduce forbidden-content adoption and produce shorter, more focused responses, yet prompt-injection following does not improve. The paired gains are modest and fail to survive Holm correction, suggesting that stronger reward shaping or additional training iterations may be needed for more robust gains. DAPO-DeepSeek exhibits no material degradation on MMLU or clean HotpotQA, indicating that the post-training procedure preserves general capabilities. These results demonstrate a directional improvement in selective evidence use, while identifying injection resistance and statistical robustness as important remaining challenges for future work.

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen3.5-4b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reinforcement-learni…] indexed:0 read:1min 2026-07-23 ·