cd /news/artificial-intelligence/deepo-dual-entropy-enhanced-policy-o… · home › topics › artificial-intelligence › article
[ARTICLE · art-139443] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

DEEPO: Dual-Entropy Enhanced Policy Optimization for Hallucination in MLLMs

Researchers introduced Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage reinforcement learning method that reduces hallucination in multimodal large language models by combining semantic-entropy-triggered expert prefixes with advantage-sign-aware Renyi gradient preconditioning. DEEPO's two branches each improved over GRPO individually, and their interaction was statistically significant on VideoMMMU, the most complex long-horizon task in the evaluation suite, at +4.0 points (95% CI [1.1, 6.9]), with additive gains elsewhere. The method reduces hallucination while preserving accuracy and training stability, according to the arXiv:2609.28570v1 paper.

by read1 min views1 publishedSep 25, 2026

arXiv:2609.28570v1 Announce Type: new Abstract: Reinforcement learning (RL) is widely used to sharpen reasoning in multimodal large language models (MLLMs), yet its effect on hallucination is uneven. We trace this to two weak points in the \emph{correction chain} from reward to parameter update. At the rollout level, hard queries---those with high semantic entropy---frequently produce unanimously wrong sample groups, collapsing the group-relative advantage to zero exactly where hallucination risk is highest. At the optimization level, confident-but-wrong tokens are gradient-invisible: a categorical policy's expected score-gradient norm vanishes as its distribution sharpens, so the predictions that most need correction receive the weakest updates. We propose Dual-Entropy Enhanced Policy Optimization (DEEPO), a dual-stage enhancement combining signal variance regularization with gradient preconditioning: semantic-entropy-triggered expert prefixes inject grounded continuations on high-uncertainty queries, providing direct supervision and restoring advantage variance, while advantage-sign-aware Renyi preconditioning counteracts logit-level saturation so correction reaches confident errors in the operational confidence regime. Both branches improve over GRPO individually; their interaction is statistically significant on VideoMMMU---the most complex long-horizon task in our evaluation suite (+4.0$, 95% CI [1.1, 6.9])---and additive elsewhere. DEEPO reduces hallucination while preserving accuracy and training stability.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @deepo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/deepo-dual-entropy-e…] indexed:0 read:1min 2026-09-25 · —