{"slug": "cigpo-contextual-information-gain-policy-optimization-for-multi-turn-evidence", "title": "CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents", "summary": "Researchers at arXiv propose Contextual Information-Gain Policy Optimization (CIGPO) to fix a zero-advantage lock-in failure in outcome-only reinforcement learning for multi-turn evidence-reading agents. In HotpotQA experiments with Qwen2.5-3B-Instruct, GRPO initially improves (standard F1 0.430) but collapses to 100% format-violating outputs, while CIGPO reaches a standard F1 of 0.518 (from 0.252 base; +105%) by assigning per-turn rewards to intermediate evidence-reading turns. The study identifies reward-variance collapse as a concrete failure mode of outcome-only GRPO and shows that turn-level information-gain rewards can prevent it.", "body_md": "arXiv:2607.16244v1 Announce Type: new\nAbstract: Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little direct credit. In HotpotQA experiments with Qwen2.5-3B-Instruct, GRPO initially improves (standard F1 0.430) but subsequently collapses to 100% format-violating outputs. Training-log diagnosis reveals a zero-advantage lock-in mechanism: all sampled trajectories receive the minimum format penalty (-2.0), group-relative advantages vanish, and the policy-gradient loss becomes zero--an optimization deadlock. We propose a variance-injection strategy: by assigning per-turn rewards to intermediate evidence-reading turns, we prevent the group reward distribution from collapsing to a single value--preserving the variation that GRPO's group-relative advantage requires. Contextual Information-Gain Policy Optimization (CIGPO) implements this strategy using the marginal increase in the frozen reference model's log-likelihood of the ground-truth answer as the per-turn signal. With separate normalization of IG and F1 rewards and an IG-weight curriculum, CIGPO reaches a standard F1 of 0.518 on HotpotQA at the 3B scale (from 0.252 base; +105%), compared with 0.430 for the best GRPO checkpoint and 0.000 for the final GRPO checkpoint. CIGPO maintains meaningful reward variance and avoids zero-advantage lock-in throughout training. These results identify reward-variance collapse as a concrete failure mode of outcome-only GRPO and show that turn-level IG rewards can prevent it in this HotpotQA setting.", "url": "https://wpnews.pro/news/cigpo-contextual-information-gain-policy-optimization-for-multi-turn-evidence", "canonical_source": "https://arxiv.org/abs/2607.16244", "published_at": "2026-07-21 04:00:00+00:00", "updated_at": "2026-07-21 04:12:59.807434+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "machine-learning"], "entities": ["arXiv", "Qwen2.5-3B-Instruct", "HotpotQA", "Contextual Information-Gain Policy Optimization", "CIGPO", "GRPO"], "alternates": {"html": "https://wpnews.pro/news/cigpo-contextual-information-gain-policy-optimization-for-multi-turn-evidence", "markdown": "https://wpnews.pro/news/cigpo-contextual-information-gain-policy-optimization-for-multi-turn-evidence.md", "text": "https://wpnews.pro/news/cigpo-contextual-information-gain-policy-optimization-for-multi-turn-evidence.txt", "jsonld": "https://wpnews.pro/news/cigpo-contextual-information-gain-policy-optimization-for-multi-turn-evidence.jsonld"}}