{"slug": "rethinking-critic-learning-in-ppo-understanding-and-mitigating-value-flattening", "title": "Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening", "summary": "Researchers identified a systematic failure mode in Proximal Policy Optimization (PPO) critics used for reinforcement learning of large language models, which they call Value Flattening, in which state values estimated from the critic collapse toward a narrow range. The finding concerns the critic component that estimates state values to reduce variance in policy updates, and the work proposes understanding and mitigating the flattening effect.", "body_md": "In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated fro", "url": "https://wpnews.pro/news/rethinking-critic-learning-in-ppo-understanding-and-mitigating-value-flattening", "canonical_source": "https://aiflash.com/news/121203/", "published_at": "2026-09-17 08:30:26+00:00", "updated_at": "2026-09-17 08:53:52.868550+00:00", "lang": "en", "topics": ["machine-learning", "large-language-models", "ai-research"], "entities": ["Proximal Policy Optimization", "PPO"], "alternates": {"html": "https://wpnews.pro/news/rethinking-critic-learning-in-ppo-understanding-and-mitigating-value-flattening", "markdown": "https://wpnews.pro/news/rethinking-critic-learning-in-ppo-understanding-and-mitigating-value-flattening.md", "text": "https://wpnews.pro/news/rethinking-critic-learning-in-ppo-understanding-and-mitigating-value-flattening.txt", "jsonld": "https://wpnews.pro/news/rethinking-critic-learning-in-ppo-understanding-and-mitigating-value-flattening.jsonld"}}