{"slug": "gcpo-diagnosing-and-constraining-subspace-geometry-in-rollout-rl-for-llms", "title": "GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs", "summary": "Researchers introduced GCPO (Geometrically Constrained Policy Optimization), a method that applies hard bilateral orthogonal projections to constrain updates in rollout reinforcement learning for large language models, preventing performance degradation. On Qwen3-8B and GLM4-9B, GCPO outperformed GRPO and variants like DAPO and GSPO, improving over base models by up to 27.69 points and over the strongest baseline by 2.37 points, while eliminating response-length inflation and stabilizing policy entropy.", "body_md": "arXiv:2608.11674v1 Announce Type: new\nAbstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior work has characterized the subspace geometry of aggregate updates, the stepwise variation of this geometry and its relationship to model performance remain unclear. We introduce Principal-Subspace Overlap, a dimension-corrected measure of individual rollout updates relative to the dominant singular subspaces of pretrained weights. Despite low average overlap, transient spikes often precede performance degradation. To address this, we propose GCPO (Geometrically Constrained Policy Optimization), which applies hard bilateral orthogonal projections to constrain updates to the complementary subspaces, preventing such excursions by construction. Across mathematical reasoning, code generation, and tool-use tasks on Qwen3-8B and GLM4-9B, GCPO consistently outperforms GRPO and recent variants, including DAPO and GSPO, improving over the base models and the strongest baseline by up to 27.69 and 2.37 points, respectively. Furthermore, GCPO preserves general capabilities, eliminates response-length inflation, and stabilizes policy entropy. Our findings provide a new diagnostic lens and a principled design perspective for stable reinforcement learning post-training.", "url": "https://wpnews.pro/news/gcpo-diagnosing-and-constraining-subspace-geometry-in-rollout-rl-for-llms", "canonical_source": "https://www.machinebrief.com/news/gcpo-diagnosing-and-constraining-subspace-geometry-in-rollou-qmax", "published_at": "2026-08-13 04:00:00+00:00", "updated_at": "2026-08-13 04:42:02.933457+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "large-language-models", "ai-research"], "entities": ["GCPO", "GRPO", "DAPO", "GSPO", "Qwen3-8B", "GLM4-9B", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/gcpo-diagnosing-and-constraining-subspace-geometry-in-rollout-rl-for-llms", "markdown": "https://wpnews.pro/news/gcpo-diagnosing-and-constraining-subspace-geometry-in-rollout-rl-for-llms.md", "text": "https://wpnews.pro/news/gcpo-diagnosing-and-constraining-subspace-geometry-in-rollout-rl-for-llms.txt", "jsonld": "https://wpnews.pro/news/gcpo-diagnosing-and-constraining-subspace-geometry-in-rollout-rl-for-llms.jsonld"}}