cd /news/artificial-intelligence/gcpo-diagnosing-and-constraining-sub… · home topics artificial-intelligence article
[ARTICLE · art-94797] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

GCPO: Diagnosing and Constraining Subspace Geometry in Rollout RL for LLMs

Researchers introduced GCPO (Geometrically Constrained Policy Optimization), a method that applies hard bilateral orthogonal projections to constrain updates in rollout reinforcement learning for large language models, preventing performance degradation. On Qwen3-8B and GLM4-9B, GCPO outperformed GRPO and variants like DAPO and GSPO, improving over base models by up to 27.69 points and over the strongest baseline by 2.37 points, while eliminating response-length inflation and stabilizing policy entropy.

read1 min views1 publishedAug 13, 2026

arXiv:2608.11674v1 Announce Type: new Abstract: On-policy rollout methods such as GRPO are central to post-training of large language models, yet they frequently suffer from training instabilities, cross-task capability degradation, and response-length inflation. Although prior work has characterized the subspace geometry of aggregate updates, the stepwise variation of this geometry and its relationship to model performance remain unclear. We introduce Principal-Subspace Overlap, a dimension-corrected measure of individual rollout updates relative to the dominant singular subspaces of pretrained weights. Despite low average overlap, transient spikes often precede performance degradation. To address this, we propose GCPO (Geometrically Constrained Policy Optimization), which applies hard bilateral orthogonal projections to constrain updates to the complementary subspaces, preventing such excursions by construction. Across mathematical reasoning, code generation, and tool-use tasks on Qwen3-8B and GLM4-9B, GCPO consistently outperforms GRPO and recent variants, including DAPO and GSPO, improving over the base models and the strongest baseline by up to 27.69 and 2.37 points, respectively. Furthermore, GCPO preserves general capabilities, eliminates response-length inflation, and stabilizes policy entropy. Our findings provide a new diagnostic lens and a principled design perspective for stable reinforcement learning post-training.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @gcpo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gcpo-diagnosing-and-…] indexed:0 read:1min 2026-08-13 ·