cd /news/artificial-intelligence/planpo-group-planning-aware-policy-o… · home topics artificial-intelligence article
[ARTICLE · art-102431] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

Researchers propose PlanPO, a group planning-aware policy optimization method that improves multi-turn agentic large language models by distinguishing advantages among successful trajectories based on interaction efficiency. PlanPO outperforms GRPO by 27.2% on average across ALFWorld, WebShop, and SciWorld benchmarks with negligible additional training cost.

read1 min views2 publishedAug 19, 2026

arXiv:2608.17289v1 Announce Type: new Abstract: Group-relative policy optimization has emerged as a key paradigm for training agentic large language models (LLMs) on multi-turn interactive tasks. However, most existing variants fail to distinguish advantages among successful trajectories even when these trajectories differ substantially in their interaction efficiency. For instance, circuitous successes are often assigned the identical outcome reward, causing advantage collapse and severe performance bottlenecks. To this end, we propose Group Planning-aware Policy Optimization (PlanPO), a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns. Specifically, PlanPO introduces coarse-to-fine advantage signals, which capture the relative differences in trajectory-level lengths and turn-level response lengths conditioned on successful trajectories sampled for the same task. Within the group-relative optimization structure, this enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generation from high-quality rollouts, without degenerating into vanilla length minimization. Experimentally, PlanPO improves over GRPO by 27.2% on average across the challenging multi-turn benchmarks ALFWorld, WebShop, and SciWorld, outperforming recent powerful baselines while incurring negligible additional training cost.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @planpo 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/planpo-group-plannin…] indexed:0 read:1min 2026-08-19 ·