04:00
2026-07-23
arxiv.org
artificial-intelligence
From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation
Researchers propose Prefix-GRPO, a reinforcement learning framework that improves small language model agents by decomposing teacher trajectories into replay-aligned prefix queries and online continuaβ¦