arXiv:2610.06861v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with \emph{external guidance} - expert traces, self-explanations, or retrieved thought patterns. Although each method reports empirical gains, none provides convergence rates, bias bounds, or an optimal weighting rule for the guidance signal. We close this gap with \emph{Guidance-Augmented GRPO} (GA-GRPO), a unified theoretical framework that casts external guidance as a stochastic guidance operator G re-writing the question distribution, and analyses the resulting policy-gradient estimator as a biased on-policy estimator whose bias is bounded by the total-variation guidance divergence delta_G between the guidance-augmented sampling distribution and the policy's own distribution. The framework subsumes vanilla GRPO, LUFFY, ExPO, PAPO, and TAPO as special cases obtained by particular choices of G. Under smoothness and bounded-divergence assumptions we prove that GA-GRPO converges at rate O(1/sqrt(T)) to an O(delta sqrt(T))-neighbourhood of the GRPO stationary point, derive the closed-form MSE-optimal guidance weight lambda-star(T, delta, sigma_0 squared) = sigma_0 squared / (sigma_0 squared + R_max squared delta squared T), and prove a matching minimax lower bound showing the Omega(delta squared T) bias term is unavoidable. Experiments on Qwen2.5-Math-7B-Base across nine math and OOD benchmarks confirm that optimal-weight GA-GRPO matches or surpasses TAPO, LUFFY, ExPO, and vanilla GRPO while requiring 31% fewer GPU-hours, and eight analysis experiments validate each theoretical prediction.
When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
A new arXiv paper (2610.06861v1) introduces Guidance-Augmented GRPO (GA-GRPO), a theoretical framework that treats external guidance in reinforcement learning with verifiable rewards as a stochastic guidance operator G and derives a closed-form MSE-optimal guidance weight lambda-star(T, delta, sigma_0 squared) = sigma_0 squared / (sigma_0 squared + R_max squared delta squared T). The authors prove GA-GRPO converges at rate O(1/sqrt(T)) to an O(delta sqrt(T))-neighbourhood of the GRPO stationary point and establish a matching minimax lower bound showing the Omega(delta squared T) bias term is unavoidable, subsuming vanilla GRPO, LUFFY, ExPO, PAPO, and TAPO as special cases. Experiments on Qwen2.5-Math-7B-Base across nine math and out-of-distribution benchmarks show optimal-weight GA-GRPO matches or surpasses TAPO, LUFFY, ExPO, and vanilla GRPO while using 31% fewer GPU-hours, with eight analysis experiments validating each theoretical prediction.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.