cd /news/machine-learning/saf-opd-stable-advantage-fusion-for-… · home topics machine-learning article
[ARTICLE · art-84328] src=machinebrief.com ↗ pub= topic=machine-learning verified=true sentiment=↑ positive

SAF-OPD: Stable Advantage Fusion for On-Policy Distillation

Researchers propose SAF, a Stable Advantage Fusion framework that combines reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) for training language models, addressing magnitude and temporal mismatches in advantage fusion. Instantiating RLVR with GRPO, SAF avoids entropy collapse and outperforms fixed-coefficient fusion across seven mathematical reasoning and code generation benchmarks with Qwen3-1.7B/4B/8B, improving aggregate scores by 0.51-2.70% across all six model-domain settings.

read1 min views1 publishedAug 3, 2026

arXiv:2607.29209v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) broadcasts a single response-level reward to every token, while on-policy distillation (OPD) scores each token against a stronger teacher for a dense advantage but caps performance at teacher quality and discourages exploration beyond it. Their complementarity makes combining RLVR and OPD promising, but we find that fusing the two advantages with a fixed coefficient triggers entropy collapse from two miscalibrations: a magnitude mismatch, where token-level OPD advantages can spike far beyond the bounded RLVR advantage and erase its signal, and a temporal mismatch, where sustained full-strength OPD keeps pulling the student toward the teacher and limits exploration needed to surpass it. We propose SAF, a Stable Advantage Fusion framework that resolves both issues via a lightweight, four-stage pipeline applied only to the OPD advantage: a sparsify-then-compress mechanism for magnitude control paired with a warm-up-then-anneal mechanism for temporal control, with each stage independently switchable and adding negligible overhead. Instantiating RLVR with GRPO, we evaluate SAF across seven mathematical reasoning and code generation benchmarks with Qwen3-1.7B/4B/8B: SAF avoids entropy collapse and consistently outperforms fixed-coefficient GRPO+OPD fusion, improving the aggregate score by 0.51-2.70% across all six model-domain settings while achieving more stable training.

── more in #machine-learning 4 stories · sorted by recency
── more on @saf 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/saf-opd-stable-advan…] indexed:0 read:1min 2026-08-03 ·