{"slug": "slpo-scaling-latent-reasoning-via-a-surrogate-policy", "title": "SLPO: Scaling Latent Reasoning via a Surrogate Policy", "summary": "Researchers introduce Surrogate Latent Policy Optimization (SLPO), a method that brings outcome-reward reinforcement learning to autoregressive latent reasoners, enabling test-time scaling in latent reasoning without decoding intermediate steps as language tokens. SLPO improves Pass@k under parallel sampling and allocates longer latent computation to harder instances with higher deterministic accuracy, addressing the limitation that latent reasoners previously remained imitation-bound while explicit Chain-of-Thought reasoners advanced via outcome-reward RL.", "body_md": "arXiv:2607.19691v1 Announce Type: new\nAbstract: Reinforcement learning with verifiable rewards has become the predominant recipe for eliciting test-time scaling in explicit Chain-of-Thought reasoners. Yet this scaling path remains computationally costly, since every intermediate step must be decoded as a language token. Latent reasoning instead carries intermediate computation as continuous vectors and already matches or surpasses explicit CoT at far shorter horizons. Despite this promise, latent reasoners remain largely imitation-bound, while explicit CoT has already moved past imitation via outcome-reward RL. Latent trajectories lack a tractable per-step likelihood and an adaptive stopping interface under fixed thinking budgets, so outcome rewards cannot elicit latent test-time scaling. We introduce Surrogate Latent Policy Optimization (SLPO) to bring outcome-reward RL to autoregressive latent reasoners: an empirical surrogate policy density over latent transitions for trajectory-level credit assignment, and a correctness-supervised stopping head that outcome-reward optimization refines into a variable-horizon policy. Across continuous and soft thinking settings, SLPO improves Pass@$k$ under parallel sampling and allocates longer latent computation to harder instances with higher deterministic accuracy.", "url": "https://wpnews.pro/news/slpo-scaling-latent-reasoning-via-a-surrogate-policy", "canonical_source": "https://www.machinebrief.com/news/slpo-scaling-latent-reasoning-via-a-surrogate-policy-i9ym", "published_at": "2026-07-23 04:00:00+00:00", "updated_at": "2026-07-23 04:04:31.286033+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "ai-research"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/slpo-scaling-latent-reasoning-via-a-surrogate-policy", "markdown": "https://wpnews.pro/news/slpo-scaling-latent-reasoning-via-a-surrogate-policy.md", "text": "https://wpnews.pro/news/slpo-scaling-latent-reasoning-via-a-surrogate-policy.txt", "jsonld": "https://wpnews.pro/news/slpo-scaling-latent-reasoning-via-a-surrogate-policy.jsonld"}}