{"slug": "convex-concave-reinforcement-learning", "title": "Convex-Concave Reinforcement Learning", "summary": "A new arXiv paper (2610.09108v1) shows that the exact per-iteration policy-learning objective in reinforcement learning, written in log-density-ratio coordinates y := log[π/π_n] and computed via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex program, allowing it to be solved directly rather than through surrogate approximations. The authors solve the per-iteration program with sequential convex programming (SCP), recover CPI, NPG, TRPO and AWR as special cases, and add a multi-step axis k coupling consecutive decisions. Their multi-step Convex-Concave RL (CCRL) converges markedly faster than tuned PPO on a stochastic mid-horizon healthcare domain, reaching the same near-optimal survival with 11.3% higher area under the training curve, and is competitive with tuned PPO on classic control.", "body_md": "arXiv:2610.09108v1 Announce Type: new \nAbstract: Policy learning drives many of the most consequential and heavily-invested applications of reinforcement learning today. Yet the core optimization problem it rests on (maximizing expected return) is notoriously non-convex, even under a direct policy parameterization, and the field has largely responded by avoiding it: optimizing convex surrogate approximations of the return under trust-region constraints (NPG, TRPO, PPO, AWR). We show that this seemingly unstructured problem is not actually structureless. In log-density-ratio coordinates $y := \\log[\\pi/\\pi_n]$, the exact per-iteration objective, computable via per-decision importance sampling (PDIS), is a difference-of-convex-constrained difference-of-convex (DC-constrained DC) program. This structure lets us move beyond surrogate approximations: it recovers CPI, NPG, TRPO, and AWR as special cases along interpretable axes, and it opens a multi-step axis $k$ that couples consecutive decisions. We solve the per-iteration program with sequential convex programming (SCP), the standard solver for difference-of-convex problems, and give convergence guarantees under mild conditions, bridging the difference-of-convex optimization and RL literatures. Empirically, multi-step Convex-Concave RL (CCRL) wins on diagnostic MDPs where credit must propagate across a horizon (its advantage growing with the dependency length), is competitive with a tuned PPO on classic control, and on a realistic, stochastic, mid-horizon healthcare domain converges markedly faster than tuned PPO to the same near-optimal survival, with an 11.3% higher area under the training curve.", "url": "https://wpnews.pro/news/convex-concave-reinforcement-learning", "canonical_source": "https://www.machinebrief.com/news/convex-concave-reinforcement-learning-dukr", "published_at": "2026-10-08 04:00:00+00:00", "updated_at": "2026-10-08 05:16:45.586433+00:00", "lang": "en", "topics": ["machine-learning", "ai-research", "artificial-intelligence"], "entities": ["Convex-Concave Reinforcement Learning", "CCRL", "PPO", "TRPO", "NPG", "AWR", "CPI", "arXiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/convex-concave-reinforcement-learning", "markdown": "https://wpnews.pro/news/convex-concave-reinforcement-learning.md", "text": "https://wpnews.pro/news/convex-concave-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/convex-concave-reinforcement-learning.jsonld"}}