Convex-Concave Reinforcement Learning
A new arXiv paper (2610.09108v1) shows that the exact per-iteration policy-learning objective in reinforcement learning, written in log-density-ratio coordinates y := log[π/π_n] and computed via per-d…
A new arXiv paper (2610.09108v1) shows that the exact per-iteration policy-learning objective in reinforcement learning, written in log-density-ratio coordinates y := log[π/π_n] and computed via per-d…
A paper on arXiv (2610.08819v1) presents HydroSphere, a governed, data-driven framework for real-time water quality monitoring, forecasting, treatment optimization and fault recovery, evaluated on 2.8…
Developer manogyasingh released transformer-visualisation, a browser-based walkthrough of a toy GPT-2-like model that computes every matrix live and shows the exact math for each cell on hover, coveri…
An open-source project called ClashRoyaleAi published a browser demo in which a neural network with a few thousand weights, trained with the REINFORCE policy-gradient method, learns to place a single …
Researchers introduced Code to Control, an approach that synthesizes Python controllers executing directly as policies, separating program structure from parameters so an LLM builds the controller str…
A new arXiv paper (2609.35897v1) presents the first causal mechanistic audit of a self-discovered reinforcement learning rule, Disco103, which previously surpassed PPO to reach state-of-the-art benchm…
A developer's attempt to train a reinforcement-learning agent to complete the 1983 ZX Spectrum game Manic Miner failed, first with PPO over joystick inputs and then with move macros, according to a fi…
A developer published a practical guide to GRPO (Group Relative Policy Optimization), an RL method for fine-tuning and aligning large language models that avoids the separate critic/value model requir…
Researchers introduced REFINEPPO, a reinforcement learning method that combines Iterative Action Refinement (IAR) with Proximal Policy Optimization (PPO), building control actions through a sequence o…
A new arXiv paper (2609.21327v1) introduces Deep-BQRL, a model-free distributional reinforcement-learning framework that extends buffered-quantile learning to neural function approximation, learning c…
A developer detailed how frontier LLM development is shifting from pre-training scaling to test-time compute scaling, outlining three architectural regimes: sequential chain-of-thought expansion, leaf…
Researchers identified a systematic failure mode in Proximal Policy Optimization (PPO) critics used for reinforcement learning of large language models, which they call Value Flattening, in which stat…
An open-source developer published an arXiv paper (2503.23303), a Hugging Face model (DeepMostInnovations/sales-conversion-model-reinf-learning), and a training dataset (DeepMostInnovations/saas-sales…
A benchmark of seven reinforcement-learning trainer algorithms on Qwen3-14B in a decentralized exchange (DEX) arbitrage gym found that CISPO posted the highest training reward curve in thinking mode (…
Long-horizon AI agents fail at an exponential rate because a 95% per-step success probability yields only a 35.8% chance of completing 20 turns and 7.7% at 50 steps, and the chance of a second error j…
A developer explains the evolution of large language model training from supervised fine-tuning to reinforcement learning from human feedback and reinforcement learning with verifiable rewards, noting…
A technical analysis of multi-reward reinforcement learning for LLM agents finds that summing multiple reward channels into standard algorithms like PPO or vanilla GRPO causes scale dominance, where t…
A new study from arXiv (2609.04880v1) demonstrates that reinforcement learning can design sequential solar PV incentive policies, balancing adoption and cost under uncertainty. Using PPO, SAC, and TD3…
A new study from arXiv (2609.00718v1) finds that structured pruning is the first stage where driving capability is lost in compressed autonomous driving policies, while distillation improves the prune…
Pollen Robotics released microduck_rl on GitHub, a set of reinforcement learning training environments for its 25 cm, 800 g bipedal robot Microduck. The environments, built on mjlab (MuJoCo Warp) and …