arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models
Tropical Reinforcement Learning
A new arXiv paper (2610.02478v1) proposes Tropical Reinforcement Learning, which replaces the standard sum of trajectory probabilities with a maximum over the tropical semiring, and reports that its TROPIC training algorithm beats the strongest on-policy baselines by up to 16 percentage points on four agentic tasks: Sokoban, Countdown, FrozenLake and WebShop. The authors argue that expected return only reports how often a policy succeeds, not which solution worked, and that summing probabilities causes the model to forget alternative solutions that were never shown to be wrong. TROPIC targets deterministic, resettable environments with verifiable outcomes, where the best prefix and best suffix meeting at a shared state can be joined even when they come from different rollouts.
Run your AI side-project on zahid.host
EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.