cd /news/artificial-intelligence/tropical-reinforcement-learning · home › topics › artificial-intelligence › article
[ARTICLE · art-145131] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Tropical Reinforcement Learning

A new arXiv paper (2610.02478v1) proposes Tropical Reinforcement Learning, which replaces the standard sum of trajectory probabilities with a maximum over the tropical semiring, and reports that its TROPIC training algorithm beats the strongest on-policy baselines by up to 16 percentage points on four agentic tasks: Sokoban, Countdown, FrozenLake and WebShop. The authors argue that expected return only reports how often a policy succeeds, not which solution worked, and that summing probabilities causes the model to forget alternative solutions that were never shown to be wrong. TROPIC targets deterministic, resettable environments with verifiable outcomes, where the best prefix and best suffix meeting at a shared state can be joined even when they come from different rollouts.

by read1 min views20 publishedOct 5, 2026

arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compositional reasoning, where a solution must be assembled from reasoning steps that the model produces in separate, often failed, attempts but rarely produces together. To address this, we propose Tropical Reinforcement Learning, which rests on a simple change of algebra: instead of adding the probabilities of alternative solutions, we take their maximum, which yields the tropical semiring. The value of a state then becomes the log-probability of its most likely verified solution, together with an explicit path that can be replayed and reused. This enables true composition, since the best prefix and the best suffix meeting at a shared state can be joined even when they come from different rollouts. To put this into practice, we introduce TROPIC, a training algorithm for deterministic, resettable environments with verifiable outcomes. On four agentic tasks (Sokoban, Countdown, FrozenLake, WebShop), TROPIC outperforms the strongest on-policy baselines by up to 16 percentage points. Changing the algebra of reinforcement learning, not just its estimators, can thus substantially improve compositional reasoning in language models

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @tropical reinforcement learning 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/tropical-reinforceme…] indexed:0 read:1min 2026-10-05 · —