cd /news/machine-learning/reversal-bench-a-reversibility-axis-… · home topics machine-learning article
[ARTICLE · art-132229] src=arxiv.org ↗ pub= topic=machine-learning verified=true sentiment=· neutral

REVERSAL-BENCH: A Reversibility Axis and Reset Oracle for Measuring the Reset-Free RL Cliff

Researchers introduced REVERSAL-BENCH, a benchmark that controls environmental reversibility via a continuous parameter ρ ∈ [0, 1] and provides a reset oracle to test state recoverability across eight manipulation settings in five physics engines. Evaluating standard actor-critic algorithms, safe RL, and specialized reset-free frameworks revealed a sharp reversibility cliff: reset-free agents are consistently absorbed into irrecoverable states as ρ increases, while episodic agents maintain steady learning, a failure mode confirmed to be causally driven by irreversibility rather than obstacle complexity. The team released the benchmark suite, a large multi-simulator dataset labeled with recoverability, and the reset oracle, and found that a safety shield can predict recoverability accurately but active recovery primarily succeeds only when the agent can physically steer clear of the trap.

by read1 min views1 publishedSep 17, 2026

arXiv:2609.17745v1 Announce Type: new Abstract: A central goal of autonomous reinforcement learning is continuous policy training without external resets. However, existing paradigms largely depend on underlying environmental reversibility, a property absent in real world manipulation, where events such as pushing objects off tables or spilling granular substances cannot be undone. We introduce REVERSAL-BENCH, a benchmark that controls reversibility via a continuous parameter $\rho \in [0, 1]$ and provides a reset oracle, a ground-truth verification mechanism to test state recoverability across eight manipulation settings in five physics engines. Evaluating a broad spectrum of policy architectures, including standard actor-critic algorithms, safe RL, and specialized reset-free frameworks, reveals a sharp reversibility cliff: reset-free agents are consistently absorbed into irrecoverable states as $\rho$ increases, whereas episodic agents maintain steady learning. We see this failure mode across autonomous reset-free baselines and constrained RL. Because reset-free agents lack external resets, any transition into an irrecoverable state results in permanent absorption, leaving the agent trapped where further learning halts. We show that this absorption phenomenon persists in full physics simulations under learned manipulation policies. By evaluating against geometrically identical reversible counterparts, we confirm that this breakdown is causally driven by irreversibility rather than obstacle complexity. We release the benchmark suite, a large multi-simulator dataset labeled with recoverability and a reset oracle. We also evaluate a safety shield that intervenes before irreversible failures occur, showing that while recoverability can be predicted accurately, active recovery primarily succeeds only when the agent can physically steer clear of the trap

── more in #machine-learning 4 stories · sorted by recency
── more on @reversal-bench 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/reversal-bench-a-rev…] indexed:0 read:1min 2026-09-17 ·