04:00
2026-07-13
arxiv.org
artificial-intelligence
SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions
Researchers have developed SafeExplorer, an unbiased policy gradient modification for proximal policy optimization (PPO) that reduces training-time falls by up to 233x on physical robots while matchinβ¦