{"slug": "robust-peak-cost-constrained-reinforcement-learning", "title": "Robust Peak-cost Constrained Reinforcement Learning", "summary": "A new study on robust peak-cost constrained reinforcement learning (RP-CRL) reveals that peak-cost constrained MDPs may not admit zero duality gap, unlike standard CMDPs. The researchers propose a surrogate optimization framework and a robust value estimation method based on integral probability metrics to address simulator-to-real-world mismatch, proving that the surrogate solution achieves the same robust reward with at most epsilon constraint violation. Experiments demonstrate effective safety enforcement under dynamics perturbations while maintaining strong reward performance.", "body_md": "arXiv:2607.15457v1 Announce Type: new\nAbstract: We study robust peak-cost constrained reinforcement learning (RP-CRL), where the objective is to maximize expected reward while controlling the maximum cost encountered along a trajectory. This setting is motivated by safety-critical applications in which a single large violation can be catastrophic and therefore cannot be adequately captured by the standard CMDP framework based on expected cumulative cost. Existing reachability-constrained RL methods adopt Lagrangian-based approaches, yet the underlying duality properties of peak-cost constrained MDPs remain unclear. We show that, unlike standard CMDPs, peak-cost constrained MDPs may not admit zero duality gap. We further consider a robust formulation to address simulator-to-real-world mismatch in the transition dynamics. To solve this problem, we develop a surrogate optimization framework and a robust value estimation method based on integral probability metrics. We prove that, with appropriate hyperparameter choices, the surrogate solution attains the same robust reward value as the original problem while violating the constraint by at most epsilon. Experiments show that the proposed method effectively enforces safety under dynamics perturbations while retaining strong reward performance.", "url": "https://wpnews.pro/news/robust-peak-cost-constrained-reinforcement-learning", "canonical_source": "https://arxiv.org/abs/2607.15457", "published_at": "2026-07-20 04:00:00+00:00", "updated_at": "2026-07-20 14:07:27.473447+00:00", "lang": "en", "topics": ["ai-safety", "machine-learning"], "entities": [], "alternates": {"html": "https://wpnews.pro/news/robust-peak-cost-constrained-reinforcement-learning", "markdown": "https://wpnews.pro/news/robust-peak-cost-constrained-reinforcement-learning.md", "text": "https://wpnews.pro/news/robust-peak-cost-constrained-reinforcement-learning.txt", "jsonld": "https://wpnews.pro/news/robust-peak-cost-constrained-reinforcement-learning.jsonld"}}