04:00
2026-08-12
arxiv.org
machine-learning
Boundary-Seeking Policy Gradient for Safe Reinforcement Learning
Researchers introduced Boundary-Seeking Policy Gradient (BSPG), a first-order method for safe reinforcement learning that drives policies to the constraint boundary when constraints are active, achievβ¦