Preprint
Sep 2026
SafePG: Safe and Globally Optimal Reinforcement Learning with Hard Constraints
An optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) reinforcement learning framework for controlling nonlinear dynamical systems under hard safety constraints is presented and a model-free PG algorithm based on stochastic gradient ascent is developed.
Vipul K. Sharma, Wesley A. Suttle, S. Sivaranjani
· 0 citations