An optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) reinforcement learning framework for controlling nonlinear dynamical systems under hard safety constraints is presented and a model-free PG algorithm based on stochastic gradient ascent is developed.
Abstract
We present an optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) framework for controlling nonlinear dynamical systems under hard safety constraints. We first construct a class of stochastic wrapper policies centered around a deterministic controller, thereby enabling exploration in unknown environments while preserving the underlying deterministic control structure. We then define a class of parameterized safe-by-construction control policies by truncating these stochastic policies onto hard safety constraints. We next establish, via measure-theoretic arguments, that the potentially nonconvex RL objective under the truncated policy class, as well as its policy gradients, are well-defined. We then develop a model-free PG algorithm based on stochastic gradient ascent that directly searches over these truncated policies and leverage gradient dominance to establish convergence and optimality guarantees. Finally, we validate this framework through simulations on a safe quadrotor navigation problem.
This work introduces Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side.
Chenhua Fan, Jiahui Zhu, Yuhang Zhang et al.· 0 citations
Non-stationary environments pose significant challenges for reinforcement learning (RL), particularly in safety-critical applications like robotics and energy systems, where adaptability, stability, and robustness to uncertainty are essential. This paper introduces Lyapunov-Guided Distributional Reinforcement Learning...
M. Khaniki, Morteza Mirzaee, Elahe Moradi· Scientific Reports· 0 citations
Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective fr...
A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...
A. Choudhury, J. Brahmanage, Akshat Kumar et al.· Proceedings of the Thirty-Fi...· 0 citations
Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and...
We present an end-to-end QP-based policy framework that enables systematic policy construction with minimal domain-specific design, while preserving the transparency and interpretability of model-based control. The proposed policy representation supports both domain-randomized model-based auto-tuning, where policy para...