Skip to content
Preprint

SafePG: Safe and Globally Optimal Reinforcement Learning with Hard Constraints

Sep 2026 · 0 citations · 21 references
Engineering Computer Science

TL;DR

An optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) reinforcement learning framework for controlling nonlinear dynamical systems under hard safety constraints is presented and a model-free PG algorithm based on stochastic gradient ascent is developed.

Abstract

We present an optimal and convergent model-free policy gradient (PG) reinforcement learning (RL) framework for controlling nonlinear dynamical systems under hard safety constraints. We first construct a class of stochastic wrapper policies centered around a deterministic controller, thereby enabling exploration in unknown environments while preserving the underlying deterministic control structure. We then define a class of parameterized safe-by-construction control policies by truncating these stochastic policies onto hard safety constraints. We next establish, via measure-theoretic arguments, that the potentially nonconvex RL objective under the truncated policy class, as well as its policy gradients, are well-defined. We then develop a model-free PG algorithm based on stochastic gradient ascent that directly searches over these truncated policies and leverage gradient dominance to establish convergence and optimality guarantees. Finally, we validate this framework through simulations on a safe quadrotor navigation problem.

View source

Similar papers

Preprint Aug 2026

Boundary-Seeking Policy Gradient for Safe Reinforcement Learning

This work introduces Boundary-Seeking Policy Gradient (BSPG), a first-order method whose update combines a tangential component that improves reward while preserving cost to first order with a signed, residual-driven normal component that regulates the policy toward the active boundary from either side.

Chenhua Fan, Jiahui Zhu, Yuhang Zhang et al. · 0 citations
Open access Aug 2026

Safe and adaptive control of non-stationary stochastic systems via Lyapunov-constrained distributional reinforcement learning

Non-stationary environments pose significant challenges for reinforcement learning (RL), particularly in safety-critical applications like robotics and energy systems, where adaptability, stability, and robustness to uncertainty are essential. This paper introduces Lyapunov-Guided Distributional Reinforcement Learning...

M. Khaniki, Morteza Mirzaee, Elahe Moradi · 0 citations
#artificial intelligence Preprint Sep 2026

Sufficiency of Zeroth-Order Reward Shaping for Policy Gradient in Stabilization Control

Reward shaping is fundamental to modern robotic control with deep reinforcement learning (RL), yet practitioners still rely heavily on heuristic principles borrowed from classical optimal control and trajectory optimization. Existing methods rarely distinguish reward terms that are intrinsic to the control objective fr...

Yi-Sheng Zhang, Tao Wang, Si-Cun Gao · 0 citations
Conference Open access Sep 2026

Persistent Safety Set Guided Offline Safe Reinforcement Learning

A framework for learning control barrier functions (CBFs) using a novel generalized Bellman operator is developed, yielding a persistent safety set from which the agent can remain safe indefinitely, and a new reward maximization algorithm is proposed that effectively exploits the learned persistent safety set for rewar...

A. Choudhury, J. Brahmanage, Akshat Kumar et al. · 0 citations
Preprint Sep 2026

Parallel Policy-Gradient Methods for Parameter Optimization of Nonlinear Feedback Controllers

Structured feedback controllers provide rigorous stability guarantees, but often require manual parameter tuning to achieve good closed-loop performance. Policy-gradient methods offer a systematic approach to parameter optimization; however, conventional gradient evaluation requires sequential forward state rollout and...

A. Nguyen, Leilei Cui · 0 citations
Preprint Sep 2026

End-to-end QP-based policies: A unified perspective on robust control and robot learning

We present an end-to-end QP-based policy framework that enables systematic policy construction with minimal domain-specific design, while preserving the transparency and interpretability of model-based control. The proposed policy representation supports both domain-randomized model-based auto-tuning, where policy para...

Fausto Vega, Priyanka Supraja Balaji, Chase Dunaway et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.