A Lyapunov-Guided Post-Action Shield for Stability-Aware Deep Reinforcement Learning
This paper proposes a Lyapunov-guided post-action shielding mechanism for deep reinforcement learning (DRL) controllers under bounded actuation disturbances. In addition, an energy-safety requirement is formulated as a one-step energy threshold constraint that keeps the predicted next-state energy proxy within a prescribed limit, using a bounded-disturbance worst-case check. Simulation results show that the proposed mechanism substantially reduces constraint violations under actuation noise while preserving the nominal policy behavior whenever possible.