Jul 2026· ACM Transactions on Autonomous and Adaptive Systems· 0 citations· 73 references
TL;DR
A counterexample-guided reinforcement learning method that navigates safe exploration in autonomous systems without prior knowledge, even when safety and optimality conflict, and a novel belief-based regularization method to address the distributional shift between online and offline learning and to balance optimization and safety.
Abstract
Safe exploration in reinforcement learning remains a critical challenge for safety-critical autonomous systems, where the typical trial-and-error learning process can lead to hazardous outcomes. While several existing approaches incorporate kinematic models or external knowledge to limit the exploration of unsafe behaviors, their effectiveness is significantly weakened in the presence of incomplete or sparse knowledge. This paper introduces a counterexample-guided reinforcement learning method that navigates safe exploration in autonomous systems without prior knowledge, even when safety and optimality conflict. Our method geometrically abstracts discrete and continuous state-space systems into compact, PAC-learnable models that capture safety-relevant information. We then generate probabilistic counterexamples of the safety requirement to regulate online exploration toward minimizing safety violations, relying on minimal offline counterexample-guided simulations. We further propose a novel belief-based regularization method to address the distributional shift between online and offline learning and to balance optimization and safety, ensuring conservative behavior with theoretical guarantees. Our evaluations demonstrate the effectiveness of the method in significantly reducing safety violations without compromising cumulative rewards when benchmarked against other Q-learning or actor-critic methods with unconstrained or safety-constrained exploration.
This work proposes an extension of the ATACOM framework, a state-of-the-art reliable safety layer that can be integrated with existing Reinforcement Learning algorithms to enforce constraints derived from prior knowledge of the system or learned directly from data.
Paolo Magliano, Puze Liu, Jan Peters et al.· arXiv.org· 0 citations
Safe model-based reinforcement learning (RL) often bridges control-theoretic analysis and RL for robots to safely explore (partially) unknown system dynamics while deriving control actions for task efficiency. The control performance and safety assurance typically rely on prior knowledge of partially modeled nominal system dynamics and the data-driven models that compensate for residual model uncertainties. However, existing methods often overlook the structure of residual model uncertainties (e.g., components affine in control), which could lead to overly conservative robot behaviors or invalid safety guarantees under the safe learning-based controllers. This paper proposes a safe reinforcement learning framework that learns control-affine dynamics with a certifiable data-driven safe policy using control barrier functions (CBF). Specifically, we first use Control-Affine Random Fourier Features (ARFF) to model robot dynamics in a control-affine form, which offers computational efficiency that scales with dataset size and reduces potential model bias for model-based reinforcement learning. Then, a model-free, efficient uncertainty quantification method using adaptive conformal prediction (ACP) is applied to quantify the uncertainty in the safety constraint arising from the learned control-affine dynamics. This allows for data-driven safety assurance amenable to principled and efficient controller synthesis with CBF. Simulation results on the cartpole and the 3D quadrotor platforms demonstrate the effectiveness of the proposed framework.
A generalized framework that combines the adaptive, high-performance nature of deep reinforcement learning (DRL) with the formal safety guarantees of model predictive control (MPC) is proposed, demonstrating successful exploration and stable policy convergence on physical hardware.
George Schafer, Jakob Rehrl, Stefan Huber et al.· 0 citations
DynBudget, a closed-loop Safe RL framework integrating a learned safety critic, temperature-calibrated risk estimation, and a dynamic safety budget, is proposed and it is shown that shielding with dynamic budgets is an interpretable and viable approach to Safe RL in autonomous systems.
It is shown that fuzzing-generated crashes can meaningfully improve agent robustness and enable accurate safety monitoring with strong cross-method generalization, and the benefits of combining complementary fuzzing strategies and adopting multi-level diversity analysis to achieve more comprehensive and practical RL testing.
Zhibin Kang, Hanmo You, Dong Wang et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.