This work-in-progress paper embeds a differentiable physics model directly into the proximal policy optimization (PPO) actor loss function, by simulating short-horizon future trajectories during training, to reduce constraint violations while maintaining reliable target tracking.
Abstract
Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but its"black-box"exploration risks violating strict hardware safety limits. Typically, these constraints are managed through complex reward shaping. In this work-in-progress paper, we embed a differentiable physics model directly into the proximal policy optimization (PPO) actor loss function. By simulating short-horizon future trajectories during training, the policy is penalized for anticipated safety violations independent of the task-reward signal. Evaluated on a simulated 1-degree-of-freedom helicopter testbed with strict pitch constraints, our physics-informed soft regularizations substantially reduce constraint violations while maintaining reliable target tracking.
A generalized framework that combines the adaptive, high-performance nature of deep reinforcement learning (DRL) with the formal safety guarantees of model predictive control (MPC) is proposed, demonstrating successful exploration and stable policy convergence on physical hardware.
George Schafer, Jakob Rehrl, Stefan Huber et al.· 0 citations
Comparative control theory analysis reveals a fundamental trade-off: transitioning from continuous to discrete rule-based control induces high-frequency Bang-Bang actuation and a stable bimodal limit cycle.
This study proposes an explainable, data-driven framework integrating active-reward proximal policy optimization (AR-PPO), which successfully distills black-box AI strategies into verifiable, physics-informed standard operating procedures (SOPs), providing a highly transparent and robust solution for autonomous windshear escape and future competency-based flight training.
This paper presents an advanced deep reinforcement learning (DRL) framework for precise trajectory tracking control of an underactuated 2-degree-of-freedom (2-DOF) helicopter system using the twin delayed deep deterministic policy gradient (TD3) algorithm. The 2-DOF helicopter serves as a benchmark for nonlinear, coupled, and underactuated systems, posing significant challenges for conventional control approaches. Both classical linear and nonlinear control methods provide baseline solutions; however, their performance often degrades in the presence of parameter variations, uncertainties, and external disturbances. To overcome the severe value overestimation errors caused by aerodynamic cross-coupling in standard actor-critic architectures, a model-free TD3-based controller is developed, incorporating an artificial potential field-inspired reward function to simultaneously optimize tracking accuracy, energy efficiency, and control smoothness. Compared with standard DRL approaches such as the deep deterministic policy gradient (DDPG), the TD3 algorithm addresses key limitations by employing twin critics to reduce overestimation bias, delayed policy updates to improve training stability, and target policy smoothing to enhance robustness. Comprehensive simulations conducted in a MATLAB/Simulink environment demonstrate the superior performance of the proposed TD3 controller compared to classical and intelligent approaches, including proportional-integral-derivative (PID), fuzzy PD + I, and fuzzy PD + FF controllers. For multi-step trajectory tracking, TD3 reduces overshoot to 4.6% (pitch) and 3.8% (yaw) compared to 22.4% and 18.7% for PID, while decreasing settling time by up to 66%. The steady-state error is reduced to 0.18° (pitch) and 0.15° (yaw), representing improvements exceeding 80% over PID. In addition, TD3 minimizes cross-coupling effects by over 60%, enabling effective decoupled control of pitch and yaw dynamics. Under complex trajectories and disturbance conditions, including ± 10% parametric uncertainties and external torque disturbances, the TD3 controller consistently achieves the lowest tracking errors, fastest convergence, and smoothest control signals, reducing control variation by up to 65% compared to conventional methods. These results highlight the effectiveness of TD3 for controlling nonlinear and underactuated systems and provide a solid foundation for future experimental validation and real-world deployment in aerial robotic platforms operating in uncertain environments.
Zied Ben Hazem, Muhammed Özdemir, Firas Saidi et al.· Discover Robotics· 1 citation
A physics-aware, end-to-end deep reinforcement learning (DRL) approach that acts directly on low-level body inputs, total thrust and body torques, and closes the loop through a high-fidelity Simulink environment is investigated.
This work proposes an extension of the ATACOM framework, a state-of-the-art reliable safety layer that can be integrated with existing Reinforcement Learning algorithms to enforce constraints derived from prior knowledge of the system or learned directly from data.
Paolo Magliano, Puze Liu, Jan Peters et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.