Skip to content
Preprint

Integrating Physics-Informed Neural Networks for Safe Reinforcement Learning in a 1-DoF Helicopter System

Jul 2026 · 0 citations · 11 references
Computer Science

TL;DR

This work-in-progress paper embeds a differentiable physics model directly into the proximal policy optimization (PPO) actor loss function, by simulating short-horizon future trajectories during training, to reduce constraint violations while maintaining reliable target tracking.

Abstract

Deep reinforcement learning (DRL) offers powerful control for industrial cyber-physical systems (ICPSs), but its"black-box"exploration risks violating strict hardware safety limits. Typically, these constraints are managed through complex reward shaping. In this work-in-progress paper, we embed a differentiable physics model directly into the proximal policy optimization (PPO) actor loss function. By simulating short-horizon future trajectories during training, the policy is penalized for anticipated safety violations independent of the task-reward signal. Evaluated on a simulated 1-degree-of-freedom helicopter testbed with strict pitch constraints, our physics-informed soft regularizations substantially reduce constraint violations while maintaining reliable target tracking.

View source

Similar papers

Preprint Jul 2026

Safe Reinforcement Learning using Ideas from Model Predictive Control

A generalized framework that combines the adaptive, high-performance nature of deep reinforcement learning (DRL) with the formal safety guarantees of model predictive control (MPC) is proposed, demonstrating successful exploration and stable policy convergence on physical hardware.

George Schafer, Jakob Rehrl, Stefan Huber et al. · 0 citations
Preprint Jul 2026

Explainable Reinforcement Learning via Physics-Aware Policy Distillation

Comparative control theory analysis reveals a fundamental trade-off: transitioning from continuous to discrete rule-based control induces high-frequency Bang-Bang actuation and a stable bimodal limit cycle.

Shaker Al-Tamari, Waled Kadour · 0 citations
Open access Aug 2026

Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation

This study proposes an explainable, data-driven framework integrating active-reward proximal policy optimization (AR-PPO), which successfully distills black-box AI strategies into verifiable, physics-informed standard operating procedures (SOPs), providing a highly transparent and robust solution for autonomous windshear escape and future competency-based flight training.

Yitan Wang, Yangyang Zhang, Zhenxing Gao · 0 citations
Open access Jul 2026

Deep reinforcement learning for accurate trajectory tracking in 2-DOF helicopter dynamics using twin delayed DDPG

This paper presents an advanced deep reinforcement learning (DRL) framework for precise trajectory tracking control of an underactuated 2-degree-of-freedom (2-DOF) helicopter system using the twin delayed deep deterministic policy gradient (TD3) algorithm. The 2-DOF helicopter serves as a benchmark for nonlinear, coupled, and underactuated systems, posing significant challenges for conventional control approaches. Both classical linear and nonlinear control methods provide baseline solutions; however, their performance often degrades in the presence of parameter variations, uncertainties, and external disturbances. To overcome the severe value overestimation errors caused by aerodynamic cross-coupling in standard actor-critic architectures, a model-free TD3-based controller is developed, incorporating an artificial potential field-inspired reward function to simultaneously optimize tracking accuracy, energy efficiency, and control smoothness. Compared with standard DRL approaches such as the deep deterministic policy gradient (DDPG), the TD3 algorithm addresses key limitations by employing twin critics to reduce overestimation bias, delayed policy updates to improve training stability, and target policy smoothing to enhance robustness. Comprehensive simulations conducted in a MATLAB/Simulink environment demonstrate the superior performance of the proposed TD3 controller compared to classical and intelligent approaches, including proportional-integral-derivative (PID), fuzzy PD + I, and fuzzy PD + FF controllers. For multi-step trajectory tracking, TD3 reduces overshoot to 4.6% (pitch) and 3.8% (yaw) compared to 22.4% and 18.7% for PID, while decreasing settling time by up to 66%. The steady-state error is reduced to 0.18° (pitch) and 0.15° (yaw), representing improvements exceeding 80% over PID. In addition, TD3 minimizes cross-coupling effects by over 60%, enabling effective decoupled control of pitch and yaw dynamics. Under complex trajectories and disturbance conditions, including ± 10% parametric uncertainties and external torque disturbances, the TD3 controller consistently achieves the lowest tracking errors, fastest convergence, and smoothest control signals, reducing control variation by up to 65% compared to conventional methods. These results highlight the effectiveness of TD3 for controlling nonlinear and underactuated systems and provide a solid foundation for future experimental validation and real-world deployment in aerial robotic platforms operating in uncertain environments.

Zied Ben Hazem, Muhammed Özdemir, Firas Saidi et al. · 1 citation
Jul 2026

Directional Constraints for Efficient Exploration in Safe Reinforcement Learning

This work proposes an extension of the ATACOM framework, a state-of-the-art reliable safety layer that can be integrated with existing Reinforcement Learning algorithms to enforce constraints derived from prior knowledge of the system or learned directly from data.

Paolo Magliano, Puze Liu, Jan Peters et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.