A systematic assessment framework is presented that compares four prominent DRL controllers with a classical control baseline across a diverse set of applied control problems, including non-minimum phase dynamics, flexible mechanical systems, nonlinear marine control, and aerial robotics, and clarifies the trade-offs between learning-based and conventional control.
A generalized framework that combines the adaptive, high-performance nature of deep reinforcement learning (DRL) with the formal safety guarantees of model predictive control (MPC) is proposed, demonstrating successful exploration and stable policy convergence on physical hardware.
George Schafer, Jakob Rehrl, Stefan Huber et al.· 0 citations
Advanced control of heating and cooling systems can substantially reduce energy costs and pollution. However, real-world adoption of popular algorithms among researchers, such as model predictive control (MPC) and reinforcement learning (RL), remains limited due in part to their high deployment and commissioning costs. Here, we develop two nearly commissioning-free controllers tailored to objectives that depend linearly on the controlled thermal load, such as energy costs and pollution. The controllers require at most two thermal parameters. In representative heating simulations, controller performance is robust to large parameter specification errors, suggesting potential for deployment with no tuning. The controllers maintain good occupant comfort while achieving 43 to 98% (depending on the electricity pricing and controller variant) of the performance improvement achieved by an omniscient policy with perfect model information and forecasts. These results suggest that simple, structure-exploiting controllers may capture most of the attainable value of advanced control while avoiding the data, modeling, tuning, and computational burdens that can arise with conventional MPC or RL.
W. G. Dierking, Arash Khabbazi, Levi D. Reyes Premer et al.· arXiv.org· 0 citations
This work begins with dynamic modeling using the Euler-Lagrange formulation and demonstrates that the proposed hybrid controller outperforms traditional methods in terms of tracking accuracy, settling time, and disturbance rejection.
Anita Verma· International Journal of Int...· 0 citations
PEARL employs an actor-adjoint algorithm that leverages automatic differentiation to compute policy gradients over short horizons and adjoint-based sensitivities of future returns approximated via neural networks, significantly reducing the number of environment interactions, while mitigating long-term gradient instabilities.