A model-free nested co-design framework for aeroelastic systems using deep reinforcement learning, in which a design-conditioned control policy is trained with proximal policy optimisation while an outer loop updates a distribution over candidate design parameters, highlights the potential of model-free co-design for complex aeroelastic systems in which design, control, and mission objectives are tightly coupled.
Abstract
Control co-design considers the physical system and its controller together, enabling the strong coupling between system design and control to be uncovered and exploited. This is especially relevant in aeroelastic flight systems, where structural, aerodynamic, and control design choices jointly determine manoeuvrability and efficiency. This paper presents a model-free nested co-design framework for aeroelastic systems using deep reinforcement learning, in which a design-conditioned control policy is trained with proximal policy optimisation while an outer loop updates a distribution over candidate design parameters. The approach is evaluated on three case studies of increasing complexity: a spring-mass-damper system, a pitch-plunge-flap aerofoil, and a highly flexible high-aspect-ratio glider performing a thermal-soaring mission in a stochastic environment. Across these case studies, the framework is shown to progressively concentrate the design search towards high-performing regions and to outperform policies trained on randomly sampled designs. The results also show that reward shaping plays an important role in enabling stable learning in partially observed and stochastic environments. In the final glider case, the method jointly addresses wing design, flight control, and mission-level behaviour in the presence of aeroelastic coupling and atmospheric uncertainty. These results highlight the potential of model-free co-design for complex aeroelastic systems in which design, control, and mission objectives are tightly coupled.
Abstract. This study introduces a DRL-based controller for multiple fixed-wing aircraft without model-specific tuning. The proposed system is trained using Proximal Policy Optimization (PPO) and is enhanced by algorithmic strategies such as multi-model training and kinematics-based reward shaping to ensure generalization across diverse platforms. It is compared with a conventional method to see if the generalized controller maintains high control performance at the level of an aircraft-specific design. This approach presents an adaptable alternative to conventional control design, offering resilience to configuration changes, and providing reduced dependency on the model.
Fatih Ahmet Sarigul· Materials Research Proceedin...· 0 citations
This work investigates Deep Reinforcement Learning (DRL) as a tool for model-free closed-loop active separation control in a fully turbulent wind tunnel flow over a one-sided diffuser. The agent controls an array of magnetic valves (on/off) that eject compressed air into the boundary layer, while the environmental state is reduced to the signal from a single wall-shear-stress sensor placed near the natural transitory detachment point. The control law is learned in real time using Proximal Policy Optimization. Compared to the standard learning design based on the weighted sum of all rewards following an action, we demonstrate that a horizon aligned with the convective time of the flow leads to faster convergence and a more robust control strategy. The resulting control law corresponds to a low-duty-cycle actuation pattern that yields a forward-flow fraction of approximately $53\%$. This compares favorably with conventional and optimized periodic open-loop control ($\sim 40\%$ and $\sim 51\%$, respectively). The findings of this article indicate that, when embedded into an online experiment, DRL represents an efficient tool to identify robust and interpretable active separation control strategies.
Sofia Avdiiv, Andre Weiner, B. Steinfurth· 0 citations
Experimental results in simulation and on real hardware demonstrate that the decentralized–supervised architecture can achieve comparable or improved aggregate tracking performance relative to a centralized policy, while preserving decentralized proposal generation and enabling execution-time supervisory coordination under partial observability.
A. Bozzi, Matteo Aicardi, E. Zero et al.· IEEE Access· 0 citations
A systematic assessment framework is presented that compares four prominent DRL controllers with a classical control baseline across a diverse set of applied control problems, including non-minimum phase dynamics, flexible mechanical systems, nonlinear marine control, and aerial robotics, and clarifies the trade-offs between learning-based and conventional control.
Klinsmann Agyei, Pouria Sarhadi, Daniel Polani· 0 citations
This article proposes a closed-loop active flow control framework based on proximal policy optimization (PPO) algorithm and a synthetic jet to suppress severe flow separation of EH1590 airfoil at high angles of attack. Aerodynamic characteristics of airfoils under different jet parameters are obtained through computational fluid dynamics (CFD) simulation, and a deep neural network surrogate model is trained to achieve rapid prediction. Based on this alternative model, the PPO algorithm is used to train the agent, taking the pressure distribution of the flow field around the airfoil as the state, to optimize and improve the weighted objectives of lift-to-drag ratio and jet energy consumption. Finally, the agent is tested in both fixed and variable angle of attack tasks, and its adaptability and control effectiveness in a higher-resolution CFD environment are verified by coupling with CFD. The results show that at a fixed angle of attack of 15°, the lift-to-drag ratio increased from 4.56 to 7.34. Under variable angles of attack, the agent can automatically adapt and maintain a high lift-to-drag ratio. The error between the lift-to-drag ratio calculated directly by coupling with CFD and the test results of the surrogate model is within 5%, which verifies the effectiveness of the training strategy based on the surrogate model in transferring to the real flow field. This study provides an efficient and feasible technical path for the reinforcement learning application of airfoil flow control under high Reynolds number and high angle of attack conditions.
Yuhang Su, Yanping Song, Fu Chen et al.· The Physics of Fluids· 0 citations
Deep reinforcement learning for autonomous unmanned aerial vehicle control has largely been demonstrated with multirotor platforms and high-level machine-learning frameworks. This study presents a Dueling Double Deep Q-Network (D3QN) training pipeline implemented in Rust without an external machine-learning library and integrated with Godot 4 through GDExtension for fixed wing flight control. The controller addresses fixed-wing requirements, including airspeed maintenance, lift management, throttle regulation, stall avoidance, and coordinated turning. The network combines a duelling architecture, double Q-learning, prioritised experience replay, and three-step returns in a 512→256 hidden-layer configuration containing 142,088 parameters for the 16-dimensional input case. The agent selects among seven discrete actions and supports both 12-dimensional and 16-dimensional observation spaces through a cross-dimensional weight-transfer procedure. Training was conducted for 200 episodes using four random seeds. Across seeds, the mean best episodic reward was 4825±40, while the coefficient of variation for best reward was 0.8%. In the final 30 episodes, no crashes were recorded, although completion rates varied substantially between seeds.
Airspeed remained within ±8 m/s of the 50 m/s target. Batch-64 gradient updates required less than 1 ms, representing an approximately 35-fold reduction in latency relative to the preceding GDScript implementation, and the reported runtime memory footprint remained below 50 MB. These findings support the feasibility of native Rust-based D3QN training for real-time fixed-wing simulation, while the observed inter-seed variability indicates that reward shaping and convergence robustness require further evaluation.
Saugat Chaudhary Tharu, Shrutika Ojha, Rija Bhomi et al.· Journal of Advances in Mathe...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.