Single Controller for Multiple Aircraft by Deep Reinforcement Learning
Abstract. This study introduces a DRL-based controller for multiple fixed-wing aircraft without model-specific tuning. The proposed system is trained using Proximal Policy Optimization (PPO) and is enhanced by algorithmic strategies such as multi-model training and kinematics-based reward shaping to ensure generalization across diverse platforms. It is compared with a conventional method to see if the generalized controller maintains high control performance at the level of an aircraft-specific design. This approach presents an adaptable alternative to conventional control design, offering resilience to configuration changes, and providing reduced dependency on the model.