Jul 2026· 2026 IEEE International Workshop on Metrology for Living Environment (MetroLivEnv)· pp. 101-106· 0 citations· 16 references
Abstract
Unmanned Aerial Vehicles (UAVs) have gained widespread attention in diverse applications like military, medical, aerial surveillance and many more. Presently, the problem of limited bandwidth and geographic factors has raised the need for effective and timely data transfer. Training UAVs with reinforcement learning-based algorithms facilitates autonomous decision-making capabilities. In this paper, we proposed an intelligent system for the optimal UAV selection process by evaluating the continuous performance of each UAV. The analyzing factors are based on the real-world factors affecting the quality of signals, such as noise interference, relative motion between source and wave, and transmission power. Based on the systematic conditions observed, the system provides efficient rewards. To promote the selection of the optimal UAV and enhance the learning process, the state information of the UAV is fed into a deep neural network (DQN), which predicts the 'Q-values'. Our system implements a deep Q-learning algorithm, which enhances the agent's performance by systematically learning from its experience. The model operates accurately by selecting the most reliable UAV, thus, enhancing the throughput by optimal power allocation. It outperforms other conventional models in terms of timely data delivery and energy utilization. The system adapts various complex patterns by analyzing the historical and present scenarios. Empowered by this intelligent system, time-critical decision-making can be achieved with minimal energy consumption.
This work investigates a reinforcement learning-based control framework for the autonomous movement and coordination of multiple Unmanned Aerial Vehicles (UAVs) in a wireless communication environment. The considered system includes UAVs performing sensing and relaying tasks, where mobility decisions directly affect the overall network performance. The main objective is to improve the communication quality of ground users by maximizing aggregate network throughput. To achieve this objective, a Double Deep Q-Network (DDQN) architecture is employed, where each UAV is assigned an individual learning agent. The agents learn role-specific movement policies while coordinating through interactions with the shared environment. Learning performance is further improved by using adaptive scaling and a custom reward function designed to capture variations in network utility. Simulation results show that the proposed approach outperforms baseline movement strategies in terms of utility. In addition, different task configurations, agent behaviors, and hyperparameter selections are examined to improve convergence speed and training stability. Overall, the results indicate that reinforcement learning is a promising method for cooperative UAV positioning in dynamic and interference-sensitive wireless communication scenarios.
Berke Kilinç, M. Ö. Efe· International Conference on...· 0 citations
Future transportation networks may require aerial communication platforms capable of providing flexible and reliable services to vehicular terminals. In conventional unmanned aerial vehicle (UAV) communication systems, the antenna geometry is commonly treated as fixed, which limits the attainable directional gain when the relative geometry between the UAV and users changes significantly. This work considered a UAV equipped with a mechanically reconfigurable antenna array and studied its joint motion and transmission control under finite-blocklength communication. A sequential optimization problem was formulated to maximize the accumulated user throughput by jointly optimizing the UAV trajectory, the array orientations, and the transmit beamforming vectors, subject to the UAV kinematic constraints, the UPA orientation constraints, and the transmission energy budget. The resulting problem involves nonlinear coupling among platform motion, antenna pointing, beamforming, and finite-blocklength rate expressions, making conventional optimization computationally demanding. To obtain an adaptive control policy, a soft actor–critic-based deep reinforcement learning method was developed. The simulation results showed that jointly controlling the UAV mobility, array orientation, and beamforming improves the achievable finite-blocklength transmission performance compared with benchmark schemes, demonstrating the effectiveness of the proposed framework in enhancing reliable data delivery for UAV-assisted transportation infrastructure applications.
Chen Zhang, Yi Xiong· Infrastructures· 0 citations
Solar-powered Unmanned Aerial Vehicles (UAVs) and High-Altitude Pseudo-Satellites (HAPS) offer significant potential for persistent intelligence, surveillance and reconnaissance operations, but their endurance remains constrained by variable solar irradiance, atmospheric turbulence, battery limitations and payload power demand. This review examines the role of Artificial Intelligence (AI) and Machine Learning (ML) in improving UAV decision-support optimization through predictive solar forecasting, reinforcement-learning-based flight-path optimization, adaptive Maximum Power Point Tracking (MPPT), battery state estimation and intelligent payload load balancing. The paper synthesizes current approaches using Long Short-Term Memory networks, gradient-boosting models, convolutional neural networks, graph neural networks and reinforcement learning architectures for autonomous energy-aware flight. Simulation-based analyses suggest that AI-assisted control can improve energy utilization by 15-25% and extend mission persistence by 30-40%, particularly under variable irradiance and wind conditions. However, practical deployment requires robust validation, certifiable AI architectures, adversarial resilience and reliable edge-computing implementation. The strategic deployment of AI-enabled autonomous UAVs directly supports Saudi Arabia's Vision 2030 objectives for indigenous defence technology development, artificial intelligence research, and advanced aerospace systems. By establishing domestic expertise in AI-driven decision-support optimization, the Kingdom advances its defence industrial sovereignty while creating high-value technical employment in autonomous systems engineering and aerospace artificial intelligence.
Sulaman Rafiq· Journal of Intelligent Decis...· 0 citations
This paper constructs a reinforcement learning framework based on the PPO algorithm for drone air combat to solve 1v1 pursuit-evasion in 2D beyond-visual-range air combat. Firstly, the mission scenario is modeled, defining key roles of ATA and AA. Then, state transition models of pursuer and evader are built based on flight kinematics. To handle reward sparsity in policy network training, a dense reward function combining distance and angle rewards is designed to guide the agent in learning tail-chasing and interception strategies. Using the Actor-Critic architecture, deep neural networks implement the decision-making and evaluation modules. The PPO algorithm trains the pursuing drone in a simulation. Results show that after ~5 million steps, the agent learns a stable strategy, completing tasks promptly and generalizing well in unseen scenarios. This research offers ideas for drone combat and guidance, and supports autonomous decision-making in complex air battles.
Kangjie Yu, Zheng Gong, Runchang Hu et al.· SAE technical paper series· 0 citations
Unmanned Aerial Vehicles (UAVs) are promising relay platforms due to their flexible deployment and high probability of line-of-sight (LoS) connectivity. This paper compares three deep reinforcement learning (DRL) algorithms-Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Recurrent PPO with LSTM memory-for joint UAV trajectory and energy optimization in UAV based relay systems. The problem formulated is a non-convex optimization problem that minimizes UAV propulsion energy while satisfying Quality of Service (QoS) and mobility constraints under realistic 3GPP channel conditions. Simulation results show that all methods achieve over 99% QoS satisfaction. SAC exhibits the fastest convergence, whereas the proposed Recurrent PPO achieves the lowest energy consumption (44.72 kJ), reducing energy usage by 5.1% compared with PPO. These results highlight the trade-off between convergence speed and energy efficiency in DRL-based UAV relay optimization.
Aniket Subbanwar, Ojas Joshi, Amit Agarwal· International Conference on...· 0 citations
Simulation results confirm the effectiveness of distributed optimization and DRL-based coordination for scalable, resilient, and adaptable UAV deployment in disaster response and other mission-critical scenarios.
A. Abdellatif, Amr E. Aboeleneen, Mohamed M. Abdallah et al.· IEEE Open Journal of the Com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.