Skip to content

3-D Trajectory Design Based on Deep Reinforcement Learning for UAV-Assisted Communication Networks

2026 · IEEE Transactions on Network Science and Engineering · Vol 13, pp. 10551-10571 · 0 citations · 121 references
Computer Science

Abstract

Most of the existing UAV-assisted communication networks provide service only for static users or deterministically moving ones. In fact, for some complex and dynamically changing scenarios, the users communicating to the UAV may move randomly, with unpredictable mobility. The uncertainty of users’ movements poses a challenge to the guarantee of stable network performance. To tackle this, the paper investigates a UAV-assisted communication network, where a UAV provides communication service for ground users which are moving randomly. We collectively factor in ground user mobility, task duration, and UAV flight restrictions to design precise 3D trajectory for UAV, and formulate them into an optimization problem, aiming to maximize the network throughput while minimizing UAV energy consumption. Considering the dynamics caused by users’ uncertain movement, we transform the optimization problem into a Markov decision process (MDP), then improve the twin-delayed deep deterministic policy gradient (TD3) to design UAV’s 3D trajectory. By utilizing the prior knowledge to accelerate the exploration efficiency, we propose a trajectory design algorithm based on prior knowledge-TD3 (PKTD3-TD), enabling UAV to autonomously adjust flight parameters by leveraging environmental observations under dynamic conditions for enhancing flexibility and intelligence. Simulation results show that our proposed scheme outperforms the compared ones in terms of communication link quality, network throughput and UAV’s energy consumption.

View source

Similar papers

2026

Optimization of 3-D Trajectory and Resource Allocation in Multi-UAV Communications Under a Probabilistic Channel Model

This study investigates the optimization of three-dimensional (3D) trajectory planning and resource allocation in unmanned aerial vehicle (UAV)-enabled wireless networks with no-fly zones (NFZs) using a deep learning framework. The objective is to maximize the minimum average spectral efficiency (SE) among mobile users served by multiple UAVs while addressing key challenges, including interference from concurrent UAV transmissions, collision avoidance, and NFZ constraints. A realistic probabilistic channel model is considered, where the likelihood of a line-of-sight (LoS) condition is modeled as a function of the elevation angle in the air-to-ground (A2G) link. To solve the formulated optimization problem, a novel deep learning framework with specialized deep neural network (DNN) structures is developed. This framework jointly optimizes 3D UAV trajectory planning and resource allocation, employing an unsupervised learning-based training approach that eliminates the need for labeled data. Performance evaluations demonstrate that the proposed scheme effectively accounts for the probabilistic channel model and co-channel interference while accounting for collision avoidance and NFZ-related constraints. Moreover, it outperforms baseline methods by achieving a higher minimum average SE with real-time computational efficiency, making it practical for UAV-assisted wireless networks.

Woongsup Lee, Howon Lee, Kisong Lee · 0 citations
Open access Aug 2026

Deep Reinforcement Learning-Based Joint Control for Rotatable-Array UAV Transportation Communications

Future transportation networks may require aerial communication platforms capable of providing flexible and reliable services to vehicular terminals. In conventional unmanned aerial vehicle (UAV) communication systems, the antenna geometry is commonly treated as fixed, which limits the attainable directional gain when the relative geometry between the UAV and users changes significantly. This work considered a UAV equipped with a mechanically reconfigurable antenna array and studied its joint motion and transmission control under finite-blocklength communication. A sequential optimization problem was formulated to maximize the accumulated user throughput by jointly optimizing the UAV trajectory, the array orientations, and the transmit beamforming vectors, subject to the UAV kinematic constraints, the UPA orientation constraints, and the transmission energy budget. The resulting problem involves nonlinear coupling among platform motion, antenna pointing, beamforming, and finite-blocklength rate expressions, making conventional optimization computationally demanding. To obtain an adaptive control policy, a soft actor–critic-based deep reinforcement learning method was developed. The simulation results showed that jointly controlling the UAV mobility, array orientation, and beamforming improves the achievable finite-blocklength transmission performance compared with benchmark schemes, demonstrating the effectiveness of the proposed framework in enhancing reliable data delivery for UAV-assisted transportation infrastructure applications.

Chen Zhang, Yi Xiong · 0 citations
Conference Jul 2026

Reinforcement Learning-Based Decode-and-Forward UAV Relay Trajectory Optimization

Unmanned Aerial Vehicles (UAVs) are promising relay platforms due to their flexible deployment and high probability of line-of-sight (LoS) connectivity. This paper compares three deep reinforcement learning (DRL) algorithms-Proximal Policy Optimization (PPO), Soft Actor-Critic (SAC), and Recurrent PPO with LSTM memory-for joint UAV trajectory and energy optimization in UAV based relay systems. The problem formulated is a non-convex optimization problem that minimizes UAV propulsion energy while satisfying Quality of Service (QoS) and mobility constraints under realistic 3GPP channel conditions. Simulation results show that all methods achieve over 99% QoS satisfaction. SAC exhibits the fastest convergence, whereas the proposed Recurrent PPO achieves the lowest energy consumption (44.72 kJ), reducing energy usage by 5.1% compared with PPO. These results highlight the trade-off between convergence speed and energy efficiency in DRL-based UAV relay optimization.

Aniket Subbanwar, Ojas Joshi, Amit Agarwal · 0 citations
Conference Jul 2026

Double Deep Reinforcement Learning–Based UAV Positioning for Throughput Optimization in Wireless Networks

This work investigates a reinforcement learning-based control framework for the autonomous movement and coordination of multiple Unmanned Aerial Vehicles (UAVs) in a wireless communication environment. The considered system includes UAVs performing sensing and relaying tasks, where mobility decisions directly affect the overall network performance. The main objective is to improve the communication quality of ground users by maximizing aggregate network throughput. To achieve this objective, a Double Deep Q-Network (DDQN) architecture is employed, where each UAV is assigned an individual learning agent. The agents learn role-specific movement policies while coordinating through interactions with the shared environment. Learning performance is further improved by using adaptive scaling and a custom reward function designed to capture variations in network utility. Simulation results show that the proposed approach outperforms baseline movement strategies in terms of utility. In addition, different task configurations, agent behaviors, and hyperparameter selections are examined to improve convergence speed and training stability. Overall, the results indicate that reinforcement learning is a promising method for cooperative UAV positioning in dynamic and interference-sensitive wireless communication scenarios.

Berke Kilinç, M. Ö. Efe · 0 citations
Jul 2026

CRB-Driven Beamforming and Trajectory Optimization for UAV-assisted ISAC System

Simulation results demonstrate that the proposed method significantly reduces the time-averaged CRB by over 10%, compared with the ISAC system without UAV assistance, and also achieves a higher sensing accuracy than both the fixed-UAV-trajectory and the maximum-ratio-transmission-based beamforming benchmarks.

Yi Yang, Qianqian Zhang, Huaxia Wang · 0 citations
Preprint Sep 2026

Enhancing UAV Trajectory and Communications Through Vision-Inertial Tracking

In this paper, we propose an energy-efficient and reliable communication system for non-terrestrial networks deployed in dynamic GPS-denied wireless environments, enabled by a Vision--Inertial Tracking-Assisted UAV Communication (VIT-UAVCom) system. To the best of our knowledge, this is the first work to exploit onboard UAV cameras and IMU sensors for UAV-assisted communications. We consider a complete VIT-UAVCom system that incorporates the key design parameters while explicitly accounting for system noise and residual tracking inaccuracies. Building on this framework, we formulate an optimization problem for jointly designing the UAV trajectory and communication performance to improve propulsion energy efficiency, reduce outage probability, and enhance physical-layer security. We then develop a dedicated solution framework to efficiently compute near-optimal trajectory and communication control actions in dynamic scenarios. Furthermore, to enable real-time implementation, we propose and evaluate three optimizers, namely linear search (LS), binary search (BS), and genetic search. Our numerical results demonstrate that our proposed VIT-UAVCom framework significantly outperforms the K-means benchmark in terms of energy consumption while maintaining robust secrecy performance and reliable user coverage. Specifically, our proposed framework improves the energy efficiency by 144% compared to the benchmark. Interestingly, our results also show that, compared with the LS, the BS reduces the computational time by approximately 50%.

Abdallah S. Ghazy, Hussein A. Ammar, James H. Bayes et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.