Skip to content
Open access

A forecast-guided reinforcement learning approach for trajectory planning of unmanned aerial base stations

Sep 2026 · Scientific Reports · Vol 16 · 0 citations · 34 references

TL;DR

The proposed Forecast-SAC framework demonstrates that unified predictive-control learning enables safe and efficient UAV-BS navigation under dynamic uncertainty, achieving a strong safety–throughput balance that reactive methods cannot match in high-risk environments.

Abstract

The rapid growth of wireless devices and the emergence of dynamic traffic hotspots have increased the need for intelligent trajectory planning in unmanned aerial vehicle base stations (UAV-BSs) operating in complex urban environments, where conventional reactive methods relying only on current system states cannot anticipate near-future changes in demand, obstacle risk, and communication conditions, often resulting in inefficient coverage, higher energy consumption, and unsafe flight behavior. To address this limitation, this paper proposes Forecast-SAC, a forecasting-enhanced deep reinforcement learning framework for proactive 3D UAV-BS trajectory planning that integrates a CNN-LSTM forecasting module with a Soft Actor-Critic (SAC) controller, where the forecasting module learns spatiotemporal demand and obstacle-risk patterns from historical map sequences and the SAC agent uses both predicted and current states to generate continuous control actions under joint service, safety, and energy constraints. The framework is evaluated in a grid-based urban simulation environment with mobile Gaussian hotspot demand, spatial obstacle-risk fields, queue-based service accumulation, and an air-to-ground channel model dependent on elevation angle, while training is performed using sequence-based forecasting and environment interaction under collision-avoidance and energy constraints. Experimental results show stable convergence and demonstrate that Forecast-SAC achieves an average served traffic of 4439.26 ± 532.96 Mbit, energy efficiency of 0.0623 Mbit/J under a revised rotary-wing energy model, average risk of 2.85, a baseline fairness index of 0.229 (a limitation addressed by the proposed fairness-enhancement term, which raises the index to 0.318–0.401; Sect. 4.7), and battery retention of 96.23%, while maintaining low-latency inference (3.07 ± 0.21 ms) compatible with real-time control loops; deployment readiness beyond simulation would additionally require hardware-in-the-loop validation. Comparative results indicate that Forecast-SAC produces smoother and safer trajectories than reactive baselines while maintaining competitive throughput. Ablation studies further show that replacing SAC with heuristic forecast-driven rules increases risk by 57–157%, confirming the necessity of learned continuous control, while PPO and DDPG baselines without forecasting are outperformed by 25.7–41.3% in episodic return, validating the contribution of spatiotemporal prediction. Sensitivity analysis over the violation penalty (λv ∈ {100–1000}) identifies 700 as a near-optimal tradeoff point, achieving 85.7% fewer violations than the least-penalized setting with only a small throughput loss. Multi-step forecasting analysis shows that single-step (t + 1) prediction is adopted as the primary operating point for its favourable accuracy–latency tradeoff, while the two-step (t + 1 + t + 3) variant yields a marginal return improvement at increased inference cost, and longer horizons (t + 5) degrade performance due to error accumulation, and a fairness-oriented variant increases the fairness index to 0.318–0.401. The proposed framework demonstrates that unified predictive-control learning enables safe and efficient UAV-BS navigation under dynamic uncertainty, achieving a strong safety–throughput balance that reactive methods cannot match in high-risk environments.

Read PDF

Similar papers

Open access Aug 2026

Physics-guided residual learning for phase-aware UAV trajectory prediction in urban environments

Results indicate that combining physical structure with learned residual correction provides a more accurate, physically consistent, and operationally interpretable approach for UAV trajectory forecasting.

Md Ashraful Islam, Stanley Förster, Tianxiong Zhang et al. · 0 citations
Open access 2026

A Hybrid Diffusion World Model for UAV Trajectory Forecasting and Collision Risk Estimation in Dense 3D Environments

: Autonomous unmanned aerial vehicles (UAVs) operating in dense three-dimensional environments require predictive models that can represent multiple plausible futures while estimating the safety consequences of these futures. This paper presents a hybrid diffusion world model for short-horizon UAV trajectory forecastin...

B. Nguyen, Ngan Nguyen Xuan Phuong · 0 citations
Open access Aug 2026

A Simulation-Based Dynamic Path Planning Approach for Low-Altitude Unmanned Aerial Vehicles in Inspection Scenarios

A dynamic path planning method for low-altitude Unmanned Aerial Vehicles (UAVs) tailored for urban inspection missions and constrains the average response latency for high-priority emergency tasks to within 40 s even under 50 concurrent dynamic tasks is proposed.

Changqi Yang, Hongjie Hu, Yi Ai · 0 citations
Sep 2026

Dynamic unmanned aerial vehicle path planning for rescue missions using trajectory-predictive and attention-enhanced deep recurrent SARSA

This paper proposes a trajectory-predictive and attention-enhanced deep recurrent State-Action-Reward-State-Action (SARSA) algorithm (PA-DR-SARSA), which incorporates future motion information into on-policy reinforcement learning through a Gated Recurrent Unit–based trajectory prediction module that learns temporal mo...

Shan Dan, Meng Zhang, Dong-Ming Liu et al. · 0 citations
Conference Aug 2026

A Multimodal Deep Reinforcement Learning Framework for Autonomous UAV Navigation in Gazebo–ROS Environments

Unmanned Aerial Vehicle (UAV) autonomous navigation is a key capability for UAVs, particularly when in complex and dynamic environments where continuous human control is impractical. While commonly used rule-based navigation and path-planning techniques can be effective in structured environments, they can be ineffecti...

B. S. Santhoshi, B. S, R. Shankar · 0 citations
Conference Open access 2026

UAV Path Planning and Multi-Aircraft Collaboration Based on Machine Learning

Unmanned Aerial Vehicle applications are expanding into dense urban airspace, and path planning needs to meet the needs of safe, efficient, and autonomous flight in complex dynamic environments. Traditional path planning methods have high computational complexity, poor real-time performance, and are difficult to deal w...

Xingran Du, Xi Hong, Min-Rui Li · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.