Jul 2026· International Conference on Control, Decision and Information Technologies· pp. 813-818· 0 citations· 11 references
Abstract
Trajectory planning for unmanned aerial vehicles (UAVs) in dynamic and partially observable environments becomes more complex when extended from two-dimensional to three-dimensional navigation. Although Deep Reinforcement Learning (DRL) methods have shown strong performance in 2D scenarios, their application to 3D spaces requires redesigned observation models, action representations, and safety mechanisms. This paper extends a 2D DRL-based trajectory planning framework to 3D environments using Proximal Policy Optimization (PPO), Deep Q-Network (DQN) and Deep Deterministic Policy Gradient (DDPG). UAV agents are trained to reach randomly placed 3D targets while avoiding static and dynamic obstacles using only local sensory information. The observation space combines a local 3D occupancy representation with a relative 3D goal vector, preserving partial observability and avoiding reliance on a global map. This article proposes that the simulation results demonstrate robust, collision-aware navigation and improved safety and trajectory efficiency in each one of the DRL algorithms implemented, each having positive and negative specifics.
This work proposes a modified Multi-Agent Twin-Delayed Deep Deterministic Policy Gradient (M-MATD3) algorithm, specifically designed to mitigate common issues such as overestimation bias and high variance observed in standard MATD3.
This paper presents an observation-only autonomy framework for Unmanned Underwater Vehicles (UUVs) navigation in dynamic underwater environments that integrates persistent occupancy mapping, global clearance-aware planning, and risk-aware local control. The proposed pipeline constructs occupancy maps solely from onboard sonar and depth image observations, adapts a clearance-constrained global planner (GP) to provide long-horizon structure, and integrates a reinforcement learning (RL) policy to handle short-range tracking and reactive avoidance. To further support decision-making under partial observability, the system learns a compact latent state representation from onboard sensor data, encoding environmental structure, obstacle dynamics, and uncertainty. Behavior tree (BT) distillation with staged supervision is introduced to improve safety and training stability, while an uncertainty-calibrated distillation mechanism reweights teacher guidance using online latent-model uncertainty, emphasizing uncertain regimes during learning, with time-to-collision (TTC) and clearance cues remaining explicit in planning and local policy features. To demonstrate the efficacy of the framework, a reproducible multi-seed evaluation protocol is established in high-fidelity GPU-accelerated simulation using NVIDIA Isaac Sim, and performance is benchmarked against BT-only and standard RL baselines. The results obtained demonstrate improved robustness and safety under dynamic conditions, thus providing a general pipeline with a unified hybrid planning learning architecture and a reproducible methodology for robust UUV autonomy under partial observability.
This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.
Yuting Cao, Zheng Zhao, Jiekai Wu et al.· Journal of King Saud Univers...· 0 citations
Efficient 3-D path planning for autonomous underwater vehicles (AUVs) in dynamic submarine environments presents a significant challenge due to complex seabed terrain, ocean currents, and obstacles. In view of the adaptability and generalization limitations of traditional methods, this article proposes the reward-adaptive prioritized experience replay (RAPER) mechanism to dynamically adjust the experience sampling priorities by evaluating the temporal-difference errors and the weights of critical reward events, including success rate, goal proximity, route length, obstacle avoidance, and ocean current utilization. Then, a deep reinforcement learning framework, the DSAC-T-RAPER algorithm, is established for underactuated AUVs by combining the RAPER mechanism and distributional soft actor–critic with three refinements (DSAC-T) algorithm. Meanwhile, the finite-step evolution and boundedness of the adaptive event weight are theoretically analyzed to guarantee the reliability and stability of the proposed algorithm. The 3-D simulation environment is constructed by considering complicated submarine seafloor topography, real ocean current data from the Copernicus Marine Environment Monitoring Service, and randomly distributed obstacles. Simulations under different scenarios demonstrate that, compared with some other algorithms, DSAC-T-RAPER achieves higher success rates and better performances.
Xin Cheng, Hai Jin, Yun Chen et al.· IEEE Systems Journal· 0 citations
Autonomous navigation of unmanned aerial vehicles (UAVs) in constrained indoor environments remains a challenging problem due to limited maneuvering space and high collision risk. This paper presents an empirical evaluation of a reinforcement learning-based approach for UAV path planning using Proximal Policy Optimization (PPO) in a high-fidelity AirSim simulation environment. The UAV operates in a cluttered warehouse setting with static obstacles, where it must reach predefined goal locations while minimizing collisions and trajectory inefficiencies. We conducted extensive experiments to analyze the reward function shaping, training dynamics, hyperparameter sensitivity, and navigation performance across multiple goal configurations. Results show that the PPO-based agent achieves an overall success rate of (78%) with stable convergence and improved path efficiency, while maintaining low collision rates in less constrained regions. However, performance degrades in obstacledense areas requiring sharp maneuvers, highlighting limitations in generalization. The study provides insights into the impact of reward design and environment complexity on reinforcement learning performance for indoor UAV navigation.
Ashraf Suyyagh, Tasneem Al-Qat, Hala Mukheimer et al.· IEEE Jordan Conference on Ap...· 0 citations
The problems of partial observability and sensor shortage pose a significant challenge for autonomous Unmanned Aerial Vehicles (UAVs) as they prove to be challenging for conventional Deep Reinforcement Learning (DRL) methods to undertake well under such conditions. In this paper, a memory-augmented Proximal Policy Optimization (PPO) model extended using a Long Short-Term Memory (LSTM) network is proposed as a solution to such challenges. The observation space is constructed from 2D LiDAR and Inertial Measurement Unit (IMU) data to sense simultaneously external observation and internal state of motion, whereas the action space consists of continuous velocity commands. A shaped reward function is optimized for encouraging safe target approaching, obstacle avoidance, and convergence speed. Experimental outcomes show that the PPO-LSTM described herein achieves smoother paths, more robust reward convergence, and a much lower rate of collision than regular PPO. It also generalizes to new environments with movable obstacles. Qualitatively, the success rate increased from 64.5% to 83.9%, collision frequency reduced by over 70%, and path efficiency increased from 0.60 to 0.85, without suffering from unstable training behavior
M. Haddad, Dhayaa Khudher· Kufa journal of Engineering· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.