An improved DRL algorithm, Dropout-based Prioritized Soft Actor-Critic (DPSAC), which integrates the Soft Actor-Critic (SAC) algorithm with Prioritized Experience Replay (PER) and the Dropout technique is proposed, and two innovative approaches are introduced to enhance the algorithm’s performance.
Abstract
In uncrewed aerial vehicle (UAV)-assisted Internet of Things (IoT) networks, UAVs often need to fly at low altitudes and navigate through obstacles to maintain reliable communication with IoT nodes during data collection missions. This paper proposes a novel deep reinforcement learning (DRL)-based approach for 3D UAV path planning and obstacle avoidance, with the objective of minimizing data collection time from IoT nodes distributed across complex urban environments. To address this problem, we propose an improved DRL algorithm, Dropout-based Prioritized Soft Actor-Critic (DPSAC), which integrates the Soft Actor-Critic (SAC) algorithm with Prioritized Experience Replay (PER) and the Dropout technique. Furthermore, two innovative approaches are introduced to enhance the algorithm’s performance. First, the Episodic Environment (EN) training approach introduces random variations in obstacle number, position, and height across training episodes, thereby enhancing the agent’s ability to generalize its learned policy to new and unknown environments. Second, the Switching Reward mechanism reduces penalties for collisions and boundary violations in the reward function during the early stages of training, thereby facilitating exploration and accelerating the agent’s learning of IoT-related tasks. Simulation results demonstrate that the proposed DRL-based approach achieves faster convergence and greater stability during the training process compared to baseline algorithms. Specifically, experiments conducted in new and complex environments show that this method can collect data with an average collision-free success rate of 98% from 10 IoT nodes and 95% from 20 IoT nodes, confirming its remarkable superiority over the baseline algorithms.
Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.
Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al.· arXiv.org· 0 citations
To address the challenges of three-dimensional (3D) flight path planning for Unmanned Aerial Vehicles (UAVs) in complex urban environments, this paper proposes a reinforcement learning approach based on the Deep Q-Network (DQN) algorithm. The method enables intelligent flight path planning within a discretized 3D urban space, dynamically avoiding obstacles in real-time through the UAV's sensory perception. The UAV agent is trained in a simulated 100×100×20 virtual urban environment, with training scenarios categorized into high, medium, and low difficulty levels to progressively enhance the agent's decision-making capabilities. Throughout the training process, a greedy strategy is adopted to balance the exploration of new potential paths and the exploitation of known optimal routes. Once over 80% of the UAV agents successfully reach their designated target points, the training program automatically advances to the next difficulty level. Experimental results validate the effectiveness of the proposed method, demonstrating its superior obstacle avoidance capabilities and exceptional energy optimization performance in complex urban settings.
Yang Li, Xinjie Qian, Yanxiu Wang et al.· International Conference on...· 0 citations
To address the challenge of rapid and precise obstacle avoidance for unmanned aerial vehicles (UAVs) in complex urban environments, rugged canyons, and other unstructured environments, this paper proposes a vision-based navigation algorithm. By combining the strengths of deep reinforcement learning (DRL) and convolutional neural networks (CNNs), this algorithm enables efficient navigation and obstacle avoidance in dynamic environments. First, to improve training efficiency, an autoencoder is used to extract latent spatial vectors from depth images, which are then used as input features for DRL. Second, an artificial potential field (APF) is introduced into the reward function to enhance obstacle avoidance performance in dynamic environments. Third, a CNN-based adaptive mode-switching mechanism is designed to meet navigation requirements under different environmental conditions. This mechanism can automatically identify environmental features based on real-time input data and dynamically adjust the UAV’s navigation strategy. To evaluate the proposed method, simulation experiments were conducted in static and dynamic scenarios, together with a preliminary indoor flight test. Under the evaluated conditions, the proposed method achieved favorable navigation success rates and path efficiency compared with the selected visual DRL baselines. The results also indicate cross-scenario transferability to the tested environments without environment-specific retraining.
Dongliang Wang, Yong-Qiang Jin, Weicheng Luo et al.· Italian National Conference...· 0 citations
Urban fire rescue poses severe challenges to the real-time performance and obstacle avoidance capabilities of unmanned aerial vehicle (UAV) path planning. Existing methods (such as A*, RRT, and standard DQN) have problems such as low search efficiency, insufficient obstacle avoidance ability, or slow convergence in complex environments. This paper proposes an improved deep Q-network (DQN) algorithm, introducing a priority experience replay mechanism to improve sample utilization, and designing a composite reward function including arrival reward, step penalty, direction guidance, and safety penalty to guide the UAV to plan safe and efficient flight paths in complex urban environments. A threedimensional grid simulation environment was constructed based on the real fire incident at Chongqing California Garden. Experimental results show that the improved DQN algorithm outperforms the traditional DQN and RRT algorithms. This method provides a feasible technical solution for multi-UAV collaborative rescue in urban fire scenarios.
Rui Qin, Han-Jin Zhou· International Conference on...· 0 citations
In recent years, Unmanned Aerial Vehicles (UAVs) have gradually been widely used in various fields such as regional search and disaster relief, and the development of related technologies has also experienced unprecedented growth. Compared to individual UAVs, the collaborative execution of tasks by UAV swarms has more advantages, but it is often difficult to achieve fast and accurate autonomous navigation and obstacle avoidance capabilities in unknown complex obstacle environments due to limitations in computing load and inter-UAV communication capabilities. To solve this problem, a hierarchical navigation decision-making framework for non-communication UAVs in unknown environments is proposed in this paper, which decomposes the UAV navigation planning task into an upper-layer global planning module and a lower-layer autonomous navigation and obstacle avoidance module. For the lower-layer module, an enhanced hybrid feature extraction network is designed, accompanied by a dual-stage training strategy that integrates traditional optimization methods with reinforcement learning. The upper-layer module incorporates a hybrid control strategy combining conventional search methods. Based on the ROS framework, simulation experiments for UAV swarm navigation were systematically conducted. The experimental results demonstrate that the proposed algorithm achieves autonomous navigation decision-making for multiple UAVs in complex unknown obstacle environments without relying on inter-UAV communication, showing significant advantages compared to existing approaches.
This work provides a feasible technical pathway and reproducible evaluation benchmark for the collaborative deployment of lightweight LLM planner, sub-goal guidance, sensor observations, cooperative reward, and reward shaping components and quantifies the indispensability of the LLM planner.
Yuting Cao, Zheng Zhao, Jiekai Wu et al.· Journal of King Saud Univers...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.