Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.
This paper investigates the problem of cooperative multiple unmanned aerial vehicles (UAVs) data collection for Internet of Things (IoT) networks in dense urban environments. Unlike existing studies that predominantly rely on idealized spatial models and average-based probabilistic channel models, this work explicitly accounts for realistic 3-D building distributions and deterministically models ground-to-air (G2A) channel blockages. We formulate a joint optimization problem to minimize the total task completion time, subject to stringent system throughput, flight dynamics, and energy constraints. To tackle the highly coupled challenges of node scheduling and trajectory planning, we propose a lightweight two-stage heuristic strategy for dynamic access control, along with a multi-agent reinforcement learning for trajectory planning. Crucially, to overcome the severe sparse-reward bottleneck inherent in complex 3-D obstacle avoidance, we introduce a Pheromone-based Reward Shaping (PRS) mechanism. By mathematically integrating the UAV’s kinematic state with deterministic environmental feedback, PRS effectively transforms the sparse-reward navigation challenge into a dense and smooth gradient, thereby profoundly accelerating policy convergence. Extensive simulations demonstrate that the proposed MATD3-PRS framework significantly outperforms representative baselines, achieving superior performance in task completion time, flight trajectory efficiency, and overall energy saving.
Haitao Chen, Xinfeng Deng, Zhe Wang et al.· IEEE Transactions on Cogniti...· 0 citations
A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.
M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al.· 0 citations
Unmanned aerial vehicles (UAVs) are emerging as critical enablers of next generation Internet of Things (IoT) infrastructures, supporting real-time data collection, wireless relaying, and agile operations in dynamic environments. However, achieving safe and energy efficient multi UAV navigation under sixth generation (6G) communication constraints remains a significant challenge due to dynamic obstacles, limited on board energy, and the high cost of centralized coordination. This study introduces a multi agent soft actor-critic (MASAC) framework for UAV path planning and energy aware coordination in a fixed RIS assisted IoT grid. MASAC integrates entropy regularized actor-critic learning with reconfigurable intelligent surface (RIS) aware reward shaping to support energy aware navigation, RIS assisted recharging, and connectivity guided trajectory optimization. A lightweight convolutional policy network is used to encode spatial information from the grid environment, including obstacle locations, dynamic obstacle states, exploration memory, and UAV position, enabling efficient policy learning under constrained navigation settings. Extensive simulations in RIS assisted, 6G enabled IoT environments demonstrate that MASAC achieves a 100% mission success rate, where mission success is defined as reaching the fixed goal cell before energy depletion and within the maximum episode horizon of 500 steps. Compared with the strongest baseline success rate of 75%, this corresponds to a 25%-point absolute improvement and a 33.3% relative improvement under the same evaluation protocol and identical environmental settings. Within the adopted grid level energy abstraction, MASAC also achieves approximately 33% higher RIS recharge utilization. It also provides 6% greater grid level 6G connectivity and 23% higher cumulative reward. Meanwhile, it maintains a low simulation time evaluation latency of approximately 38 ms per UAV. Statistical analysis confirms these gains as significant ([Formula: see text]). The proposed framework offers a simulation level benchmark for energy efficient UAV navigation in RIS assisted IoT environments. It also supports future deployment oriented research under realistic operational constraints.
Md. Najmul Mowla, D. Asadi, Khaled M. Rabie et al.· Scientific Reports· 0 citations
Unmanned Aerial Vehicles (UAVs) are increasingly deployed as embodied aerial agents in low-altitude economies, forming mobile aerial edge networks that enable flexible computation offloading for vehicles. However, their limited endurance and frequent join/leave behaviours result in highly dynamic topologies, undermining long-term resource availability. Moreover, existing vehicle-centric task scheduling strategies cause resource contention and decision complexity in dense environments. To address these challenges, this paper proposes a hierarchical and scalable reinforcement learning-based scheduling framework (SkySched). In SkySched, UAVs collaboratively make deployment and task scheduling decisions. The framework consists of two tightly coupled modules. First, an adaptive UAV deployment module introduces a capability encoding mechanism that compresses heterogeneous UAV attributes into a unified one-dimensional capability index. This compact representation enables a Scalable Proximal Policy Optimization (SPPO) algorithm to efficiently coordinate UAV positioning, maximizing task coverage and sustaining network-wide computing availability under dynamic topology variations. Second, a hierarchical task scheduling module is designed, where K-means-based Roadside Unit (RSU) clustering enables vertical task offloading, while a SPPO-driven horizontal UAV-to-UAV task redistribution mechanism achieves fine-grained load balancing across the UAV swarm. Simulations demonstrate that SkySched consistently outperforms state-of-the-art methods in terms of task coverage and load fairness, validating its effectiveness as an agentic AI-driven embodied networking solution for UAV-assisted vehicular edge computing.
Meng Yi, V. Lee, Miao Du et al.· IEEE Transactions on Cogniti...· 0 citations
Unmanned Aerial Vehicle (UAV) networks are increasingly deployed in dynamic environments where reliable and low‐latency communication is critical. However, high mobility, intermittent connectivity, spectrum limitations, and energy constraints make conventional static communication protocols inadequate for maintaining stable, dependable links. To address these challenges, this paper proposes TRC‐MAPPO, a topology‐aware reliability‐constrained multi‐agent deep reinforcement learning framework for adaptive UAV swarm communication. The routing problem is formulated as a constrained decision‐making task that jointly considers packet delivery reliability, end‐to‐end delay, link stability, bandwidth usage, and energy consumption. Unlike single‐agent DRL baselines, TRC‐MAPPO represents the swarm as a dynamic communication graph and combines local relay selection with centralized training, enabling cooperative routing decisions under time‐varying network conditions. Simulations are conducted in a controlled, dynamic UAV environment, with the same mobility and traffic settings used for all methods. The proposed framework is compared with DQN and PPO over different swarm densities. Results show that TRC‐MAPPO achieves a higher packet delivery ratio, lower end‐to‐end delay, and more efficient energy behaviour, particularly in dense deployments. These findings indicate that topology‐aware cooperative learning can provide a scalable and practical solution for reliable UAV communication in future intelligent aerial and 6G‐enabled networks.
V. Nam, A. Chehri, Weiwei Jiang et al.· Expert systems· 0 citations
The rapid expansion of the Internet of Things (IoT) and the emergence of sixth-generation (6G) wireless networks have created unprecedented opportunities for large-scale intelligent sensing, real-time data collection, and ubiquitous connectivity. However, the deployment of massive IoT sensor networks faces significant challenges, including limited energy resources, dynamic network topologies, communication reliability issues, routing inefficiencies, and coverage constraints, particularly in remote, disaster-stricken, and infrastructure-deficient environments where conventional terrestrial communication systems often fail to provide reliable services. Unmanned Aerial Vehicles (UAVs) have emerged as a promising solution for enhancing network coverage, improving data collection efficiency, and supporting communication services in IoT ecosystems; nevertheless, their integration introduces additional challenges related to energy consumption, trajectory planning, routing optimization, and resource allocation. To address these issues, this paper proposes an AI-Driven Energy-Efficient Routing and UAV Trajectory Optimization Framework for UAV-assisted IoT sensor networks operating in 6G environments. The proposed framework integrates intelligent routing, adaptive energy management, and dynamic UAV trajectory optimization within a unified cross-layer architecture and develops a comprehensive mathematical model to characterize the relationships among energy consumption, communication delay, packet delivery performance, routing decisions, and UAV mobility. Furthermore, a Deep Reinforcement Learning (DRL)-based optimization algorithm is introduced to enable autonomous decision-making and adaptive network control under dynamic environmental conditions. The proposed approach continuously monitors key network parameters, including residual sensor energy, link quality, transmission distance, traffic load, UAV battery status, and data collection requirements, and dynamically determines optimal routing paths and UAV flight trajectories to minimize overall energy consumption while maximizing network lifetime, packet delivery ratio, and data collection efficiency. In addition, the framework leverages the ultra-reliable low-latency communication capabilities envisioned for future 6G infrastructures to facilitate intelligent coordination between UAV platforms and IoT sensor nodes. Performance evaluation under various network densities, mobility scenarios, and communication conditions demonstrates that the proposed framework significantly reduces energy consumption, improves routing efficiency, extends network lifetime and enhances packet delivery performance, and decreases communication overhead and data collection latency compared with conventional approaches. The results confirm that the integration of artificial intelligence, energy-aware routing, and UAV trajectory optimization provides an effective and scalable solution for next-generation UAV-assisted IoT systems and establishes a robust foundation for intelligent 6G-enabled wireless sensor networks.
Mojtaba Nasehi· Internet of Things and Cloud...· 0 citations