Back to feed

Similar papers

Preprint Jul 2026

Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

Unmanned Aerial Vehicle (UAV)-enabled Mobile Edge Computing (MEC) offers flexible capacity provisioning for heterogeneous network slices, including Hyper-Reliable and Low-Latency Communication (HRLLC), Enhanced Mobile Broadband (eMBB), and Massive Machine-Type Communications (mMTC). However, guaranteeing slice-level Service-Level Agreements (SLAs) under dynamic user mobility, stochastic task arrivals, and constrained onboard energy and computing resources remains a fundamental challenge. This paper proposes a predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation. A lightweight prediction module forecasts near-future user mobility, enabling UAVs to anticipate congestion and reposition before SLA violations occur. We design an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices, alongside total energy consumption. UAV agents are trained using Multi-Agent Proximal Policy Optimization (MAPPO) with centralized training and decentralized execution, enabling scalable online decision-making. Event-driven simulations with realistic mobility traces demonstrate that the proposed framework significantly improves SLA stability compared with baselines while maintaining competitive energy efficiency and delay performance, approaching oracle-level performance with sufficiently accurate predictive information.

M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al. · 0 citations
2026

Optimizing Information Freshness in Satellite-UAV IoRT Networks: A Heterogeneous Multi-Agent Approach

In satellite-UAV assisted communication networks, jointly optimizing the UAV’s trajectory and the multi-agent scheduling decisions to minimize the age of information (AoI) is a notoriously challenging problem. The complexity is compounded by the fundamental heterogeneity between the satellite and UAV agents, including their disparate action spaces, partial observations, and differing energy-consumption and communication-cost penalties. To address this, we formulate the problem as a decentralized partially observable Markov decision process (Dec-POMDP) and propose a novel heterogeneous multi-agent compound-action proximal policy optimization (HMACPPO) algorithm. HMACPPO leverages a centralized training with decentralized execution (CTDE) framework, using role-specific decentralized actors together with agent-specific centralized critics conditioned on the global state. Specifically, the UAV employs a compound PPO (CPPO) actor for its hybrid action space, while the satellite uses a PPO actor for discrete scheduling. Extensive simulations show that HMACPPO outperforms the compared baselines, and that the resulting coordinated policy effectively manages the trade-off between AoI, UAV energy consumption, and operational cost.

Weijie Zhou, Mengjie Yi, Yan Zhang et al. · 0 citations
2026

SkySched: A Hierarchical and Scalable Reinforcement Learning Framework for Multi-UAV Vehicular Edge Computing Network

Unmanned Aerial Vehicles (UAVs) are increasingly deployed as embodied aerial agents in low-altitude economies, forming mobile aerial edge networks that enable flexible computation offloading for vehicles. However, their limited endurance and frequent join/leave behaviours result in highly dynamic topologies, undermining long-term resource availability. Moreover, existing vehicle-centric task scheduling strategies cause resource contention and decision complexity in dense environments. To address these challenges, this paper proposes a hierarchical and scalable reinforcement learning-based scheduling framework (SkySched). In SkySched, UAVs collaboratively make deployment and task scheduling decisions. The framework consists of two tightly coupled modules. First, an adaptive UAV deployment module introduces a capability encoding mechanism that compresses heterogeneous UAV attributes into a unified one-dimensional capability index. This compact representation enables a Scalable Proximal Policy Optimization (SPPO) algorithm to efficiently coordinate UAV positioning, maximizing task coverage and sustaining network-wide computing availability under dynamic topology variations. Second, a hierarchical task scheduling module is designed, where K-means-based Roadside Unit (RSU) clustering enables vertical task offloading, while a SPPO-driven horizontal UAV-to-UAV task redistribution mechanism achieves fine-grained load balancing across the UAV swarm. Simulations demonstrate that SkySched consistently outperforms state-of-the-art methods in terms of task coverage and load fairness, validating its effectiveness as an agentic AI-driven embodied networking solution for UAV-assisted vehicular edge computing.

Meng Yi, V. Lee, Miao Du et al. · 0 citations
2026

Agentic AI-Empowered Reliable Routing in Interference-Aware UAV Networks

The burgeoning low-altitude economies place stringent demands on reliable communication within uncrewed aerial vehicle (UAV) networks. However, in complex electromagnetic environments, the topology and communication link quality of highly maneuverable UAVs exhibit strong random fluctuations, leading to frequent routing path failures. Agentic AI integrated with embodied intelligence is leveraged to empower UAV networks, thereby significantly enhancing communication reliability and autonomous adaptability. Specifically, we propose an interference-aware multi-agent cooperative routing optimization framework HRC-QMIX for embodied-enhanced communication, integrating the hypergraph module to characterize the cooperative dependencies among multiple embodied nodes, thereby supporting more stable forwarding decisions. Within a value decomposition learning framework, we further introduce a dynamic mixing mechanism based on recursive atrous self-attention (RASA) to enhance the expressive ability of the joint value function for complex cooperative relationships. Furthermore, causal-inspired regularization term is designed to alleviate the instability of credit allocation in strongly non-stationary scenarios and improve training stability. Simulation results demonstrate that the proposed method exhibits superior communication performance and stronger anti-interference robustness in scenarios with multiple interference sources and varying maneuverability levels.

Wenjing Wei, Tianyu Wang, Hanze Liu et al. · 0 citations
Jul 2026

Coverage-aware offloading in multi-UAV-aided terrestrial MEC networks using quantum-inspired particle swarm optimization

An MECN that integrates UAVs as the aerial layer and TESs as the terrestrial layer is introduced, and a quantum-inspired particle swarm optimization-based offloading strategy (QIPSO-TOS) is proposed to facilitate coverage-aware task offloading.

Marlom Bey, P. Kuila, Biswadip Bandyopadhyay et al. · 0 citations
Preprint Jul 2026

TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.

Muhammad Umar Farooq Qaisar, Lin Zhang, Zhen Chen et al. · 0 citations