Skip to content

Causality-Aware Scheduling and Mobility Control in Multi-UAV Cooperative MEC via Graph-Enhanced Dual-Timescale Learning

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 9917-9934 · 0 citations · 36 references
Computer Science

Abstract

Multi-uncrewed aerial vehicle (UAV) cooperative mobile edge computing (MEC) systems present significant challenges owing to task causal dependencies, dynamic channel variations, and multi-dimensional resource coupling. In this study, a multi-UAV cooperative MEC system with task causality constraints and time-varying wireless channels is considered, and the joint optimization of task offloading, task migration, dynamic UAV clustering, and continuous UAV trajectory planning is investigated. The objective is to minimize the long-term weighted sum of system latency and energy consumption while ensuring task queue stability. Lyapunov optimization is first introduced to transform the formulated stochastic mixed-integer problem into a deterministic per-slot optimization. Afterward, a dual-timescale graph-enhanced multi-agent proximal policy optimization (DT-HGMAPPO) framework is proposed to coordinate long-timescale UAV clustering and trajectory planning with short-timescale task offloading and migration. Specifically, this framework decouples the problem by utilizing dual-layer weighted hypergraph matching (DL-WHM) for joint clustering and association, a dynamic priority scoring (DPS) mechanism for intra-cluster load balancing, and a graph-enhanced MAPPO algorithm for trajectory optimization. Simulation results reveal that the proposed DT-HGMAPPO algorithm outperforms conventional multi-agent deep reinforcement learning baselines in terms of both convergence speed and policy stability. It achieves a final reward that is at least 18.0% greater than that of other multi-agent algorithms. Moreover, the proposed framework reduces total system cost by 33.3% compared with MASAC and 18.9% compared with MATD3, thereby achieving a superior delay-energy trade-off while ensuring queue stability.

View source

Similar papers

Open access Aug 2026

Drift-Plus-Penalty-Based Joint Optimization of Computational Resource Scheduling, Power Control, and UAV Flight Decisions in UAV-Enabled Mobile Edge Computing

A Lyapunov-based joint optimization framework for UAV-enabled MEC systems achieves a balanced tradeoff between delay, energy consumption, and UAV flight activity, supporting energy-efficient and delay-aware UAV-MEC operation.

Lei Li, Xue Gao, Quansheng Guan · 0 citations
#edge computing Open access Aug 2026

Collaborative resource allocation in UAV-assisted MEC networks: A heterogeneous MAPPO scheme

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) is a key enabler for meeting the stringent low-latency and energy-efficiency requirements of emerging low-altitude economy applications. However, achieving these objectives remains challenging due to dynamic environments, limited communication and computation resources, and the heterogeneity of network entities. This paper investigates the long-term joint optimization framework that minimizes system-wide latency and energy consumption simultaneously by coordinating UAV association, subchannel selection, uplink/downlink power allocation, and computational resource distribution. This sequential decision-making process is formulated into a partially observable Markov decision process (POMDP) to account for localized observations and dynamic channel states. To solve it, we propose a heterogeneous multi-agent proximal policy optimization (MAPPO)-based framework where both user devices (UDs) and UAVs act as heterogeneous agents. This architecture utilizes a centralized training and decentralized execution (CTDE) paradigm to enable collaborative strategies between computing requesters and providers. Numerical results demonstrate that the proposed scheme effectively navigates the high-dimensional action space and achieves superior convergence and cost reduction compared to benchmarks, including PPO, independent PPO (iPPO), Q-learning multi-agent extension (QMIX), value decomposition networks (VDN), independent deep Q-network (iDQN), and genetic algorithm (GA).

Ming Cheng, Canlin Zhu, Jianghang Tang et al. · 0 citations
#edge computing Preprint Aug 2026

Distributed Trajectory Planning and Resource Allocation for Dynamic Multi-UAV Collaborative Computing

This paper investigates a multiple uncrewed aerial vehicles (UAVs)-enabled distributed mobile edge computing (MEC) framework, where the set of collaborative UAVs dynamically varies over time due to their energy states and service loads. The joint optimization of trajectory planning and resource allocation is formulated as a Stackelberg game, where UAVs and mobile terminals (MTs) are modeled as leaders and followers, respectively. UAVs aim to maximize their benefits by balancing executed workload, energy cost, and resource allocation revenue, while MTs seek to minimize their total overhead, composed of computing delay and resource costs, through offloading and resource-request decisions. A hierarchical joint optimization algorithm is developed within a multi-agent deep reinforcement learning (MADRL) framework to coordinate UAVs and MTs in a distributed manner. At the leader level, UAVs jointly determine their trajectories, task migration ratios, MT-UAV association, and unit computing resource pricing. Each UAV is modeled as an agent in a partially observable Markov decision process, and the agents are jointly trained via multi-agent proximal policy optimization (MAPPO) under the centralized-training-and-decentralized-execution paradigm. At the follower level, MTs determine their optimal task offloading ratios and requested computing resources using a two-stage iterative algorithm. Simulation results demonstrate stable convergence under dynamic UAV participation. Compared to the no-collaboration benchmark, the proposed algorithm improves UAV efficiency by 18.58% through inter-UAV task migration and reduces average MT overhead by 33.77% over the fully offloading scheme. It also outperforms other benchmarks under varying network scales and capabilities by jointly optimizing UAV operations and resource utilization.

Tiankui Zhang, Wenlong Xu, Tianyi Shi et al. · 0 citations
Open access Jun 2026

Distributed Task Allocation and Trajectory Planning for Heterogeneous UAV Swarms in Multi-Constraint Environments

Owing to the stringent spatio-temporal coupling and kinematic constraints, the task allocation problem for heterogeneous unmanned aerial vehicle (UAV) swarms is generally regarded as an NP-hard problem. To address this, this paper proposes the Sequentially Extended Consensus-Based Bundle Algorithm (SECBBA), a deadlock-free distributed scheduling framework. First, a multi-task allocation model is established by incorporating constraints associated with payload resources, task scheduling, and threat zone. Subsequently, the conventional Consensus-Based Bundle Algorithm (CBBA) is extended through the integration of a deadlock detection and resolution mechanism based on directed graph Depth-First Search (DFS), thereby guaranteeing conflict-free task allocation. Furthermore, a sequential hierarchical strategy is introduced to transform global temporal dependencies into tractable soft time-window constraints. Finally, to ensure physical feasibility, Dubins curves are tightly coupled with the allocation process, enabling nonholonomic path planning for fixed-wing UAVs. Simulation results demonstrate that SECBBA reduces global task costs by 13.3%, 22.7%, and 39.4% compared to the Consensus-Based Bundle Algorithm with Temporal Consistency Constraints (CBBA-TCC), Improved Genetic Algorithm (IGA) and Q-Learning baselines, respectively. It consistently maintains performance advantage of 9.8%, 23.2% and 19.0% under variable weights with high computational efficiency, significantly enhancing swarm timeliness in complex, coupled multi-task scenarios.

Bochang Yu, Feng Gao, Wen Wu et al. · 0 citations
Preprint Jul 2026

Multi-Agent Reinforcement Learning for SLA-Aware Network Slicing in UAV-Enabled MEC

A predictive multi-agent Reinforcement Learning (RL) framework that proactively maintains SLA stability in UAV-enabled MEC through coordinated trajectory control and computation resource allocation and designs an SLA-aware reward function that explicitly penalizes both violation probability and duration across slices.

M. Farhoudi, Zeinab Sasan, Masoud Shokrnezhad et al. · 0 citations
2026

Collaborative Trajectory and Resource Optimization in Multi-UAV MEC Under Jamming: An LLM-Guided MARL Framework

Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.

Yeguang Qin, Jie Tang, Fengxiao Tang et al. · 0 citations