A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.
Abstract
Unmanned Aerial Vehicle (UAV)-assisted Mobile Edge Computing (MEC) has emerged as a promising paradigm for supporting computation-intensive and delay-sensitive applications. However, efficient task offloading and resource allocation remain a challenging problem due to the need to jointly minimize the maximum processing delay and energy consumption of User Device (UD) in dynamic environments. Existing solutions often suffer from training instability, limited exploration capabilities, and slow convergence, limiting their ability to achieve optimal task offloading and resource allocation decisions. To address these challenges, this paper proposes a Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory. The proposed algorithm introduces a state-aware normalization mechanism to stabilize the learning process, a new pre-training initialization technique that populates the Experience Replay Buffer (ERB) before learning to accelerate convergence and improve policy quality, and a hybrid noise exploration scheme that enhances exploration efficiency. Furthermore, to achieve an effective balance between delay and energy consumption, a novel adaptive weighting mechanism based on a modified Exponential Moving Average (EMA) algorithm is proposed. Simulation results demonstrate that the proposed PAW-DDPG algorithm outperforms DDPG and all baseline algorithms, achieving performance gains over DDPG of 4–18%, 9–18%, 8–26%, and 7–18% under varying task sizes, UAV computing capabilities, user device computing capabilities, and numbers of user devices, respectively.
The rapid growth of Internet of Vehicles (IoV) applications has imposed strict requirements on low-latency and energy-efficient computing services. This letter investigates a multi-Uncrewed Aerial Vehicle (UAV)-assisted IoV system, where multiple Mobile Edge Computing (MEC)-enabled UAVs (MUs) collaboratively provide computing services for vehicular terminals (VTs). To improve service capability, we propose an energy-efficient task offloading and load balancing scheme that jointly considers vehicle mobility, task offloading and migration, and computing resource allocation to formulate an optimization problem. To solve this problem, a collective learning (CL)-enabled multi-agent reinforcement learning (CL-MARL) algorithm is proposed, where each agent learns optimal policies through centralized training and collective cooperative learning. Simulation results demonstrate that the proposed scheme outperforms benchmark strategies in terms of energy efficiency, task completion rate, and load balancing.
Yongbin Wang, Peng Lin, Yan Liu et al.· IEEE Wireless Communications...· 0 citations
A Lyapunov-based joint optimization framework for UAV-enabled MEC systems achieves a balanced tradeoff between delay, energy consumption, and UAV flight activity, supporting energy-efficient and delay-aware UAV-MEC operation.
Lei Li, Xue Gao, Quansheng Guan· Electronics· 0 citations
With the rapid development of the Internet of Things (IoT) and mobile computing, edge computing has emerged as a promising paradigm for providing low-latency and energy-efficient services. However, in some extremely computation-intensive scenarios, conventional terrestrial edge computing may fail due to the insufficient computing capability of ground base stations. Fortunately, multi-UAV-assisted edge computing offers a promising solution to this challenge. Nevertheless, existing methods often struggle to provide efficient horizontal cooperative deployment for multiple UAVs with low computational overhead. To address this issue, this paper considers user randomness and inter-UAV collaboration, and proposes a low-complexity yet highly adaptive approach for cooperative deployment and task-scheduling optimization in multi-UAV-assisted edge computing systems. Specifically, we formulate the problem as a stochastic optimization problem that minimizes the energy consumption of ground users while ensuring UAV battery endurance and overall system performance. We then propose a dynamic cooperative deployment and task scheduling (DCDTS) algorithm that integrates K-means clustering with the Lyapunov optimization framework. Through Lyapunov optimization, the original dynamic optimization problem is transformed into a deterministic problem and further decomposed into multiple subproblems that can be solved in parallel. K-means is exploited to enable cooperative UAV deployment and user offloading decisions, while non-convex optimization and nonlinear programming are employed to solve the task-scheduling and resource-allocation subproblem. Extensive parameter analysis and comparative experiments demonstrate that the proposed dynamic cooperative deployment algorithm can effectively reduce user energy consumption while maintaining UAV energy constraints and system performance.
Mobile edge computing (MEC) supports computation-intensive and latency-sensitive Internet of Things (IoT) applications. However, collaborative task offloading in dynamic heterogeneous environments remains challenging due to coupled physical constraints, shared resource competition, and high-dimensional decision spaces. Existing multi-agent deep reinforcement learning (MADRL) approaches often rely on static penalties or centralized action truncation for constraint handling. These methods may lead to unstable training, conservative strategies, and limited collaboration. To address these limitations, this paper proposes a constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE). The offloading problem is formulated as a constrained Markov decision process and addressed under a centralized training and decentralized execution (CTDE) framework. Dynamic Lagrange multipliers replace fixed penalties to improve training stability and support smoother exploration near constraint boundaries. A multi-threshold-guided Lagrangian constraint regulation mechanism further coordinates heterogeneous constraints, including energy consumption, latency, and server capacity. In addition, a congestion-driven cost allocation method transforms global resource competition into dynamic cost signals, guiding agents toward more coordinated offloading decisions. The simulation results show that CARE-CTDE achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.
Yuxuan Yang, Hexing Wang, Yang Zhou· Mathematics· 0 citations
This paper proposes a joint optimization algorithm for trajectory control and task offloading ratios based on multi-agent deep reinforcement learning. By jointly optimizing the flight trajectories of unmanned aerial vehicles (UAVs), user scheduling strategies, and task offloading ratios, the decoupled coordination of resource allocation and trajectory planning is achieved, thereby minimizing system delay and weighted energy consumption. An enhanced multi-agent proximal policy optimization algorithm, named FMAHPPO, is designed. Compared with existing benchmark algorithms, the FMAHPPO algorithm significantly reduces the total system overhead and effectively improves the energy efficiency and task processing success rate of multi-UAV swarms. This research provides a valuable theoretical foundation and algorithmic support for the collaborative management of edge resources in future space-air-ground integrated networks (SAGIN).
Unmanned aerial vehicle (UAV)-assisted mobile edge computing (MEC) systems provide flexible computing services for resource-constrained devices, but malicious jamming attacks introduce dynamic channel conditions and resource competition, making joint trajectory and resource optimization challenging. This paper investigates this problem in multi-UAV MEC systems under jamming, aiming to minimize delay and energy consumption while ensuring anti-jamming robustness. The problem is formulated as a decentralized partially observable Markov decision process (Dec-POMDP). However, traditional multi-agent reinforcement learning (MARL) approaches struggle with high exploration costs and low sampling efficiency in high-dimensional hybrid action spaces. To overcome these limitations, we propose an LLM-guided MARL framework instantiated with the multi-agent deep deterministic policy gradient (MADDPG), which leverages LLM-generated semantic trajectory prompts to dynamically constrain exploration within the continuous action space, effectively compressing the policy search space and accelerating convergence. Simulation results demonstrate that the proposed method achieves $3.4\times $ to $5\times $ faster convergence over hierarchical MADDPG, MADDPG, and independent soft actor-critic (ISAC) baselines, significantly reducing training costs while maintaining superior performance and anti-jamming robustness.
Yeguang Qin, Jie Tang, Fengxiao Tang et al.· IEEE Transactions on Communi...· 0 citations