A constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE) that achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.
Abstract
Mobile edge computing (MEC) supports computation-intensive and latency-sensitive Internet of Things (IoT) applications. However, collaborative task offloading in dynamic heterogeneous environments remains challenging due to coupled physical constraints, shared resource competition, and high-dimensional decision spaces. Existing multi-agent deep reinforcement learning (MADRL) approaches often rely on static penalties or centralized action truncation for constraint handling. These methods may lead to unstable training, conservative strategies, and limited collaboration. To address these limitations, this paper proposes a constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE). The offloading problem is formulated as a constrained Markov decision process and addressed under a centralized training and decentralized execution (CTDE) framework. Dynamic Lagrange multipliers replace fixed penalties to improve training stability and support smoother exploration near constraint boundaries. A multi-threshold-guided Lagrangian constraint regulation mechanism further coordinates heterogeneous constraints, including energy consumption, latency, and server capacity. In addition, a congestion-driven cost allocation method transforms global resource competition into dynamic cost signals, guiding agents toward more coordinated offloading decisions. The simulation results show that CARE-CTDE achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.
This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.
Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al.· International journal of Com...· 0 citations
An adaptive Beta-policy and delayed-update multi-agent soft actor-critic method, abbreviated as ABDMASAC, which uses a Beta policy to model bounded actions and achieves a better overall trade-off than the selected MASAC-backbone and on-policy MARL baselines under the considered simulation settings.
Zheng Yao, Jie Liu, Changjun Deng et al.· Computers, Materials & C...· 0 citations
A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.
W. Saber, Hanan Algamil, Fifi Farouk et al.· Future Internet· 0 citations
A task-driven offloading algorithm based on Balanced Multi-Agent Deep Deterministic Policy Gradient (BMADDPG) that reduces average task processing latency by approximately 22.67% and decreases total system cost by at least 18.32% under high-load scenarios.
With the rapid growth of Vehicular Edge Computing (VEC) and Mobile Edge Computing, efficient task offloading is essential for enhancing the computing and communication capabilities in vehicular networks. However, many existing methods suffer from slow convergence, load imbalance, and instability in dynamic, latency-sensitive environments. To address these challenges, we propose MAPPO-Lyapunov (MAPPO-L), a multi-agent offloading framework that integrates Multi-Agent Proximal Policy Optimization (MAPPO) with Lyapunov optimization. MAPPO-L enables distributed coordination among vehicles, roadside units (RSUs), and cloud servers, minimizing delay, improving resource utilization, and ensuring long-term stability. Lyapunov theory transforms long-term stability into per-slot optimizations, while MAPPO ensures efficient policy learning. An adaptive exploration mechanism dynamically adjusts exploration rates based on network dynamics, accelerating convergence and stabilizing training. Extensive simulations with real-world data show that MAPPO-L maintains task completion rates above 80%, converges 25%–37.5% faster than baselines, and reduces training fluctuations to 2.3%. Ablation studies confirm the critical roles of location, channel, and queue information, validating the robustness of MAPPO-L in practical VEC environments.
Lu Wei, Yong Yu, Jie Cui et al.· IEEE Transactions on Network...· 0 citations
The rapid growth of Internet of Vehicles (IoV) applications has imposed strict requirements on low-latency and energy-efficient computing services. This letter investigates a multi-Uncrewed Aerial Vehicle (UAV)-assisted IoV system, where multiple Mobile Edge Computing (MEC)-enabled UAVs (MUs) collaboratively provide computing services for vehicular terminals (VTs). To improve service capability, we propose an energy-efficient task offloading and load balancing scheme that jointly considers vehicle mobility, task offloading and migration, and computing resource allocation to formulate an optimization problem. To solve this problem, a collective learning (CL)-enabled multi-agent reinforcement learning (CL-MARL) algorithm is proposed, where each agent learns optimal policies through centralized training and collective cooperative learning. Simulation results demonstrate that the proposed scheme outperforms benchmark strategies in terms of energy efficiency, task completion rate, and load balancing.
Yongbin Wang, Peng Lin, Yan Liu et al.· IEEE Wireless Communications...· 0 citations