Back to feed
Open access

MULTI-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING AND RESOURCE ALLOCATION IN MEC SYSTEMS

Jul 2026 · International journal of Computer Networks & Communications · 0 citations · 30 references

TL;DR

This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.

Abstract

Multi-Access Edge Computing (MEC) brings computation closer to end users to reduce latency and energy consumption for compute-intensive mobile applications. In this paper, we address the joint task offloading and resource allocation problem in multi-user MEC systems and propose a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL). Each user–base station link is modeled as an autonomous agent that decides whether to execute tasks locally or offload them to the edge, and how much computing resource to request. The learning framework follows a centralized training and decentralized execution paradigm and integrates Q-Learning, Deep Q-Network (DQN), and Double DQN (DDQN) algorithms. Extensive simulations show that the DDQN-based approach achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines. The results confirm that lightweight MARL is a practical and scalable solution for dynamic MEC environments under realistic resource constraints.

Read PDF

Similar papers

Open access Jul 2026

Constraint-Aware Resource Exploration for Multi-Agent Collaborative Offloading in Mobile Edge Computing

Mobile edge computing (MEC) supports computation-intensive and latency-sensitive Internet of Things (IoT) applications. However, collaborative task offloading in dynamic heterogeneous environments remains challenging due to coupled physical constraints, shared resource competition, and high-dimensional decision spaces. Existing multi-agent deep reinforcement learning (MADRL) approaches often rely on static penalties or centralized action truncation for constraint handling. These methods may lead to unstable training, conservative strategies, and limited collaboration. To address these limitations, this paper proposes a constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE). The offloading problem is formulated as a constrained Markov decision process and addressed under a centralized training and decentralized execution (CTDE) framework. Dynamic Lagrange multipliers replace fixed penalties to improve training stability and support smoother exploration near constraint boundaries. A multi-threshold-guided Lagrangian constraint regulation mechanism further coordinates heterogeneous constraints, including energy consumption, latency, and server capacity. In addition, a congestion-driven cost allocation method transforms global resource competition into dynamic cost signals, guiding agents toward more coordinated offloading decisions. The simulation results show that CARE-CTDE achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.

Yuxuan Yang, Hexing Wang, Yang Zhou · 0 citations
Open access Jul 2026

TWO-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING IN IOT-MEC NETWORKS

The rapid proliferation of Internet of Things (IoT) devices has placed unprecedented pressure on the network edge, where applications such as augmented reality, real-time analytics, and autonomous navigation demand low latency and tight energy budgets that traditional cloud-centric architectures cannot meet. Multi-access Edge Computing (MEC) addresses this gap by relocating computation closer to end users, but the core question of where and how each task should be executed remains open: rulebased and single-objective offloading strategies fail to simultaneously balance service latency, energy efficiency, and user experience under dynamic, large-scale conditions. In this paper we propose TARLOT (Two-Agent Reinforcement Learning Offloading Tasks), a cooperative framework for threetier IoT–MEC–Cloud environments. TARLOT decouples the offloading decision from the resourceallocation problem and assigns each to a dedicated Q-learning agent, so that the two subproblems are specialised independently while still being optimised jointly. The framework is evaluated on PureEdgeSim under heterogeneous IoT workloads, device densities ranging from 200 to 2,400, and diverse application profiles, and is compared against five widely-used baselines (Random, Round-Robin, Trade-Off, Pure-Edge, and Pure-Cloud). At 2,400 devices, TARLOT delivers an average service time of 1.1 s (against 4.3 s for Pure-Cloud), a Quality of Experience of 0.77 (against 0.22 for Pure-Cloud), a task-failure rate below 2 % (against nearly 14 % for Pure-Cloud), and a per-device energy consumption of only 3.6 W (against 11.2 W for Pure-Cloud) — roughly a 68 % reduction. Balanced CPU utilisation across the local, edge, and cloud tiers further confirms that TARLOT prevents resource bottlenecks, establishing it as a practical solution for next-generation large-scale IoT deployments.

Oussama Lagnfdi, Marouane Myyara, A. Darif · 0 citations
2026

Energy-Efficient Task Offloading and Load Balancing for Multi-UAV-Assisted Vehicular Networks

The rapid growth of Internet of Vehicles (IoV) applications has imposed strict requirements on low-latency and energy-efficient computing services. This letter investigates a multi-Uncrewed Aerial Vehicle (UAV)-assisted IoV system, where multiple Mobile Edge Computing (MEC)-enabled UAVs (MUs) collaboratively provide computing services for vehicular terminals (VTs). To improve service capability, we propose an energy-efficient task offloading and load balancing scheme that jointly considers vehicle mobility, task offloading and migration, and computing resource allocation to formulate an optimization problem. To solve this problem, a collective learning (CL)-enabled multi-agent reinforcement learning (CL-MARL) algorithm is proposed, where each agent learns optimal policies through centralized training and collective cooperative learning. Simulation results demonstrate that the proposed scheme outperforms benchmark strategies in terms of energy efficiency, task completion rate, and load balancing.

Yongbin Wang, Peng Lin, Yan Liu et al. · 0 citations
Open access Jul 2026

Toward Low-Delay and Energy-Efficient UAV-Assisted MEC Systems Through Intelligent Resource Allocation

A Prioritized Adaptive Weighting based on Deep Deterministic Policy Gradient (PAW-DDPG) as an enhanced Deep Deterministic Policy Gradient (DDPG) algorithm to minimize both processing delay and energy consumption by jointly optimizing user scheduling, partial-task offloading, and UAV trajectory is proposed.

W. Saber, Hanan Algamil, Fifi Farouk et al. · 0 citations
2026

Toward Stable and Low-Latency Task Offloading: A Multi-Agent Framework for Vehicular Edge Computing

With the rapid growth of Vehicular Edge Computing (VEC) and Mobile Edge Computing, efficient task offloading is essential for enhancing the computing and communication capabilities in vehicular networks. However, many existing methods suffer from slow convergence, load imbalance, and instability in dynamic, latency-sensitive environments. To address these challenges, we propose MAPPO-Lyapunov (MAPPO-L), a multi-agent offloading framework that integrates Multi-Agent Proximal Policy Optimization (MAPPO) with Lyapunov optimization. MAPPO-L enables distributed coordination among vehicles, roadside units (RSUs), and cloud servers, minimizing delay, improving resource utilization, and ensuring long-term stability. Lyapunov theory transforms long-term stability into per-slot optimizations, while MAPPO ensures efficient policy learning. An adaptive exploration mechanism dynamically adjusts exploration rates based on network dynamics, accelerating convergence and stabilizing training. Extensive simulations with real-world data show that MAPPO-L maintains task completion rates above 80%, converges 25%–37.5% faster than baselines, and reduces training fluctuations to 2.3%. Ablation studies confirm the critical roles of location, channel, and queue information, validating the robustness of MAPPO-L in practical VEC environments.

Lu Wei, Yong Yu, Jie Cui et al. · 0 citations
2026

Collaborative Task Offloading in Space Computing Power Network: A World Model-Based Multi-Agent Reinforcement Learning Approach

Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-agent reinforcement learning (MARL) provides an adaptive solution, existing model-free MARL methods often suffer from slow convergence, insufficient foresight, and limited robustness in dynamic satellite environments. To address these challenges, this paper proposes a World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework for space computing power networks. The offloading problem is formulated as a partially observable multi-agent decision-making process, where LEO satellites make decentralized decisions under incomplete local observations. A predictive world model learns latent transition dynamics of network states and provides future context for proactive planning. Meanwhile, a Transformer-based policy architecture captures inter-agent dependencies and supports cooperative scheduling under centralized training and decentralized execution (CTDE). Simulation results show that WM-MAPPO achieves higher task completion ratios, lower average latency, improved energy efficiency, and stronger robustness than model-free MARL baselines, heuristic methods, Lyapunov-based scheduling, and MINLP-inspired optimization.

Yuqi Cong, Zhiwei Wei, Jiarui Chen et al. · 0 citations