Skip to content
Open access

Master-Refined MAPPO for Long-Term Joint Resource Scheduling in NOMA-MEC Systems

Jul 2026 · Symmetry · 0 citations · 31 references

TL;DR

This study jointly optimizes task offloading and system resource scheduling to minimize the long-term delay–energy cost of NOMA-MEC systems using a master-refined multi-agent proximal policy optimization algorithm.

Abstract

Mobile edge computing (MEC) enables resource-constrained user devices (UDs) to obtain low-latency computing services by offloading computational tasks to the network edge. Non-orthogonal multiple access-enabled mobile edge computing (NOMA-MEC) systems feature asymmetric states across UDs, dynamic task arrivals, and competition for wireless and edge computing resources. Under these conditions, offloading decisions affect device energy consumption, task delay, and edge computing resource allocation, making long-term system optimization difficult. This study jointly optimizes task offloading and system resource scheduling to minimize the long-term delay–energy cost. The problem is formulated as a partially observable Markov decision process (POMDP) and addressed using a master-refined multi-agent proximal policy optimization (MR-MAPPO) algorithm. MR-MAPPO combines continuous action relaxation, master action refinement, and a behavior cloning auxiliary term to learn policies in a hybrid discrete–continuous action space. A marginal congestion delay term is also introduced to capture the impact of newly admitted tasks on existing edge workloads. Simulation results show that MR-MAPPO outperforms the considered baselines, while ablation studies verify the effects of its key components. Under the main experimental setting, MR-MAPPO reduces the system cost by 17.9% and 22.9% relative to standard MAPPO and particle swarm optimization (PSO), respectively.

Read PDF

Similar papers

Open access 2026

Cooperative Task Offloading in Mobile Edge Computing via an Improved MASAC Framework

An adaptive Beta-policy and delayed-update multi-agent soft actor-critic method, abbreviated as ABDMASAC, which uses a Beta policy to model bounded actions and achieves a better overall trade-off than the selected MASAC-backbone and on-policy MARL baselines under the considered simulation settings.

Zheng Yao, Jie Liu, Changjun Deng et al. · 0 citations
Open access Jul 2026

MULTI-AGENT REINFORCEMENT LEARNING FOR TASK OFFLOADING AND RESOURCE ALLOCATION IN MEC SYSTEMS

This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.

Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al. · 0 citations
Conference Aug 2026

Resource Allocation and Task Offloading for MEC-Enabled 6G Networks Using DRL

A new paradigm for satisfying the ever-growing demands of real-time Sixth Generation (6G) applications is Mobile Edge Computing (MEC). Additionally, base stations and Internet of Things devices that incorporate renewable energy harvesting capabilities have the potential to lower grid energy use. To maximize system potential and lower carbon emissions, it is crucial to make effective decisions about job offloading and resource allocation. A carbon-aware MEC architecture that uses both grid and renewable energy sources is proposed in this paper. Our goal is to jointly manage resource allocation and task offloading while monitoring carbon emissions and task queue delays to optimize system behavior under uncertainty, specifically for stochastic workloads and variable renewable generation. To balance these two cost components (emissions and queue length), we create a combined optimization problem. We develop a deep deterministic policy gradient (DDPG)-based joint optimization technique to address this issue in a constantly changing environment. In the optimization, we consider greedy policy (GP) and full offloading (FO), as well as time-average carbon emission (TACE) and time-average queue length (TAQL) as performance metrics, and time-average queue length (TAQL) and full execution (FE) as baseline strategies; we also evaluate normalized time-average cumulative reward (NTACR). This method uses continuous-action reinforcement learning to generate efficient, real-time control policies. For the proposed MEC network, numerical statistics show that our approach can lead to effective offloading and lower carbon emissions.

M. Saeed, Rashid A. Saeed, M. A. Ahmed et al. · 0 citations
Open access Jul 2026

Constraint-Aware Resource Exploration for Multi-Agent Collaborative Offloading in Mobile Edge Computing

A constraint-aware multi-agent edge collaborative offloading algorithm (CARE-CTDE) that achieves better scheduling performance, resource utilization, and constraint satisfaction than baseline methods in dynamic heterogeneous MEC scenarios, demonstrating its effectiveness and robustness for constrained edge computing systems.

Yuxuan Yang, Hexing Wang, Yang Zhou · 0 citations
Jul 2026

Intelligent Cooperative Computation Offloading and Resource Allocation for Dual-Dependency Tasks in Edge Computing

Mobile edge computing (MEC) has accelerated the development of artificial intelligence and Internet of Things technologies, leading to the explosive growth of intelligent applications characterized by resource intensity and latency sensitivity, such as image processing and smart home. In practice, an application typically consists of multiple tasks with execution dependencies, where the output of some tasks serves as the input for specific others. Recently, the design of computation offloading methods for such execution-dependent tasks has received extensive research. However, computation offloading for execution-dependent tasks with service dependencies in resource-constrained multi-user, multi-edge-server cooperative MEC systems has not been thoroughly studied. In this paper, we formulate a cooperative computation offloading problem for dual-dependency tasks in multi-edge-server scenarios with limited service and computing resources, aiming to minimize the long-term average service delay for multiple users. To solve this problem, we propose a recurrent multi-agent reinforcement learning-based dual-dependency task offloading (RMA-DepO) algorithm, which enables users to communicate during training to explore and learn optimal joint task offloading and computing resource allocation strategies, and to make distributed offloading decisions at execution time. Simulation results demonstrate that the proposed RMA-DepO algorithm outperforms several baselines under different network settings, demonstrating its effectiveness in coordinating edge resources for cooperative computation of dual-dependency tasks.

Zhixiu Yao, Yun Li, Qilie Liu et al. · 0 citations
Open access Jul 2026

Multi-Objective Balanced Optimization Task Offloading Algorithm Based on Multi-Agent Collaboration

A task-driven offloading algorithm based on Balanced Multi-Agent Deep Deterministic Policy Gradient (BMADDPG) that reduces average task processing latency by approximately 22.67% and decreases total system cost by at least 18.32% under high-load scenarios.

Hui Li, Zhilong Zhu, Wanwei Huang et al. · 0 citations