This paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent and the Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues.
Abstract
Time-sensitive networking (TSN) is increasingly integrated into mobile edge computing (MEC) to support applications with stringent latency requirements, such as extended reality (XR). However, existing TSN scheduling solutions predominantly rely on static optimization techniques or centralized learning models that are based on fixed traffic patterns, limiting their effectiveness in dynamic environments. In practice, MEC environments often host multiple co-located XR traffic flows whose characteristics evolve over time, creating complex inter-queue dependencies that current schedulers fail to capture. Addressing these challenges requires adaptive, decentralized scheduling mechanisms capable of coordinating multiple TSN queues under varying traffic conditions. To this end, this paper proposes a multi-agent reinforcement learning (MARL) framework for TSN scheduling, where each TSN queue is modeled as an autonomous agent. The Heterogeneous-Agent Proximal Policy Optimization (HAPPO) algorithm is employed to explicitly model inter-agent dependencies and jointly optimize service delivery across queues. The simulation results demonstrate that the proposed approach reduces average frame waiting times by up to 26.8% and worst-case delays by approximately 16.8%, highlighting its effectiveness in dynamic XR-driven MEC scenarios.
Time-Sensitive Networking (TSN) and Mobile Edge Computing (MEC) hold strong potential for enabling ultra-reliable low-latency communication for time-sensitive applications, such as eXtended Reality (XR). However, the widespread adoption of XR introduces significant challenges due to co-located services in MEC environments, leading to contention for shared network resources. Moreover, XR traffic types have distinct characteristics and criticality in terms of timing requirements, further increasing the complexity and dynamics of such environments. Although reinforcement learning has shown promise for TSN scheduling optimization in dynamic network scenarios, existing approaches rely on centralized or high-level multi-agent designs and are typically tailored to periodic and predictable industrial traffic, limiting their applicability to XR workloads. As a result, these approaches suffer from (i) limited ability to capture inter-queue dependencies due to coarse-grained control, and (ii) poor adaptability to highly dynamic and heterogeneous XR traffic. To address these gaps, we propose a multi-agent reinforcement learning approach for queue-level XR traffic scheduling. We adopt the multi-agent transformer (MAT) to model inter-queue dependencies via attention over agents'observations and actions, enabling implicit coordination across heterogeneous co-located XR applications. Our simulation results show that the proposed method outperforms baselines, achieving up to 71.42% latency reduction and up to 83.2% reduction in failure rate, while consistently achieving high reliability across all queues.
Marcos Carvalho, Fatih Temiz, Shavbo Salehi et al.· 0 citations
This paper addresses the joint task offloading and resource allocation problem in multi-user MEC systems and proposes a decentralized control framework based on Multi-Agent Reinforcement Learning (MARL), which achieves lower total system cost and faster convergence than the full-local, full-offload, and heuristic baselines.
Youssef Oukissou, Mohamed Amine Meddaoui, Ayoub Belaidi et al.· International journal of Com...· 0 citations
Sixth-generation (6G) networks are expected to rely on agentic artificial intelligence for zero-touch, self-managed orchestration of heterogeneous network slices serving enhanced mobile broadband (eMBB), ultra-reliable low-latency communication (URLLC), and massive machine-type communication (mMTC). A central and under-studied challenge for adaptive multi-agent resource management (AMRM) in such settings is multi-timescale non-stationarity: channel fading evolves per time-slot, user demand shifts at the window scale, and service-level agreement (SLA) regimes change at an operational scale. Single-timescale multi-agent reinforcement learning (MARL) algorithms cannot track all three signals cleanly—a learning rate fast enough for the per-slot channel destabilises the coordination structure that governs longer-timescale policies. This paper proposes Nested-MARL, an independent-learner actor-critic algorithm in which each agent’s parameters are partitioned into three groups updated at separated rates <inline-formula> <tex-math notation="LaTeX">$\alpha _{0}\!\ll \!\alpha _{1}\!\ll \!\alpha _{2}$ </tex-math></inline-formula>, with a continuum-memory exponential moving average (EMA) anchoring the slowest group. The design is grounded in the Nested Learning paradigm of Behrouz et al. (2025) and is extended here from single-model continual learning to decentralised multi-agent coordination. We establish a finite-time convergence result in the two-timescale stochastic approximation framework showing that under standard regularity and timescale-separation conditions, Nested-MARL achieves <inline-formula> <tex-math notation="LaTeX">$O(T^{-1/2})$ </tex-math></inline-formula> fast-group convergence vs. an <inline-formula> <tex-math notation="LaTeX">$\Omega (T^{-1/3})$ </tex-math></inline-formula> lower bound for any single-timescale algorithm. An empirical study on a three-agent 6G slicing simulator with continuous multi-timescale drift shows Nested-MARL outperforms independent PPO (IPPO) in mean reward at every drift severity we test (<inline-formula> <tex-math notation="LaTeX">$\kappa \!\in \!\{0.5,1.0,1.5,2.0\}$ </tex-math></inline-formula>) and by + 8.6% in sample efficiency over the first 40 episodes at <inline-formula> <tex-math notation="LaTeX">$\kappa {=}1.5$ </tex-math></inline-formula> (<inline-formula> <tex-math notation="LaTeX">$n{=}10$ </tex-math></inline-formula> seeds, <inline-formula> <tex-math notation="LaTeX">$p\lt 0.05$ </tex-math></inline-formula>). A controlled ablation establishes that stripping timescale separation reduces performance below the IPPO baseline, isolating timescale separation as the causal mechanism. Nested-MARL also reduces policy switching cost by 16.6%, an operationally meaningful benefit for zero-touch orchestration. The complete simulator, agents, and 60 + per-seed training runs are released as open source.
Abraheem Rashid, Faisal Iradat, Waseem Iqbal et al.· IEEE Open Journal of the Com...· 0 citations
Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-agent reinforcement learning (MARL) provides an adaptive solution, existing model-free MARL methods often suffer from slow convergence, insufficient foresight, and limited robustness in dynamic satellite environments. To address these challenges, this paper proposes a World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework for space computing power networks. The offloading problem is formulated as a partially observable multi-agent decision-making process, where LEO satellites make decentralized decisions under incomplete local observations. A predictive world model learns latent transition dynamics of network states and provides future context for proactive planning. Meanwhile, a Transformer-based policy architecture captures inter-agent dependencies and supports cooperative scheduling under centralized training and decentralized execution (CTDE). Simulation results show that WM-MAPPO achieves higher task completion ratios, lower average latency, improved energy efficiency, and stronger robustness than model-free MARL baselines, heuristic methods, Lyapunov-based scheduling, and MINLP-inspired optimization.
Yuqi Cong, Zhiwei Wei, Jiarui Chen et al.· IEEE Transactions on Cogniti...· 0 citations
In edge–cloud collaborative computing, efficient scheduling of concurrent task chains is essential for reducing end‐to‐end latency. However, heterogeneous resources, dynamic link contention, and coupled routing‐computation decisions make it difficult to improve system efficiency while reducing communication conflicts. To address these challenges, we formulate concurrent heterogeneous task‐chain scheduling under partial observability as a multi‐agent partially observable Markov game and propose a MARL‐based scheduling method. Each task chain is represented by a mobile agent that makes routing, computation, data‐access, and waiting decisions based on local observations. Building upon a multi‐agent proximal policy optimization (MAPPO) architecture, we develop DC‐MAPPO, which introduces a Directional Clamp mechanism to improve policy update stability under high‐concurrency conditions. In addition, a dynamic action masking strategy is designed to ensure decision feasibility and reduce invalid exploration. Experimental results across multiple network topologies show that the proposed method outperforms the compared baselines in most tested settings. Under high‐load conditions, DC‐MAPPO reduces system makespan by 10.7%–17.1% and reduces the link reservation failure rate by 8.4–13.2 percentage points compared with MAPPO.
Xinyi Li, Chao Wang, Jiakai Liang et al.· Concurrency and Computation· 0 citations