Skip to content
Open access

ShellMean-MAPPO: A Conflict-Aware MARL Framework for Downlink Resource Allocation in Multi-Shell LEO Satellite Networks

2026 · IEEE Open Journal of the Communications Society · Vol 7, pp. 10193-10213 · 0 citations · 55 references

TL;DR

This work proposes a structured multi-agent reinforcement learning (MARL) framework based on multi-agent proximal policy optimization (MAPPO), termed ShellMean-MAPPO, for downlink resource allocation with explicit conflict resolution, and demonstrates its advantages over representative MARL schemes in terms of scheduling performance and conflict mitigation.

Abstract

Low Earth orbit (LEO) satellite communication, able to provide ubiquitous and continuous connectivity, has become a vital component of future sixth-generation global networks. To further improve service continuity and support spatially non-uniform traffic demand, LEO satellite systems are evolving toward dense multi-shell constellations, where overlapping coverage enables multiple satellites to serve the same traffic region. However, under limited onboard beams, spectrum, and power budgets, such overlap may lead multiple beams to request the same traffic cell, causing duplicate beam-cell requests (DBRs) in downlink scheduling. To address this challenge, we propose a structured multi-agent reinforcement learning (MARL) framework based on multi-agent proximal policy optimization (MAPPO), termed ShellMean-MAPPO, for downlink resource allocation with explicit conflict resolution. Specifically, we first formulate a long-term scheduling problem that separates pre-resolution beam-cell requests from post-resolution retained transmissions, enabling DBRs to be modeled together with resource block (RB) chunk and power allocation. By leveraging compact local information and per-shell summaries, a typed encoder is then designed to capture heterogeneous service and contention states without relying on a full global map. Furthermore, an autoregressive policy is developed to generate beam-cell, RB chunk, and power-share decisions in accordance with the downlink scheduling sequence, while a deterministic replay-based post-resolution credit mechanism transforms team outcomes into per-beam training signals. Extensive simulation results validate the effectiveness of ShellMean-MAPPO, and demonstrate its advantages over representative MARL schemes in terms of scheduling performance and conflict mitigation.

Read PDF

Similar papers

Preprint Aug 2026

Multi-Agent Reinforcement Learning for Joint Handover Management and Power Allocation in Multi-Orbit Satellite Networks

The proposed multi-agent reinforcement learning policy attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.

Yassine Afif, Ashutosh Balakrishnan, Philippe Martins et al. · 0 citations
2026

A Heterogeneous Multiagent Reinforcement Learning Approach for Robust Uplink Beamforming in Maritime Satellite Communications

Maritime satellite communications (SATCOMs) are expected to support high-capacity ship-to-satellite uplinks for remote maritime services beyond terrestrial coverage, with low-Earth-orbit (LEO) satellites providing wide-area connectivity. However, robust uplink beamforming in LEO maritime SATCOMs is challenging because dynamic ship–satellite geometry, wave-induced attitude motion, imperfect channel state information, and multiship interference make transmit power, ship-side transmit beamforming, and satellite-side receive combining tightly coupled. Accordingly, we formulate a long-term spectral efficiency (SE) maximization problem under transmit-power and quality-of-service constraints. An attitude-aware uplink channel model is developed by incorporating roll, pitch, and yaw motions into the effective angle-of-departure/angle-of-arrival evolution. Based on this model, the problem is cast as a heterogeneous decentralized partially observable Markov decision process. We then propose a robust heterogeneous cooperative QMIX (RHC-QMIX) framework under centralized training and decentralized execution, where type-specific recurrent local Q-networks, history-refined angular features, and centralized monotonic value mixing coordinate ship and satellite agents. Extensive simulations demonstrate that in the load-controlled scalability evaluation, RHC-QMIX achieves an average network SE of 25.32 bps/Hz, improves over alternating optimization by up to 51.00% as the satellite load increases, and outperforms heterogeneous cooperative QMIX by 16.38% on average under network-size scaling; it also maintains more stable SE under severe sea-state-induced ship motion.

Huayuan Wang, Bodong Shang, Meixia Tao · 0 citations
2026

Distributed Routing for LEO Satellite Networks: A Multi-Agent Deep Reinforcement Learning Approach With State Information Lag

Multi-agent deep reinforcement learning (MADRL) offers a promising solution for routing in low Earth orbit (LEO) satellite networks. However, large inter-satellite propagation delays lead to severe state information lag in agent interactions, giving rise to decision biases and degraded routing timeliness. To this end, this paper proposes a distributed routing algorithm named time-aware prediction and dynamic attention routing (TAP-DAR). Specifically, it constructs a delay compensation model that incorporates ephemeris data and queue prediction to generate near real-time neighbor state estimates. In addition, a multi-head attention fusion mechanism considering temporal reliability is designed to achieve adaptive aggregation of asynchronous neighbor states. Simulation results demonstrate that across various constellation configurations and network load conditions, the proposed algorithm achieves a maximum reduction of 16.16% in end-to-end (E2E) latency, an average decrease of nearly 30% in packet loss rate, and a maximum improvement of 19.41% in throughput compared to the baseline. Moreover, it substantially curtails communication overhead by more than 90% relative to the global state flooding mechanism.

Weidan Liu, Tong Liu, Li-Xia Xiao et al. · 0 citations
2026

Collaborative Task Offloading in Space Computing Power Network: A World Model-Based Multi-Agent Reinforcement Learning Approach

Low Earth Orbit (LEO) constellations are required to process increasing volumes of heterogeneous tasks from ground networks. Intermittent inter-satellite links, heterogeneous onboard resources, and time-varying traffic loads make collaborative task offloading difficult for static or reactive strategies. Although multi-agent reinforcement learning (MARL) provides an adaptive solution, existing model-free MARL methods often suffer from slow convergence, insufficient foresight, and limited robustness in dynamic satellite environments. To address these challenges, this paper proposes a World Model-Based Multi-Agent Proximal Policy Optimization (WM-MAPPO) framework for space computing power networks. The offloading problem is formulated as a partially observable multi-agent decision-making process, where LEO satellites make decentralized decisions under incomplete local observations. A predictive world model learns latent transition dynamics of network states and provides future context for proactive planning. Meanwhile, a Transformer-based policy architecture captures inter-agent dependencies and supports cooperative scheduling under centralized training and decentralized execution (CTDE). Simulation results show that WM-MAPPO achieves higher task completion ratios, lower average latency, improved energy efficiency, and stronger robustness than model-free MARL baselines, heuristic methods, Lyapunov-based scheduling, and MINLP-inspired optimization.

Yuqi Cong, Zhiwei Wei, Jiarui Chen et al. · 0 citations
2026

Distributed Cooperative Beamforming for Spectrum Sharing in GEO-LEO Heterogeneous Multi-Satellite System

Due to their resilience and global coverage, satellite networks are poised to become a key component for non-terrestrial networks in the future. However, given the scarcity of spectrum resources, the dense deployment of low Earth orbit (LEO) satellites introduces significant interference challenges. Meanwhile, the limited computing power and backhaul capacity of satellites have become bottlenecks hindering the development of advanced interference mitigation techniques. This paper studies beamforming in GEO-LEO heterogeneous multi-satellite systems. For the GEO system, we develop a multicast beamforming approach based on a nonlinear eigenvalue problem (NEPv) for beam direction design and Lagrange dual decomposition (LDD) for power allocation. For the LEO system, we propose a general distributed beamforming framework and two distributed beamforming methods. Specifically, we first leverage equivalent multi-dimensional fractional programming (FP) to decompose the objective function. The resulting subproblems are then optimized in a distributed manner across multiple satellites via the parallel block coordinate descent (PBCD) method. For the distributed optimization subproblems, we derive semi-closed-form solutions using Lagrangian dual ascent (LDA) and alternating direction method of multipliers (ADMM) for scenarios without and with GEO-LEO interference avoidance, respectively. Simulation results show that the proposed NEPv-LDD method strictly satisfies the QoS constraints of users and achieves near-optimal performance with low complexity. For the LEO beamforming, the developed distributed FP (DiFP) framework exhibits strong scalability in large-scale constellations. Built upon the DiFP framework, the proposed DiFP-NoSIA incurs almost no performance loss, while DiFP-ADMM shows only an 8.58% performance degradation compared to the centralized benchmark.

Xin Chen, Zhiyong Luo · 0 citations
Preprint Aug 2026

LEO-Aware DRL Meta-Scheduler for 5G Non-Terrestrial Network Slicing

The integration of Low Earth Orbit (LEO) Non-Terrestrial Networks (NTNs) into 5G and upcoming 6G architectures introduces various challenges, including severe propagation delays, ultra-high base station mobility, and channel non-stationarity, complicating radio resource management of heterogeneous network slices. In this paper, we propose a deep reinforcement learning (DRL) meta-scheduler for twin-timescale resource allocation. Our solution adopts a decoupled Open Radio Access Network (RAN) architecture, in which a strategic 100 ms meta-scheduler selects scheduling policies for the different network slices using stale telemetry, while a fast-timescale MAC packet scheduler processes per-TTI user requests. The resulting Markov Decision Process captures non-stationary orbital dynamics and heterogeneous SLAs constraints via a TD3 agent. Simulation results under varying traffic load show that, unlike other solutions, the proposed meta-scheduler explicitly trades a statistically insignificant 1% capacity fraction (p>0.05) to strictly bound the variance and overall magnitude of RLC-layer queuing delay for Mission-Critical (MC) traffic. Crucially, it enforces this isolation without inducing the broadband slice starvation characteristic of standard maximum-CQI heuristics, establishing a robust foundation for 6G O-RAN NTN resource allocation.

Víctor Vilchez, T. P. C. de Andrade, Edward Hinojosa et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.