2026· IEEE Open Journal of the Communications Society· Vol 7, pp. 7770-7789· 0 citations· 34 references
Computer Science
TL;DR
A hybrid LEO-terrestrial architecture that integrates Software-Defined Networking ground station clusters with repeater-assisted reception and a structured fallback mechanism is proposed and results highlight the effectiveness of the proposed framework in enabling adaptive and delay-efficient control in next-generation LEO satellite systems.
Abstract
Low Earth Orbit (LEO) satellite constellations enable global low-latency connectivity but face challenges due to weather-dependent link variability, orbital dynamics, and heterogeneous ground infrastructure. In this paper, we propose a hybrid LEO-terrestrial architecture that integrates Software-Defined Networking (SDN) ground station clusters with repeater-assisted reception and a structured fallback mechanism. We develop a stochastic model that captures weather-driven reliability, correlated repeater behavior, and multi-path reception, and formulate an optimization problem to minimize the expected communication delay. To address the intractability of this problem in large-scale, partially observable environments, we design a fully decentralized Multi-Agent Reinforcement Learning (MARL) framework based on Proximal Policy Optimization (PPO), where each satellite makes decisions using only local observations. The reward is aligned with the analytical delay objective, ensuring consistency between the model and learned policies. Simulation results across diverse scenarios demonstrate that the proposed approach reduces the mean delay by 40-60% and significantly decreases fallback usage compared to baseline methods. These results highlight the effectiveness of the proposed framework in enabling adaptive and delay-efficient control in next-generation LEO satellite systems.
Maritime satellite communications (SATCOMs) are expected to support high-capacity ship-to-satellite uplinks for remote maritime services beyond terrestrial coverage, with low-Earth-orbit (LEO) satellites providing wide-area connectivity. However, robust uplink beamforming in LEO maritime SATCOMs is challenging because dynamic ship–satellite geometry, wave-induced attitude motion, imperfect channel state information, and multiship interference make transmit power, ship-side transmit beamforming, and satellite-side receive combining tightly coupled. Accordingly, we formulate a long-term spectral efficiency (SE) maximization problem under transmit-power and quality-of-service constraints. An attitude-aware uplink channel model is developed by incorporating roll, pitch, and yaw motions into the effective angle-of-departure/angle-of-arrival evolution. Based on this model, the problem is cast as a heterogeneous decentralized partially observable Markov decision process. We then propose a robust heterogeneous cooperative QMIX (RHC-QMIX) framework under centralized training and decentralized execution, where type-specific recurrent local Q-networks, history-refined angular features, and centralized monotonic value mixing coordinate ship and satellite agents. Extensive simulations demonstrate that in the load-controlled scalability evaluation, RHC-QMIX achieves an average network SE of 25.32 bps/Hz, improves over alternating optimization by up to 51.00% as the satellite load increases, and outperforms heterogeneous cooperative QMIX by 16.38% on average under network-size scaling; it also maintains more stable SE under severe sea-state-induced ship motion.
Multi-agent deep reinforcement learning (MADRL) offers a promising solution for routing in low Earth orbit (LEO) satellite networks. However, large inter-satellite propagation delays lead to severe state information lag in agent interactions, giving rise to decision biases and degraded routing timeliness. To this end, this paper proposes a distributed routing algorithm named time-aware prediction and dynamic attention routing (TAP-DAR). Specifically, it constructs a delay compensation model that incorporates ephemeris data and queue prediction to generate near real-time neighbor state estimates. In addition, a multi-head attention fusion mechanism considering temporal reliability is designed to achieve adaptive aggregation of asynchronous neighbor states. Simulation results demonstrate that across various constellation configurations and network load conditions, the proposed algorithm achieves a maximum reduction of 16.16% in end-to-end (E2E) latency, an average decrease of nearly 30% in packet loss rate, and a maximum improvement of 19.41% in throughput compared to the baseline. Moreover, it substantially curtails communication overhead by more than 90% relative to the global state flooding mechanism.
Weidan Liu, Tong Liu, Li-Xia Xiao et al.· IEEE Transactions on Cogniti...· 0 citations
The integration of Low Earth Orbit (LEO) Non-Terrestrial Networks (NTNs) into 5G and upcoming 6G architectures introduces various challenges, including severe propagation delays, ultra-high base station mobility, and channel non-stationarity, complicating radio resource management of heterogeneous network slices. In this paper, we propose a deep reinforcement learning (DRL) meta-scheduler for twin-timescale resource allocation. Our solution adopts a decoupled Open Radio Access Network (RAN) architecture, in which a strategic 100 ms meta-scheduler selects scheduling policies for the different network slices using stale telemetry, while a fast-timescale MAC packet scheduler processes per-TTI user requests. The resulting Markov Decision Process captures non-stationary orbital dynamics and heterogeneous SLAs constraints via a TD3 agent. Simulation results under varying traffic load show that, unlike other solutions, the proposed meta-scheduler explicitly trades a statistically insignificant 1% capacity fraction (p>0.05) to strictly bound the variance and overall magnitude of RLC-layer queuing delay for Mission-Critical (MC) traffic. Crucially, it enforces this isolation without inducing the broadband slice starvation characteristic of standard maximum-CQI heuristics, establishing a robust foundation for 6G O-RAN NTN resource allocation.
Víctor Vilchez, T. P. C. de Andrade, Edward Hinojosa et al.· 0 citations
This work proposes a structured multi-agent reinforcement learning (MARL) framework based on multi-agent proximal policy optimization (MAPPO), termed ShellMean-MAPPO, for downlink resource allocation with explicit conflict resolution, and demonstrates its advantages over representative MARL schemes in terms of scheduling performance and conflict mitigation.
Li Zhen, Qi-Hao Zhang, Qing-Zhi Meng et al.· IEEE Open Journal of the Com...· 0 citations
The proposed multi-agent reinforcement learning policy attains slightly higher throughput with fewer handovers by offloading a fraction of the users to the MEO and GEO layers, an emergent multi-orbit behavior that drives its favorable throughput and handover trade-off.
Yassine Afif, Ashutosh Balakrishnan, Philippe Martins et al.· 0 citations
: Space-Air-Ground Integrated Networks (SAGIN) provide a multi-layered, wide-coverage computing infrastructure for distributed urban sensing systems. However, their heterogeneity and dynamics pose unprecedented challenges for task offloading and resource allocation. Existing methods struggle to simultaneously address the complexity of cross-layer decision-making and reliability assurance under uncertain conditions. This paper proposes a novel framework, termed DRL-RA, which synergistically integrates Deep Reinforcement Learning (DRL) with reliability-aware optimization. The framework consists of two complementary components: (1) a Dueling Double Deep Q-Network (D3QN) module that learns adaptive policies to make offloading decisions among various options including local execution, terrestrial edge, UAVs, and satellites; (2) a Reliability-Aware Multi-Objective Optimization Framework (RA-MOOF) that introduces explicit reliability guarantees through cross-layer link reliability modeling, node availability estimation, and smooth reliability proxy functions. Addressing the heterogeneous communication characteristics of the SAGIN architecture, this paper establishes a complete cross-layer delay model and composite reliability metrics. The reliability formulation is defined under explicitly stated conditional-independence assumptions, and the proposed smooth constraint terms are treated as surrogate CMDP costs rather than exact hard chance-constraint guarantees. Extensive experiments in a SAGIN simulation environment demonstrate that the proposed method improves the task completion rate by 3.8%, reduces average latency by 11.1%, and increases system reliability by 3.9% compared to state-of-the-art benchmarks. The optimization-only RA-Opt baseline is used as a non-real-time optimization reference for assessing reliability-aware offloading decision quality, while deployment-time decision-latency comparisons are interpreted primarily among learned inference policies. Comprehensive ablation studies and statistical validation across multiple random seeds confirm the contributions of each component, while cross-layer offloading decision analysis verifies the effectiveness of the method across different network layer selections.
Fei-Yan Bu, Zheng Wang, Yong Pan et al.· Computers, Materials & C...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.