Skip to content

Distributed Routing for LEO Satellite Networks: A Multi-Agent Deep Reinforcement Learning Approach With State Information Lag

2026 · IEEE Transactions on Cognitive Communications and Networking · Vol 12, pp. 11116-11130 · 0 citations · 41 references
Computer Science

Abstract

Multi-agent deep reinforcement learning (MADRL) offers a promising solution for routing in low Earth orbit (LEO) satellite networks. However, large inter-satellite propagation delays lead to severe state information lag in agent interactions, giving rise to decision biases and degraded routing timeliness. To this end, this paper proposes a distributed routing algorithm named time-aware prediction and dynamic attention routing (TAP-DAR). Specifically, it constructs a delay compensation model that incorporates ephemeris data and queue prediction to generate near real-time neighbor state estimates. In addition, a multi-head attention fusion mechanism considering temporal reliability is designed to achieve adaptive aggregation of asynchronous neighbor states. Simulation results demonstrate that across various constellation configurations and network load conditions, the proposed algorithm achieves a maximum reduction of 16.16% in end-to-end (E2E) latency, an average decrease of nearly 30% in packet loss rate, and a maximum improvement of 19.41% in throughput compared to the baseline. Moreover, it substantially curtails communication overhead by more than 90% relative to the global state flooding mechanism.

View source

Similar papers

Open access Jul 2026

Deep Reinforcement Learning-Based QoS-Aware Routing Protocol for Space–Air–Ground Integrated Networks

A deep reinforcement learning (DRL)-based adaptive routing scheme for maximizing throughput and minimizing end-to-end delay jointly in SAGIN and indicates that adaptive policy learning enables better congestion avoidance and more efficient resource utilization.

Nilu Mishra, Sanakat Bhanjan Prusty, Sachin Sharma · 0 citations
Conference Jul 2026

Load-Balanced and Congestion-Aware Routing for LEO Laser Satellite Networks Based on Deep Reinforcement Learning

To address load imbalance in low earth orbit (LEO) laser satellite networks (LSN), this paper proposes a deep reinforcement learning (DRL) based routing algorithm, which combines proximal policy optimization (PPO) and K-shortest path (KSP) strategies to transform the large-scale routing problem into a decision-making process over a small set of paths. Simulation results demonstrate that, compared with traditional Dijkstra and random routing algorithms, the proposed algorithm fully exploits network resources, effectively prevents network bottlenecks, and significantly enhances the network’s service-carrying capacity.

Haoxin Li, Junling Yuan, Xu-Hong Li et al. · 0 citations
Preprint Aug 2026

LLM-Driven Automated Reward Design for Reinforcement Learning-Based Routing in LEO Satellite Networks

Routing in Low Earth Orbit (LEO) satellite networks is challenging due to highly dynamic topologies and spatio-temporal network conditions. Reinforcement Learning (RL) has emerged as a promising approach for adaptive routing; however, its performance critically depends on reward function design, which must balance objectives such as goodput and end-to-end delay. In practice, reward design remains a complex manual process requiring significant domain expertise and extensive trial-and-error. Recent works have explored Large Language Models (LLMs) for automated reward design, but their application to highly dynamic systems such as LEO satellite networks remains largely unexplored. We propose LARGE, a framework that automates reward design for RL-based routing by combining LLM- driven generation with iterative simulator-in-the-loop evaluation. LARGE generates an initial reward from LLM prior knowledge and iteratively refines it using simulation feedback. This loop enables exploration of diverse reward formulations while aligning them with network objectives. Results show that LARGE improves reward quality within a few iterations through feedback-driven refinement. Across different backbones, the framework achieves performance comparable to an expert-designed baseline, with the best-performing configuration reaching goodput within approximately 3% of the baseline and slightly lower end-to-end delay, without manual reward engineering. These results indicate that effectiveness emerges from the iterative feedback-driven process enabled by LARGE, highlighting the potential of framework-driven LLM-in-the-loop optimization for RL-based routing in dynamic satellite networks.

W. P. Casas, N. Fonseca, and Carlos A. Astudillo · 0 citations
Conference Jul 2026

RoutePPO: eBPF-Based Proximal Policy Optimization for Adaptive Routing in UAV Swarm Networks

Unmanned Aerial Vehicle (UAV) swarm networks demand routing protocols that adapt continuously to rapid topology changes, node mobility, and fluctuating link quality. AODV may incur route-discovery overhead after topology changes, while OLSR relies on periodic topology dissemination that may lag behind fast link-quality changes; both can struggle under UAV swarm dynamics. We present RoutePPO, a closedloop adaptive routing framework that couples Proximal Policy Optimization (PPO) with eBPF-based real-time link telemetry and a P4 programmable data plane. RoutePPO-Adapt introduces a Top-K path encoder with fixed-order slot assignment and a 3-step slot-history observation, producing a topology-agnostic 30-dimensional state representation. Training uses a 9-scenario curriculum with anticipatory reward shaping and cosine learningrate decay. Across 15 deterministic routing scenarios, RoutePPOAdapt achieves a mean reward of 0.674–9.2% above the two-path baseline (RoutePPO-Base) - winning 11 of 15 scenarios while reducing latency by 31.6% and packet loss by 37.1%. A kernel-native evaluation (Linux netns + eBPF TC egress) confirms non-zero telemetry counters (0.13-0.27 Mbps), demonstrating end-to-end viability of the eBPF-PPO pipeline.

Nazım Cürmen, F. Okay, Suat Özdemir · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.