Skip to content
Conference

Energy-aware deep reinforcement learning routing algorithm for space–air–ground–sea integrated networks

Jul 2026 · International Conference on Robotics and Sensor Networks · Vol 14254, pp. 1425418 - 1425418-7 · 0 citations · 11 references
Engineering

TL;DR

An Energy-Aware routing framework based on Proximal Policy Optimization (PPO) is proposed, where a One-Hot encoding mechanism is introduced to reconstruct the network state space, enabling the accurate capture of topological structural features.

Abstract

As a core component of future 6G architectures, the Space-Air-Ground-Sea Integrated Network (SAGS) is essential for marine environmental monitoring and emergency communications. However, constrained by scarce energy replenishment and the heterogeneous distribution of marine relay nodes, traditional shortest-path protocols often induce load imbalance and central node congestion, leading to premature failure and network connectivity loss. To address this "energy hole" problem, an Energy-Aware routing framework based on Proximal Policy Optimization (PPO) is proposed. Specifically, a One-Hot encoding mechanism is introduced to reconstruct the network state space, enabling the accurate capture of topological structural features. Furthermore, a composite reward function incorporating an energy penalty term is designed to guide routing decisions toward an optimal balance between path length and residual node energy. Experimental results in a high-fidelity simulation environment characterized by severe energy constraints demonstrate that the proposed algorithm effectively bypasses low-battery nodes while maintaining a 100% Packet Delivery Ratio. Notably, compared to Dijkstra’s algorithm, the proposed method significantly increases the average residual energy of network bottleneck nodes from 45.40% to 66.80%.

View source

Similar papers

Open access 2026

Reliable Low-Latency Task Offloading and Resource Allocation Method for Space-Air-Ground Integrated Networks

: Space-Air-Ground Integrated Networks (SAGIN) provide a multi-layered, wide-coverage computing infrastructure for distributed urban sensing systems. However, their heterogeneity and dynamics pose unprecedented challenges for task offloading and resource allocation. Existing methods struggle to simultaneously address the complexity of cross-layer decision-making and reliability assurance under uncertain conditions. This paper proposes a novel framework, termed DRL-RA, which synergistically integrates Deep Reinforcement Learning (DRL) with reliability-aware optimization. The framework consists of two complementary components: (1) a Dueling Double Deep Q-Network (D3QN) module that learns adaptive policies to make offloading decisions among various options including local execution, terrestrial edge, UAVs, and satellites; (2) a Reliability-Aware Multi-Objective Optimization Framework (RA-MOOF) that introduces explicit reliability guarantees through cross-layer link reliability modeling, node availability estimation, and smooth reliability proxy functions. Addressing the heterogeneous communication characteristics of the SAGIN architecture, this paper establishes a complete cross-layer delay model and composite reliability metrics. The reliability formulation is defined under explicitly stated conditional-independence assumptions, and the proposed smooth constraint terms are treated as surrogate CMDP costs rather than exact hard chance-constraint guarantees. Extensive experiments in a SAGIN simulation environment demonstrate that the proposed method improves the task completion rate by 3.8%, reduces average latency by 11.1%, and increases system reliability by 3.9% compared to state-of-the-art benchmarks. The optimization-only RA-Opt baseline is used as a non-real-time optimization reference for assessing reliability-aware offloading decision quality, while deployment-time decision-latency comparisons are interpreted primarily among learned inference policies. Comprehensive ablation studies and statistical validation across multiple random seeds confirm the contributions of each component, while cross-layer offloading decision analysis verifies the effectiveness of the method across different network layer selections.

Fei-Yan Bu, Zheng Wang, Yong Pan et al. · 0 citations
2026

Distributed Routing for LEO Satellite Networks: A Multi-Agent Deep Reinforcement Learning Approach With State Information Lag

Multi-agent deep reinforcement learning (MADRL) offers a promising solution for routing in low Earth orbit (LEO) satellite networks. However, large inter-satellite propagation delays lead to severe state information lag in agent interactions, giving rise to decision biases and degraded routing timeliness. To this end, this paper proposes a distributed routing algorithm named time-aware prediction and dynamic attention routing (TAP-DAR). Specifically, it constructs a delay compensation model that incorporates ephemeris data and queue prediction to generate near real-time neighbor state estimates. In addition, a multi-head attention fusion mechanism considering temporal reliability is designed to achieve adaptive aggregation of asynchronous neighbor states. Simulation results demonstrate that across various constellation configurations and network load conditions, the proposed algorithm achieves a maximum reduction of 16.16% in end-to-end (E2E) latency, an average decrease of nearly 30% in packet loss rate, and a maximum improvement of 19.41% in throughput compared to the baseline. Moreover, it substantially curtails communication overhead by more than 90% relative to the global state flooding mechanism.

Weidan Liu, Tong Liu, Li-Xia Xiao et al. · 0 citations
Open access 2026

Feasibility-Aware Reinforcement Learning for Reliable Hop-Constrained Routing in Wireless Sensor Networks

The experiments show that feasibility-aware learning can approach deterministic baseline reliability while retaining learned forwarding capability under hop constraints, and confirm that action masking is the dominant mechanism for maintaining feasible routing decisions, whereas trust mainly provides reliability-aware regularization.

Adeel Iqbal, Muhammad Faisal Siddiqui · 0 citations
Conference Jul 2026

Load-Balanced and Congestion-Aware Routing for LEO Laser Satellite Networks Based on Deep Reinforcement Learning

To address load imbalance in low earth orbit (LEO) laser satellite networks (LSN), this paper proposes a deep reinforcement learning (DRL) based routing algorithm, which combines proximal policy optimization (PPO) and K-shortest path (KSP) strategies to transform the large-scale routing problem into a decision-making process over a small set of paths. Simulation results demonstrate that, compared with traditional Dijkstra and random routing algorithms, the proposed algorithm fully exploits network resources, effectively prevents network bottlenecks, and significantly enhances the network’s service-carrying capacity.

Haoxin Li, Junling Yuan, Xu-Hong Li et al. · 0 citations
Review Open access Aug 2026

AI FOR GREEN 6G: A REVIEW OF ENERGY-AWARE ROUTING TECHNIQUES

It is concluded that AI is not merely an enabler but a cornerstone for realizing sustainable and intelligent 6G networks, paving the way for an eco-friendly digital future.

Joshna M, R. K. · 0 citations
Open access Aug 2026

Quantum Federated Reinforcement Learning‐Based Traffic Offloading and Resource Allocation for RSMA‐Enabled Space–Air–Ground Integrated Networks

A Quantum Federated Reinforcement Learning (QFRL)‐based traffic offloading framework for RSMA‐enabled SAGINs is proposed, allowing distributed small cells to jointly optimize traffic offloading ratios, bandwidth allocation, RSMA power distribution, and UAV trajectory planning while satisfying stringent delay and reliability requirements.

Ishan Budhiraja, Abhay Bansal, B. Unhelkar et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.