Skip to content
Conference

Carbon-Aware Traffic Steering for TN-NTN Scenario via Graph Attention Reinforcement Learning

Jul 2026 · International Conference on Ubiquitous and Future Networks · pp. 1137-1139 · 0 citations · 4 references

Abstract

Sixth-generation (6G) wireless systems envision seamless coexistence between terrestrial networks (TNs) and non-terrestrial networks (NTNs), while the carbon footprint of dense radio access infrastructure has become a critical concern. This paper proposes carbon-aware traffic steering (CATS), a Near-RT RIC xApp for the open RAN (O-RAN) architecture. CATS introduces a network-wide virtual carbon queue into the global observation of a learning-based steering policy. The policy is implemented using a graph attention network (GAT) with type-aware attention to capture heterogeneous user equipment (UE) services and is trained via proximal policy optimization (PPO). Simulation results show that CATS significantly reduces net carbon emissions while preserving quality-of-service (QoS) satisfaction.

View source

Similar papers

Review Open access 2026

Comprehensive Review of Optimization Techniques for User-Centric Distributed Network Slicing in 5G Networks

A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.

Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang · 0 citations
Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations
Preprint Aug 2026

LEO-Aware DRL Meta-Scheduler for 5G Non-Terrestrial Network Slicing

The integration of Low Earth Orbit (LEO) Non-Terrestrial Networks (NTNs) into 5G and upcoming 6G architectures introduces various challenges, including severe propagation delays, ultra-high base station mobility, and channel non-stationarity, complicating radio resource management of heterogeneous network slices. In this paper, we propose a deep reinforcement learning (DRL) meta-scheduler for twin-timescale resource allocation. Our solution adopts a decoupled Open Radio Access Network (RAN) architecture, in which a strategic 100 ms meta-scheduler selects scheduling policies for the different network slices using stale telemetry, while a fast-timescale MAC packet scheduler processes per-TTI user requests. The resulting Markov Decision Process captures non-stationary orbital dynamics and heterogeneous SLAs constraints via a TD3 agent. Simulation results under varying traffic load show that, unlike other solutions, the proposed meta-scheduler explicitly trades a statistically insignificant 1% capacity fraction (p>0.05) to strictly bound the variance and overall magnitude of RLC-layer queuing delay for Mission-Critical (MC) traffic. Crucially, it enforces this isolation without inducing the broadband slice starvation characteristic of standard maximum-CQI heuristics, establishing a robust foundation for 6G O-RAN NTN resource allocation.

Víctor Vilchez, T. P. C. de Andrade, Edward Hinojosa et al. · 0 citations
Conference Jul 2026

Load-Balanced and Congestion-Aware Routing for LEO Laser Satellite Networks Based on Deep Reinforcement Learning

To address load imbalance in low earth orbit (LEO) laser satellite networks (LSN), this paper proposes a deep reinforcement learning (DRL) based routing algorithm, which combines proximal policy optimization (PPO) and K-shortest path (KSP) strategies to transform the large-scale routing problem into a decision-making process over a small set of paths. Simulation results demonstrate that, compared with traditional Dijkstra and random routing algorithms, the proposed algorithm fully exploits network resources, effectively prevents network bottlenecks, and significantly enhances the network’s service-carrying capacity.

Haoxin Li, Junling Yuan, Xu-Hong Li et al. · 0 citations
Conference Open access Jul 2026

Graph-Centric Deep Q-Learning for Interference-Aware Resource Allocation in Rsma-Enabled 5G Slicing

The emergence of 5G and 6G advanced ecosystems demands highly adaptive resource management to orchestrate the specialised requirements of eMBB, URLLC, and mMTC network slices. In dense multi-cell environments, capturing complex spatial interdependencies and mitigating dynamic interference is paramount for maintaining Quality of Service (QoS). This paper introduces a robust GNN-DQN framework designed for Rate Splitting Multiple Access (RSMA) based networks. By representing the network topology as a graph, the framework leverages Graph Neural Networks (GNNs) to extract highdimensional spatial features and model inter-cell interference patterns. These insights enable a Deep Q-Network (DQN) agent to perform intelligent resource partitioning and dynamic power splitting of the RSMA common stream. Experimental results demonstrate that the proposed GNN-DQN framework achieves a connectivity success ratio exceeding 90% across all slices, representing an average improvement of over 60% compared to non-graph-based reinforcement learning and supervised baselines. Notably, the framework demonstrates exceptional spectral efficiency, maintaining near-total connectivity while utilising less than 10% of the normalised system bandwidth, a 4× reduction in resource overhead compared to traditional methods. Furthermore, the GNN-driven architecture ensures stable convergence during training, yielding a 1.6× higher system reward score. Our findings validate GNN-DQN as a high-performance, scalable, and resource-efficient paradigm for intelligent orchestration in 5G and 6G networks.

Aya Kh. Ahmed, Nadia Al-Aboody, Hamed S. Al-Raweshidy · 0 citations
Jul 2026

Adaptive Offloading Control in 6G Cell-Free O-RAN Using Reinforcement Learning

Open Radio Access Networks (O-RAN) enable programmable radio access architectures in which learning-based intelligence at the RAN Intelligent Controller (RIC) supports adaptive resource management in future 6G systems. To address the dynamic and heterogeneous traffic conditions observed at the regional edge, we propose a Reinforcement Learning (RL)–driven offloading framework to adaptively offload wireless services between SDN-controlled network segments. Simulation results demonstrate that the proposed approach consistently outperforms a threshold-based heuristic baseline, achieving an average reduction of approximately 34% in overall blocking across different traffic loads. These results convey the effectiveness of learning-based offloading control for improving service performance in dynamic 6G network environments.

Pouya Mehdizadeh, Golshan Famitafreshi, John S. Vardakas et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.