Skip to content
Open access

Scalable Reinforcement Learning-Based Environment-Aware Scheduler for Enhancing Reliability and Availability in 6G Networks

2026 · IEEE Transactions on Machine Learning in Communications and Networking · Vol 4, pp. 1177-1198 · 0 citations · 26 references

TL;DR

A scalable deep reinforcement learning (DRL) framework that exploits environmental-aware knowledge to optimize multi-user scheduling under limited resources, in order to enhance reliability and availability while maintaining fairness, outperforming Round Robin and Proportional Fair schedulers.

Abstract

The sixth generation (6G) of wireless networks must ensure high reliability, availability, and fairness, even in dynamic and challenging environments. Environmental-aware knowledge, such as that obtained from Radio Environment Maps (REMs), offers predictive insights into channel conditions and can guide more effective scheduling decisions. This paper proposes a scalable deep reinforcement learning (DRL) framework that exploits such knowledge to optimize multi-user scheduling under limited resources, in order to enhance reliability and availability while maintaining fairness. Unlike standard deep Q-network (DQN), which evaluates Q-values per action, we propose a novel learning method that estimates per-user Q-values, enabling user selection with per-decision complexity that scales linearly with the number of users. Simulation results show that the proposed approach consistently balances reliability and availability while maintaining fairness, outperforming Round Robin (RR) and Proportional Fair (PF) schedulers, especially in environments with high clutter density and frequent non-line-of-sight conditions. In particular, the proposed method improves reliability by over 400% with only a 2% drop in availability, compared to the RR scheduler. Relative to PF, it improves availability and fairness by 23% and 40%, respectively, without sacrificing reliability. These improvements are observed under balanced propagation conditions, where the probabilities of line-of-sight and non-line-of-sight are equal due to a clutter density of 50%. These results highlight the potential of the proposed environmental-aware DRL scheduler to support trustworthy 6G communication in complex and dynamic environments.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

From Prior-Guided Heuristics to Deployable Agents: Accelerating Demonstration-Driven Reinforcement Learning for Deadline-Constrained Network Control

Timely delivery of delay-sensitive information over dynamic, heterogeneous networks is essential for NextG interactive applications, yet providing strict End-to-End (E2E) peak latency guarantees remains an open challenge. Two obstacles limit the adoption of learning-based network control in this setting: traditional volume-based routing metrics, while highly effective for general traffic management, are not designed to capture traffic urgency; and Deep Reinforcement Learning (DRL) controllers trained from scratch suffer from sample inefficiency, long training times, and early-stage exploration volatility. This paper introduces a deployment-focused network control framework that addresses both obstacles. First, we present Effective Congestion (EC), a deadline-aware metric family that quantifies interface congestion by packet urgency and proactively filters non-viable traffic, coupled with a Uniform Path Grouping (UPG) distribution heuristic promoting robust load-balancing; the resulting policies are embedded into Multi-Agent Deep Reinforcement Learning Effective Congestion ($p^*$) (MADRL EC ($p^*$)), a hybrid architecture combining a distributed scheduler with a centralized RL-based router. Second, we introduce a unified training objective that generalizes existing policy-learning paradigms---behavioral cloning, offline Reinforcement Learning (RL), online RL, and offline-to-online schemes---as special cases, combining a live-reward term, a pre-collected-reward term, and a policy-imitation term. From this objective, we derive the Model-Guided Annealed Reinforcement Learning (MGA-RL) protocol, instantiated on a Deep Deterministic Policy Gradient (DDPG) backbone: a deployment-oriented, demonstration-driven training approach that generalizes conventional Offline-to-Online (O2O) schemes, in which trajectories from a lightweight [...]

Vincenzo Norman Vitale, Mohammad Solki, A. Tulino et al. · 0 citations
Open access Jul 2026

Robust Offline Multi-Agent Reinforcement Learning for Latency-Aware SDN Path Control in 6G-Oriented Network Softwarization

Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.

A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al. · 0 citations
Open access 2026

Reliable Low-Latency Task Offloading and Resource Allocation Method for Space-Air-Ground Integrated Networks

: Space-Air-Ground Integrated Networks (SAGIN) provide a multi-layered, wide-coverage computing infrastructure for distributed urban sensing systems. However, their heterogeneity and dynamics pose unprecedented challenges for task offloading and resource allocation. Existing methods struggle to simultaneously address the complexity of cross-layer decision-making and reliability assurance under uncertain conditions. This paper proposes a novel framework, termed DRL-RA, which synergistically integrates Deep Reinforcement Learning (DRL) with reliability-aware optimization. The framework consists of two complementary components: (1) a Dueling Double Deep Q-Network (D3QN) module that learns adaptive policies to make offloading decisions among various options including local execution, terrestrial edge, UAVs, and satellites; (2) a Reliability-Aware Multi-Objective Optimization Framework (RA-MOOF) that introduces explicit reliability guarantees through cross-layer link reliability modeling, node availability estimation, and smooth reliability proxy functions. Addressing the heterogeneous communication characteristics of the SAGIN architecture, this paper establishes a complete cross-layer delay model and composite reliability metrics. The reliability formulation is defined under explicitly stated conditional-independence assumptions, and the proposed smooth constraint terms are treated as surrogate CMDP costs rather than exact hard chance-constraint guarantees. Extensive experiments in a SAGIN simulation environment demonstrate that the proposed method improves the task completion rate by 3.8%, reduces average latency by 11.1%, and increases system reliability by 3.9% compared to state-of-the-art benchmarks. The optimization-only RA-Opt baseline is used as a non-real-time optimization reference for assessing reliability-aware offloading decision quality, while deployment-time decision-latency comparisons are interpreted primarily among learned inference policies. Comprehensive ablation studies and statistical validation across multiple random seeds confirm the contributions of each component, while cross-layer offloading decision analysis verifies the effectiveness of the method across different network layer selections.

Fei-Yan Bu, Zheng Wang, Yong Pan et al. · 0 citations
Preprint Aug 2026

Deep Reinforcement Learning Orchestration of Game-Theoretic User Association and Resource Allocation in HetNets

A novel orchestration scheme for game-theoretic UARA in HetNets that closely approximates the optimal policy for the considered operational objectives, while delivering higher network throughput than conventional association methods.

Sotiris Kopsinos, Alexandros I. Papadopoulos, Antonios Lalas et al. · 0 citations
Open access Sep 2026

Reinforcement Learning-Based Adaptive Data Rate Control in Mobile LoRaWAN with Imperfect CSI

Low Power Wide Area Networks (LPWANs), particularly LoRaWAN, are increasingly being deployed in mobile Internet of Things (IoT) applications. In these applications, adaptive data rate (ADR) control is essential for maintaining reliable and energy-efficient communication under dynamic wireless conditions. However, ADR performance remains constrained when channel state information (CSI) is imperfect. Although reinforcement learning (RL)-based ADR methods have demonstrated notable improvement over conventional approaches, their effectiveness under imperfect CSI deployment scenarios remains insufficiently explored. In this paper, we investigate the robustness of RL-based ADR approaches with imperfect CSI for mobile LoRaWAN networks using a mobility-aware simulation framework. Standard ADR, heuristic ADR, Deep Q-Network (DQN)-based ADR, and Proximal Policy Optimisation (PPO)-based ADR are evaluated across varying mobility patterns, network layouts, and channel dynamics. Key performance metrics include packet delivery ratio, latency, runtime efficiency, and retention of baseline capability are evaluated. The Results show that PPO achieves the best performance under imperfect CSI conditions, attaining a mean packet delivery ratio of 0.9954, outperforming standard ADR, heuristic ADR, and DQN by 29.49%, 19.12%, and 1.55%, respectively. In addition, PPO reduced latency by 53.12% and runtime by 81.94% compare to DQN while preserving over 99% of baseline communication reliability under scenario variation. These findings demonstrate that policy-gradient learning provides better transferability, robustness, and deployment readiness compared with conventional and value-based ADR approaches. The study highlights the importance of evaluating intelligent wireless controllers not only for optimisation performance, but also for cross-scenario generalisation before real-world deployment in mobile LPWAN systems.

Unknown authors · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.