Jul 2026· International Conference on Computer, Information and Telecommunication Systems· pp. 1-8· 0 citations· 21 references
Abstract
The emergence of 5G and 6G advanced ecosystems demands highly adaptive resource management to orchestrate the specialised requirements of eMBB, URLLC, and mMTC network slices. In dense multi-cell environments, capturing complex spatial interdependencies and mitigating dynamic interference is paramount for maintaining Quality of Service (QoS). This paper introduces a robust GNN-DQN framework designed for Rate Splitting Multiple Access (RSMA) based networks. By representing the network topology as a graph, the framework leverages Graph Neural Networks (GNNs) to extract highdimensional spatial features and model inter-cell interference patterns. These insights enable a Deep Q-Network (DQN) agent to perform intelligent resource partitioning and dynamic power splitting of the RSMA common stream. Experimental results demonstrate that the proposed GNN-DQN framework achieves a connectivity success ratio exceeding 90% across all slices, representing an average improvement of over 60% compared to non-graph-based reinforcement learning and supervised baselines. Notably, the framework demonstrates exceptional spectral efficiency, maintaining near-total connectivity while utilising less than 10% of the normalised system bandwidth, a 4× reduction in resource overhead compared to traditional methods. Furthermore, the GNN-driven architecture ensures stable convergence during training, yielding a 1.6× higher system reward score. Our findings validate GNN-DQN as a high-performance, scalable, and resource-efficient paradigm for intelligent orchestration in 5G and 6G networks.
A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.
Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang· IEEE Access· 0 citations
Space–Air–Ground Integrated Networks (SAGINs) have been envisioned to support next-generation communication networks.Due to their heterogeneity and dynamic link characteristics, routing in SAGIN is challenging. Existing shortest-path routing mechanisms do not adapt well to the varying bandwidth and latency, leading to poor quality of service (QoS). In this paper, we propose a deep reinforcement learning (DRL)-based adaptive routing scheme for maximizing throughput and minimizing end-to-end delay jointly in SAGIN. In the proposed model, an agent learns the policy of choosing the suitable path by interacting with the network environment and obtaining rewards. The network is modeled as a weighted graph with delay and bandwidth constraints. We compare our model with a traditional delay minimization baseline over multiple independent runs. Experimental results show that our DRL approach achieves a 6.51% improvement in average throughput and a 29.90% reduction in end-to-end delay compared to the baseline strategy. Statistical analysis confirms the robustness of the delay reduction, highlighting the effectiveness of reinforcement learning in dynamic HetNets. This indicates that adaptive policy learning enables better congestion avoidance and more efficient resource utilization. Overall, the proposed DRL-based routing framework offers a scalable and intelligent solution for optimizing performance in complex SAGIN architectures, with promising potential for next generation integrated communication systems.
The emergence of 5G and the imminent deployment of 6G networks have revolutionized wireless communication by enabling ultra‐reliable low‐latency communication, enhanced mobile broadband, and massive machine‐type communication. While these diverse applications foster increasingly complex congestion scenarios due to highly dynamic traffic patterns, diverse requirements of different services, and ultra‐dense deployments of the network, efficient congestion control has become a pressing challenge. Conventional methods cannot address the nonlinear and multi‐relational characteristics of data traffic. Therefore, with the rapid scaling of connected devices with heterogeneous service demands, bottlenecks will be experienced, which leads to packet loss, increased latency, and poor QoS. Given the above challenges, this paper introduces an advanced congestion control framework developed on top of intelligent graph‐based learning and optimization strategies. The input data is from the DeepSense 6G dataset and preprocessed using Z‐score and min–max normalization to enhance the stability and uniformity of the preprocessed data. The feature selection is further fine‐tuned through the Orchard Algorithm (OA), ensuring highly relevant feature selection. An innovative combination of Multi‐Relational Skeleton Graph Attention Networks is proposed to cater to the challenge of understanding multiple relational behaviors in data traffic. Furthermore, the performance of the integrated model is optimized with the Greater Cane Rat Optimization Algorithm to enhance learning efficiency and the precision of congestion control. Experimental analysis demonstrates an outstanding accuracy of 99.9%, ensuring the proposed approach effectively mitigates congestion in 5G/6G networks.
R. Mohanapriya, M. Gunavathie, S. Susmi et al.· International Journal of Com...· 0 citations
Cell-free massive multiple-input multiple-output (mMIMO), which eliminates cell edge effects and enhances coverage and resource utilization, is suited for industrial Internet of things (IIoT) applications. In user-centric cell-free mMIMO-based IIoT networks, joint optimization of network slicing and access point (AP) selection is crucial for meeting diverse quality-of-service (QoS) requirements. However, the joint optimization is challenging due to the coupling of resource allocation decisions and typically imperfect channel state information. In this paper, we formulate the joint AP selection and network slicing problem as a constrained Markov decision process (CMDP) with a hybrid action space, and propose a deep reinforcement learning (RL) algorithm, domain-guided hybrid soft actor-critic for CMDP (DG-HSA2C), to maximize the long-term proportional fairness in UE transmission rates while ensuring their QoS across slices. DG-HSA2C integrates CMDP-based RL into a hybrid action space by extending the Lagrangian multiplier method. To mitigate reward hacking, our algorithm incrementally predicts future states and incorporates a domain-adaptation mechanism, enhancing fairness in resource allocation and balancing performance across slices. Simulations verify our algorithm’s effectiveness in achieving rate fairness among UEs and mitigating reward hacking under the balance of QoS and rewards.
Na Li, Meiyan Song, Hangguan Shan et al.· IEEE Transactions on Communi...· 0 citations
Routing optimization in cloud-edge collaborative networks faces a fundamental conflict between global strategic planning and local real-time responsiveness, further complicated by structural heterogeneity and stochastic traffic patterns. Traditional protocols lack adaptivity, while existing Deep Reinforcement Learning (DRL) approaches based on Graph Neural Networks (GNN) struggle with limited receptive fields and over-smoothing issues in large-scale topologies. In this paper, we propose HAT-Route, a Transformer-driven hierarchical routing framework supported by the Network Digital Twin (NDT). Our contributions are threefold: 1) We establish a cloud-edge collaborative architecture operating under the Centralized Training and Decentralized Execution paradigm. This architecture balances the trade-off between global optimization and real-time inference. 2) We introduce FlowFormer, a Spatiotemporal Transformer for the NDT. FlowFormer integrates a novel Edge-Conditioned Spatial Attention (EC-SAT) mechanism to capture physical link constraints and distinguish between congestion and Head-of-Line (HOL) blocking. 3) We design HAT-Route, a hierarchical DRL agent that utilizes Graph Transformers for global policy learning in the cloud, coupled with knowledge distillation to deploy lightweight policies at the network edge. Extensive experiments demonstrate that our framework outperforms traditional protocols and GNN-based baselines in terms of QoS optimization, training stability, scalability, and generalization capability on large-scale network topologies.
Bin Dai, Yuntao Wang, Jianhai Zheng· IEEE Transactions on Network...· 0 citations
Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al.· Technologies· 0 citations