Jul 2026· International Conference on Ubiquitous and Future Networks· pp. 1188-1193· 0 citations· 13 references
Abstract
This paper proposes a Service Level Agreement (SLA)-aware resource allocation framework for 6G V2X slicing, realized as a Soft Actor-Critic (SAC) based xApp within the Open-Radio Access Network (O-RAN) near-realtime-RAN Intelligent Controller (near-RT-RIC). The xApp dynamically distributes radio resources across heterogeneous slices, minimizing SLA violations while considering fairness and throughput efficiency. Unlike heuristic or single-metric Deep Reinforcement Learning (DRL) methods, our design incorporates deadline awareness and service reliability directly into the reward formulation. Simulation results show that the proposed scheme consistently outperforms fixed, random, proportional, and Exponential moving Average (EMA)-based baselines, improving average packet delivery ratio (PDR), reducing mean SLA violations, and achieving a Pareto-optimal trade-off between throughput and compliance. These findings demonstrate the potential of O-RAN-native intelligent control for future 6G networks.
Network slicing is a fundamental technology in 5G and beyond networks, enabling multiple virtual networks to coexist over a shared physical infrastructure while supporting heterogeneous services with diverse quality‐of‐service (QoS) requirements. In this context, radio access network (RAN) slicing plays a critical role in allocating radio resource units to different service slices. A fundamental challenge in RAN slicing is the dynamic allocation of radio resources under fluctuating traffic demand and user mobility, where rare but severe tail‐latency events corresponding to extreme delays experienced by a small fraction of packets have a dominant impact on performance, particularly for latency‐critical ultrareliable low‐latency communication (URLLC) services. Although reinforcement learning has been extensively studied for RAN slicing, most existing RAN‐side approaches focus on optimizing average performance metrics and permit unconstrained exploration, which can result in instability and service‐level agreement (SLA) violations in safety‐critical scenarios. This paper proposes a tail‐latency‐aware proximal policy optimization (PPO) framework for dynamic multislice RAN resource allocation. The proposed approach integrates offline behavior cloning using ns‐3‐generated new radio (NR) traces with online policy refinement to enable continuous adaptation. The framework incorporates a slice‐aware state representation capturing traffic and mobility dynamics, an explicit tail‐latency penalty targeting the 95th‐percentile URLLC delay, a URLLC‐aware adaptive exploration mechanism, and an SLA‐driven safety layer with emergency resource reallocation. System‐level evaluations across three representative scenarios demonstrate that the proposed approach consistently outperforms baseline methods in overall SLA satisfaction and tail‐latency control, achieving above 92% overall SLA compliance and maintaining balanced compliance across all service slice types, particularly during sudden traffic spikes.
Unknown authors· International Journal of Com...· 0 citations
Future sixth-generation (6G)-oriented networks require programmable control that can adapt routing to latency and congestion without unsafe online exploration. This study evaluates offline multi-agent deep deterministic policy gradient (MADDPG) with behavior-adjusted training rewards for latency-aware path control in software-defined networking (SDN). Each traffic pair is modeled as an agent selecting one of three retained candidate paths, while centralized critics learn coordinated decisions from topology-specific Ryu–Mininet transition datasets. Nine policies are compared using ten paired seeds on fat-tree, mesh-grid, and WAN-corridors topologies under a deployed utilization–latency weighting of 0.60/0.40, together with flow-completion, latency, congestion, architectural-comparison, sensitivity, robustness, statistical, and controller-overhead analyses. The utilization-aware path heuristic achieves the strongest overall reward ranking. MADDPG is the strongest learned policy on fat-tree, is not significantly outperformed by any evaluated policy on mesh-grid, and remains statistically tied with completion-matched policies on WAN-corridors. Behavior adjustment is topology-dependent rather than uniformly beneficial. The exported policy requires approximately 52μs per joint decision, whereas complete control-loop timing is dominated by network-statistics polling. These results support offline multi-agent SDN control as a competitive, low-overhead option when interpreted jointly with topology structure, flow completion, and strong heuristic baselines.
A. Kyzyrkanov, Y. Nurakhov, Zhenis Otarbay et al.· Technologies· 0 citations
The transition toward Open Radio Access Network (O-RAN) architecture has enabled unprecedented intelligence and flexibility in 5G and 6G network slicing. However, a fundamental challenge remains in managing the tension between radio unit energy efficiency and the strict Service Level Agreement (SLA) requirements of Ultra-Reliable Low-Latency Communication (URLLC) slices, particularly under highly dynamic traffic conditions. Existing O-RAN approaches suffer from a timescale conflict where Non-Real-Time (Non-RT) policy planners optimize for long-term energy but fail to react to rapid traffic surges, while Near-Real-Time (Near-RT) controllers prioritize reliability at the cost of significant energy over-provisioning. To address this, we propose H-RLS, a hierarchical multi-timescale framework that decouples control into a Non-RT Proximal Policy Optimization (PPO) agent for strategic, energy-aware policy planning and a Near-RT Recursive Least Squares (RLS)-assisted xApp. By predicting millisecond-level delay risks, the xApp acts as a mathematically constrained safety net, applying bounded tactical adjustments when critical SLA violations are detected. Extensive evaluations across dynamic traffic transitions demonstrate that H-RLS maintains zero SLA violations. By actively preventing resource over-provisioning, the framework achieves the lowest composite Energy-SLA cost across all tested regimes, significantly minimizing dynamic power consumption while preserving Enhanced Mobile Broadband (eMBB) service integrity.
A QoE-aware framework for Multi-Access Edge Computing-enabled Open Radio Access Network (O-RAN) architectures, combining a graph attention network (GAT) encoder, distributed multi-agent DRL, and privacy-preserving FL, while transitioning control from Quality of Service (QoS) to QoE metrics is proposed.
Manoj Prasad Kunasegran, Wai Leong Pang, S. K. Phang· IEEE Access· 0 citations
: Space-Air-Ground Integrated Networks (SAGIN) provide a multi-layered, wide-coverage computing infrastructure for distributed urban sensing systems. However, their heterogeneity and dynamics pose unprecedented challenges for task offloading and resource allocation. Existing methods struggle to simultaneously address the complexity of cross-layer decision-making and reliability assurance under uncertain conditions. This paper proposes a novel framework, termed DRL-RA, which synergistically integrates Deep Reinforcement Learning (DRL) with reliability-aware optimization. The framework consists of two complementary components: (1) a Dueling Double Deep Q-Network (D3QN) module that learns adaptive policies to make offloading decisions among various options including local execution, terrestrial edge, UAVs, and satellites; (2) a Reliability-Aware Multi-Objective Optimization Framework (RA-MOOF) that introduces explicit reliability guarantees through cross-layer link reliability modeling, node availability estimation, and smooth reliability proxy functions. Addressing the heterogeneous communication characteristics of the SAGIN architecture, this paper establishes a complete cross-layer delay model and composite reliability metrics. The reliability formulation is defined under explicitly stated conditional-independence assumptions, and the proposed smooth constraint terms are treated as surrogate CMDP costs rather than exact hard chance-constraint guarantees. Extensive experiments in a SAGIN simulation environment demonstrate that the proposed method improves the task completion rate by 3.8%, reduces average latency by 11.1%, and increases system reliability by 3.9% compared to state-of-the-art benchmarks. The optimization-only RA-Opt baseline is used as a non-real-time optimization reference for assessing reliability-aware offloading decision quality, while deployment-time decision-latency comparisons are interpreted primarily among learned inference policies. Comprehensive ablation studies and statistical validation across multiple random seeds confirm the contributions of each component, while cross-layer offloading decision analysis verifies the effectiveness of the method across different network layer selections.
Fei-Yan Bu, Zheng Wang, Yong Pan et al.· Computers, Materials & C...· 0 citations
This work proposes an enhanced Proximal Policy Optimization (PPO) framework for resource-aware and latency-sensitive SFC placement in edge-enabled networks, and demonstrates the applicability of the proposed framework in mission-critical and latency-sensitive service environments.
Nithin Melala Eshwarappa, Ching-Hsien Hsu, Hojjat Baghban et al.· ACM Transactions on Modeling...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.