Tail‐Latency‐Aware Reinforcement Learning for Dynamic Resource Allocation in 5G Network Slicing
Abstract
Network slicing is a fundamental technology in 5G and beyond networks, enabling multiple virtual networks to coexist over a shared physical infrastructure while supporting heterogeneous services with diverse quality‐of‐service (QoS) requirements. In this context, radio access network (RAN) slicing plays a critical role in allocating radio resource units to different service slices. A fundamental challenge in RAN slicing is the dynamic allocation of radio resources under fluctuating traffic demand and user mobility, where rare but severe tail‐latency events corresponding to extreme delays experienced by a small fraction of packets have a dominant impact on performance, particularly for latency‐critical ultrareliable low‐latency communication (URLLC) services. Although reinforcement learning has been extensively studied for RAN slicing, most existing RAN‐side approaches focus on optimizing average performance metrics and permit unconstrained exploration, which can result in instability and service‐level agreement (SLA) violations in safety‐critical scenarios. This paper proposes a tail‐latency‐aware proximal policy optimization (PPO) framework for dynamic multislice RAN resource allocation. The proposed approach integrates offline behavior cloning using ns‐3‐generated new radio (NR) traces with online policy refinement to enable continuous adaptation. The framework incorporates a slice‐aware state representation capturing traffic and mobility dynamics, an explicit tail‐latency penalty targeting the 95th‐percentile URLLC delay, a URLLC‐aware adaptive exploration mechanism, and an SLA‐driven safety layer with emergency resource reallocation. System‐level evaluations across three representative scenarios demonstrate that the proposed approach consistently outperforms baseline methods in overall SLA satisfaction and tail‐latency control, achieving above 92% overall SLA compliance and maintaining balanced compliance across all service slice types, particularly during sudden traffic spikes.