Edge AI inference is an important workload in 5G networks. Multiple classes of edge AI workloads often share edge infrastructure, yet each may require distinct latency and rate guarantees. 5G network slicing supports differentiated requirements on the network side, but shared GPU inference also needs compute-side guarantees after requests reach the Multi-access Edge Computing (MEC) host. We extend 3GPP network slicing with compute-side enforcement so that slice guarantees remain effective after traffic reaches the MEC host. To realize this extension, we design a GPU scheduler that combines Hierarchical Token Bucket (HTB)-based traffic conditioning with Earliest Deadline First (EDF) scheduling. Our scheduler enforces per-class assured goodput, defined as the committed rate of latency-compliant completions for each class. The GPU scheduler identifies request classes via tags, which are assigned during GTPU encapsulation at the 5G user plane. This integration preserves overall latency guarantees across both the network and compute domains of a 5G slice for AI.
Yu-Hong Shen, Wen-Ju Chiang, Hsiang-Ming Hung et al.· Proceedings of the ACM SIGCO...· 0 citations
DPDK-DU-NS is presented, a DPDK-enabled O-RAN DU that supports per-flow network slicing and bandwidth management in compliance with 3GPP 5G QoS flow and bearer management specifications and is validated as a high-performance and scalable foundation for next-generation O-RAN DU implementations.