Jul 2026· International Conference on Ubiquitous and Future Networks· pp. 589-592· 0 citations· 10 references
Abstract
The AI-RAN Alliance has proposed the coexistence of Artificial Intelligence (AI) workloads and Radio Access Network (RAN) functions on shared edge infrastructure, referred to as AI-and-RAN. However, practical understanding of resource contention between compute-intensive AI tasks and timing-constrained RAN operations remains limited, largely due to the lack of end-to-end platforms that enable controlled co-location and measurement across system and RAN layers. To fill this gap, this paper presents an end-to-end AI-and-RAN testbed for controlled co-location experiments on an unified CPU–GPU platform. We deploy a containerized 5G RAN stack with an Open Radio Access Network (O-RAN) compliant E2 interface for monitoring alongside edge Large Language Model (LLM) inference workloads. To expose resource sensitivity, we intentionally constrain the CPU budget and vary AI workload intensity. Through measurements of Medium Access Control (MAC) layer key performance indicators (KPIs) and edge-level resource utilization, we demonstrate that LLM inference reduces downlink (DL) throughput and increases Block Error Rate (BLER), providing empirical insights to inform future AI-and-RAN orchestration strategies.
Cellular networks are integrating Artificial Intelli- gence (AI) into radio access network control. The MAC scheduler is a promising target because it allocates a limited resource, spectrum, at every slot, under competing latency, throughput, and reliability requirements. However, most learning-based sched- ulers are evaluated only in simulation. Production schedulers are difficult to modify, and realistic stress tests require more radio hardware than most laboratories can provide. We present MAC-Gyver, an open-source framework for developing and evaluating scheduling applications that execute directly inside the OpenAirInterface scheduler. It exposes scheduler observations and controls through typed interfaces while preserving the underlying protocol and real-time execution paths. The same applications run over the air and in mac-emu, a PHY-less emulator that executes the unmodified OpenAirInterface Layer 2 stack for up to 90 users on one host at real-time slot pace, with a 3GPP-compliant channel model. To showcase the flexibility of MAC-Gyver, we evaluate two use cases. A proactive uplink scheduler predicts packet arrivals and roughly halves median round-trip latency. A frequency-selective uplink scheduler selects contiguous sub-bands from per-PRB sounding observations and is evaluated across mobility and power-limited operating points against an offline scheduling ceiling. Together, they show how the same production stack can be an AI playground that supports implementation, controlled evaluation, and over-the-air validation through complementary scheduling use cases.
M. Elkael, Reshma Prasad, Tamerlan Aghayev et al.· arXiv.org· 0 citations
O-DAG is presented, an end-to-end framework that closes the SAGA--simulation gap and evaluates five scheduling algorithms for a slice scheduling application across various configurations spanning 5K--50K UEs, 2--20 cells, and 2--10 network slices.
Y. Hwang, B. Krishnamachari· International Conference on...· 0 citations
Edge AI is evolving from isolated inference toward long-running services that coordinate model pipelines, data streams, state, and accelerators near users and physical environments. Cloud-native and edge-native platforms offer useful foundations, but their primary control objects–containers, nodes, links, and enrolled sites–generally do not expose the model, data, state, quality, and participation semantics required by these services. This paper presents eAI+, a vision for an edgeAI-native service platform built around three first-class control objects: AI service graphs, dynamic edge resource fabrics, and participant contracts. eAI+ aims to preserve service quality under latency, privacy, reliability, cost, and participation constraints through three coordinated mechanisms. Runtime would select safe execution adaptations based on current workload, environment, and contract signals. Deployment would map service-graph components and prepared fallbacks to heterogeneous resources. PolyLink is the participant-contract module for plug-and-play resource onboarding; it would register contributors, verified resource offers, capabilities, and participation terms. Once a resource is onboarded, it would become available to Deployment for placing eligible service-graph components under the registered contract, while PolyLink would maintain metering, reputation, rewards, and exit events. Migration would transfer only continuity-critical state or control when mobility, overload, policy changes, or contributor lifecycle events invalidate the current placement. This framing treats edge AI as a coordinated service-platform problem across models, data, state, resources, and contracts.
Jiannong Cao, Zhiyuan Hu, Mingjin Zhang et al.· International Symposium on S...· 0 citations
An integrated management framework powered by Tapis is shown how a unified UI-driven workflow streamlines the transition from initial evaluation to deployment, ensuring operational consistency and reproducibility without manual script porting.
Manikya Swathi Vallabhajosyula, Gautam Gururaj Molakalmuru, Samuel Khuvis et al.· Practice and Experience in A...· 0 citations
The sixth-generation (6G) of mobile networks will be shaped not only by artificial intelligence (AI)-enabled network automation and optimization, but also by the need to serve AI as a 6G-native workload. Emerging AI services introduce traffic and compute demands that differ from conventional mobile broadband. Their user experience depends on how quickly useful information is delivered, how bursty and asymmetric multimodal flows are handled, and where inference, retrieval, caching, and content processing are executed. This article presents a joint connectivity-compute view of AI-native 6G. We first characterize representative AI service traffic in terms of uplink/downlink throughput skew, burstiness, and token latency. Next, we discuss how fifth-generation extended reality awareness mechanisms can evolve toward AI traffic characteristics awareness in 6G. Finally, we introduce AI Grid as a distributed AI infrastructure platform for placing workloads according to latency, cost, policy, and service-level constraints. Together, AI-aware connectivity and AI Grid enable 6G as a distributed intelligence platform.
Lopamudra Kundu, Xingqin Lin, Shuvo Chowdhury et al.· 0 citations
An architecture that decouples complex policy enforcement from high-speed packet forwarding to support VPC semantics on back-end NICs and enable front-end/back-end integration is proposed, suggesting that commodity hardware can support both high-throughput AI training and flexible VPC features.
Yinhe Wang, Xing Li, Enge Song et al.· Asia-Pacific Workshop on Net...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.