Skip to content
Book Open access

Unlocking Software-defined GPU Fabric Scheduling in the LLM Era

Sep 2026 · Proceedings of the 17th ACM SIGOPS Asia-Pacific Workshop on Systems · 1 citation · 47 references

Abstract

Large language model (LLM) systems increasingly rely on techniques such as prefill-decode disaggregation, KV-cache offloading, and computation-communication overlap. These optimizations often treat GPU interconnects as best-effort substrates, overlooking contention across shared PCIe, NVLink, and RDMA fabrics. We characterize GPU-fabric contention in LLM workloads and find that hardware arbitration can impose implicit priorities across traffic classes, causing communication to interfere with computation and other data movements. In particular, GPU-to-GPU NVLink communication can interfere with GPU memory access, while GPU-direct RDMA can contend with host-device transfers over PCIe. Based on these observations, we advocate a vision for software-defined GPU fabric scheduling and present GPUWeaver, a context-aware communication scheduler that monitors application progress and fabric state, identifies contention, and dynamically regulates communication injection according to application performance goals. Our study highlights the opportunity to make GPU fabrics actively managed resources for efficient and predictable LLM systems.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.