Unlocking Software-defined GPU Fabric Scheduling in the LLM Era
Abstract
Large language model (LLM) systems increasingly rely on techniques such as prefill-decode disaggregation, KV-cache offloading, and computation-communication overlap. These optimizations often treat GPU interconnects as best-effort substrates, overlooking contention across shared PCIe, NVLink, and RDMA fabrics. We characterize GPU-fabric contention in LLM workloads and find that hardware arbitration can impose implicit priorities across traffic classes, causing communication to interfere with computation and other data movements. In particular, GPU-to-GPU NVLink communication can interfere with GPU memory access, while GPU-direct RDMA can contend with host-device transfers over PCIe. Based on these observations, we advocate a vision for software-defined GPU fabric scheduling and present GPUWeaver, a context-aware communication scheduler that monitors application progress and fabric state, identifies contention, and dynamically regulates communication injection according to application performance goals. Our study highlights the opportunity to make GPU fabrics actively managed resources for efficient and predictable LLM systems.