Skip to content
Preprint

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

Sep 2026 · 0 citations · 43 references
Computer Science

TL;DR

VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems that combines an offline topology-aware analyzer with an online demand-aware scheduler, and shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP.

Abstract

AlltoAllv communication is a critical primitive in distributed large-model inference, particularly for mixture-of-experts (MoE) models. The growing adoption of PCIe GPU systems for cost-efficient inference makes AlltoAllv performance on these systems increasingly important. Without a dedicated scale-up interconnect (e.g., NVLink or Infinity Fabric), PCIe GPU systems carry both intra-node and inter-node traffic through the PCIe hierarchy, where concurrent transfers can contend for PCIe link bandwidth. This link contention, compounded by skewed traffic distributions and dynamic traffic demand, makes efficient AlltoAllv scheduling challenging. Existing approaches are either poorly suited to PCIe GPU systems or incur substantial schedule synthesis overhead that reduces their practicality in real-world deployments. We present VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems. It combines an offline topology-aware analyzer with an online demand-aware scheduler. The analyzer records contention-free transfer patterns as AlltoAllv channels and exploits topology symmetry to build a compact catalog for efficient search. The online scheduler decomposes each AlltoAllv invocation's demand across a sequence of channels, adapting to rapidly changing and skewed traffic while incurring low planning overhead. Evaluation on four platforms (up to 256 GPUs) shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP. End-to-end experiments show that VarioPath reduces Qwen3 inference latency by up to 27.2% and Wan2.1 generation latency by 6.1%.

View source

Similar papers

#machine learning Preprint Sep 2026

Efficient Expert-Parallel Communication on PCIe-Connected Consumer GPUs

Expert parallelism (EP) enables inference of large Mixture-of-Experts (MoE) models by placing their experts across multiple GPUs, but requires substantial communication between GPUs at every MoE layer. As contemporary MoE models activate more experts per token, this communication accounts for a growing fraction of infe...

Jaehwan Lee, Sang-Min Lee, Chaewon Kim et al. · 0 citations
Preprint Sep 2026

Joint Effects of GPU Server Topology, Parallelism, and Congestion Control on MoE Inference: A Controlled Simulation Study

Mixture-of-experts (MoE) models expand capacity via sparse activation, but inference across GPUs introduces tensor-parallel (TP) collectives and expert-parallel (EP) dispatch and combine operations. Completion time depends not just on communication volume but on how logical groups map onto intra-server interconnects, G...

Kai-Kai Yuan, Rui Xi, Yu Liu · 0 citations
Preprint Aug 2026

ElastiCo: Elastic Configuration and Interference-Aware Orchestration for GPU Clusters

ElastiCo is presented, an elastic co-location framework that enables training and inference workloads to safely share GPUs through three integrated mechanisms that decomposes the resulting multi-resource allocation problem into per-job configuration selection subproblems via dynamic per-resource shadow prices.

Jing-Hao Wang, Yi-Hang Zhou, Xiaoyang Sun et al. · 0 citations
Book Open access Sep 2026

Unlocking Software-defined GPU Fabric Scheduling in the LLM Era

Large language model (LLM) systems increasingly rely on techniques such as prefill-decode disaggregation, KV-cache offloading, and computation-communication overlap. These optimizations often treat GPU interconnects as best-effort substrates, overlooking contention across shared PCIe, NVLink, and RDMA fabrics. We chara...

Dan-Yang Chen, Yu-Feng Gu, Yibo Huang et al. · 1 citation
#small language model Book Open access Sep 2026

MIGServe: Layout-Aware Multi-Instance GPU Management for Efficient LLM Serving

MIGServe treats the physical layout of MIG instances as a first-class scheduling dimension through three techniques: buddy-aware partition placement, which preserves large contiguous free blocks by allocating next to existing occupied buddies; proactive pair-matching migration, which consolidates fragmented half-full b...

Jian-Wen Chen, Yun-Kai Liang, Bin Gao et al. · 0 citations
Preprint Aug 2026

FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution

Frontier open-weight models are increasingly available, but serving them still largely assumes datacenter infrastructure. We present FreeToken, an edge-native MoE serving system that treats a personal machine not as a small GPU, but as a unified, elastic inference platform. FreeToken co-designs the full serving stack,...

Shuo Yang, Xiao-yun Fan, Melissa Z. Pan et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.