Skip to content
Preprint

DLB: Distributed Load Balancing at Scale for Generative AI Inference

Sep 2026 · 0 citations · 44 references
Computer Science Engineering

TL;DR

DLB, the Distributed Load Balancer is introduced, a novel system designed to minimize end-to-end user latency for large-scale, heterogeneous workloads and its design choices and practical experiences gained from the system in production are detailed.

Abstract

The reliance on scarce and expensive accelerators such as GPUs and TPUs in modern datacenters places unprecedented demands on backend infrastructure. For workloads characterized by heterogeneous service times and complex multi-stage processing, such as Generative AI, conventional load balancing techniques are often inadequate, relying heavily on costly overprovisioning to maintain service level objectives. This paper introduces DLB, the Distributed Load Balancer, a novel system designed to minimize end-to-end user latency for large-scale, heterogeneous workloads. DLB employs a scalable, distributed design with peer-to-peer probing to maintain real-time visibility into server capacity across large-scale, geographically distributed infrastructure. The system continuously learns latency models to estimate the latency impact of routing decisions, allowing it to effectively manage heterogeneous hardware and diverse model architectures. We provide a novel theoretical analysis of our routing algorithms that establishes their stability and global performance guarantees over time. We also evaluate DLB through extensive simulations, which show substantial gains compared to state-of-the-art load balancing algorithms. Finally, following a 22-month deployment of DLB at Google, where it facilitates large-scale Generative AI inference for thousands of different machine learning models and millions of requests per second, we detail the design choices and practical experiences gained from the system in production. Analysis of production migrations demonstrates that DLB yields statistically significant latency reductions compared to the legacy baseline, including a 17\% decrease in median latency and a 13\% decrease at the p95 tail.

View source

Similar papers

Open access 2026

Synchronous Distributed Training With Runtime-Adaptive Mechanisms in Hybrid Cloud Environments

Experimental evaluations demonstrate the effectiveness of the framework ASTRA, which achieves lower time-to-accuracy than a resource-heterogeneity-aware baseline and several compression-based frameworks, while preserving convergence quality and robustness across heterogeneous hybrid cloud environments.

Tuan Anh Vuong, Thanh Loi Hoang, Huan Le et al. · 0 citations
Book Open access Aug 2026

OrionInfer: Low-Overhead Parallelism Switching and Live Migration for Efficient LLM Serving

Or OrionInfer, an adaptive LLM serving system that aligns inference strategies with real-time demand and introduces three key techniques: runtime switching between data parallelism and tensor parallelism with negligible overhead, an efficient inference pipeline that preserves batching efficiency during parallelism tran...

Jingqi Feng, Guang Yang, Yukai Huang et al. · 0 citations
Preprint Sep 2026

VarioPath: Workload-Aware All-to-All Communication for PCIe GPU Clusters

VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems that combines an offline topology-aware analyzer with an online demand-aware scheduler, and shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP.

Yao Fei, Jin Fang, Si-Ze Zheng et al. · 0 citations
Book Open access Sep 2026

Taming Inference Workloads at Global Scale: Foundation Model Serving in Amazon Bedrock

Foundation model (FM) inference platforms must manage scarce accelerator capacity distributed unevenly across regions. They must also handle requests whose token consumption varies widely and may be revealed progressively during generation, while sharing capacity across workloads with different latency and throughput o...

Pratik Pankaj Raichura, Somu Perianayagam, Rama Krishna Sandeep Pokkunuri et al. · 0 citations
Preprint Aug 2026

ShardMeter: Sharded and Geo-Distributed Training Without the Guesswork

ShardMeter is introduced, a lightweight analytical performance model that predicts the end-to-end runtime of transformer-based workloads across arbitrary sharded, distributed, and even decentralized training.

Tim Beringer, Patrick Diem, Felix Wolf et al. · 0 citations
Preprint Aug 2026

EasyBalance: Cross-Layer Load Balancing in Distributed MoE Inference

EasyBalance is proposed, a cross-layer load balancing strategy that requires no modifications to the expert-device mapping, enabling instant adaptability and incurring essentially no additional overhead.

Yi-Ze Wu, Ke Gao, Ling Li et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.