Preprint
Sep 2026
DLB: Distributed Load Balancing at Scale for Generative AI Inference
DLB, the Distributed Load Balancer is introduced, a novel system designed to minimize end-to-end user latency for large-scale, heterogeneous workloads and its design choices and practical experiences gained from the system in production are detailed.
S. Balseiro, B. Wydrowski, Sameer Agarwal et al.
· 0 citations