Skip to content
Book Open access

DistDPU: A Disaggregated DPU Architecture for High-Performance and Cost-Efficient AI Clouds

Aug 2026 · Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication · pp. 519-534 · 0 citations · 41 references
Computer Science

TL;DR

DistDPU is presented, a disaggregated DPU architecture that redefines the scaling abstraction for high-bandwidth cloud networking and co-designs the EM-OM functions and the inter-module fabric to minimize virtualization overhead while enforcing security and manageability invariants equivalent to those of a monolithic DPU.

Abstract

AI training and inference are driving cloud networks toward terabit-per-second (Tbps) bandwidth per server, challenging the scalability and efficiency of today's cloud network architectures. A prevalent design scales bandwidth by stacking monolithic Data Processing Units (DPUs), but this approach tightly couples control and data plane resources, leading to excessive cost, power consumption, and operational complexity. We identify a fundamental control-data plane divergence in AI clouds: while data plane bandwidth demand grows rapidly, control plane demand remains largely flat due to the dominance of elephant flows. As a result, monolithic DPUs become systematically over-provisioned when used as bandwidth scaling primitives. We present DistDPU, a disaggregated DPU architecture that redefines the scaling abstraction for high-bandwidth cloud networking. DistDPU decomposes a monolithic DPU into lightweight, bandwidth-provisioning Execution Modules (EMs) and a shared, control-centric Orchestration Module (OM), enabling independent scaling of data and control plane resources. By scaling out low-cost EMs under a single OM, DistDPU exposes a unified, high-bandwidth logical DPU interface to the cloud management plane. To preserve RDMA performance and multi-tenant isolation at scale, we co-design the EM-OM functions and the inter-module fabric to minimize virtualization overhead while enforcing security and manageability invariants equivalent to those of a monolithic DPU. DistDPU has been deployed in production for two years. It serves more than 10,000 GPUs and delivers higher efficiency and strong performance on real-world AI workloads than state-of-the-art designs.

Read PDF

Similar papers

Book Open access Aug 2026

Pegasus: A Data Center Network for Bare-Metal AI Cloud

The experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment are shared.

Xianneng Zou, Yadong Liu, Yiran Zhang et al. · 0 citations
Book Open access Aug 2026

Spillway: Orchestrating DPU and Host into a Unified vSwitching Fabric

Spillway introduces a DPU-host hybrid data plane that repurposes idle host CPU resources to process spillover traffic when the DPU becomes the bottleneck, and decouples virtual switching capacity from static DPU hardware limits.

Xiaochong Jiang, Dian Fan, Yilong Lv et al. · 0 citations
Book Open access Aug 2026

Dynamic Compute and Network Orchestration for Disaggregated RL

This work builds Silverstone to orchestrate dynamically both compute and network in disaggregated RL, using a reconfigurable optical-electrical fabric called RFabric that achieves superior performance-cost efficiency at scale over static Fat-Tree networks.

Xin Tan, Yicheng Feng, Yu Zhou et al. · 0 citations
Book Open access Aug 2026

CSIG: Congestion Signaling for Datacenter Transports

This work introduces CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header, and proposes Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production.

Abhiram Ravi, Nandita Dukkipati, Weiwu Pang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.