Aug 2026· Conference on Applications, Technologies, Architectures, and Protocols for Computer Communication· 0 citations· 55 references
Computer Science
TL;DR
This work introduces CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header, and proposes Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production.
Abstract
Optimizing burst-heavy datacenter workloads necessitates finegrained network control and visibility. We introduce CSIG, a protocol that delivers precise, multi-bit bottleneck congestion signals via a fixed-length Ethernet header. The architecture captures μsgranularity switch metrics, such as available bandwidth, and signals them to end-hosts using in-band, line-rate operations. We propose Fast Ramp-Up, a congestion control primitive that leverages these bottleneck signals to reduce median RPC latency by 20% and unclaimed bandwidth by 60% in production. Beyond transport-level performance, CSIG enables flow-aware observability by embedding μs-scale metrics into every packet, allowing individual application transfers to pinpoint their bottleneck location, such as the topology tier limiting their performance. CSIG thus transforms network telemetry from post-hoc correlation into a real time, context-aware capability. We demonstrate CSIG's broad deployability by validating it across five generations of commodity switch hardware (up to 102.4 Tbps), four NIC generations, and five transport stacks. Our design proves that a streamlined Layer 2 approach, focusing exclusively on the principal path bottleneck, provides transport-agnostic gains without requiring forklift hardware upgrades.
InfiniFlow is presented, a credit-based hop-by-hop flow control method that supports massive VCs with a limited buffer budget via per-port buffer sharing, and introduces a paradigm shift in buffer management: Upstream Allocates Buffer for Downstream (UABD).
Zerui Tian, Sen Liu, Minkun Xue et al.· Conference on Applications,...· 0 citations
Evaluation on representative workloads demonstrates that P2CS achieves performance comparable to in-network mechanisms while significantly reducing complexity and cost, and requires minimal software changes making it readily deployable in today's datacenter infrastructure.
Ali Munir, Xiaolin Pang, Junyi Zhang· Conference on Applications,...· 0 citations
Datacenter performance is often limited by network-centric congestion controls relying on low-level metrics (e.g., packet loss, latency) that misinterpret applications needs. This work argues that applications should participate in congestion control decisions and introduces the starvation ratio (SR), a metric that detects when applications are truly limited by the network. Experimental evaluations within an 11-flow bottleneck scenario on a Linux-based prototype show that asynchronous applications can absorb network variations within a newly identified “silence zone” without degradation, proving conventional controls are overly restrictive. By deploying proactive and reactive mechanisms, our approach consistently reduces Flow Completion Time (FCT) for network-sensitive workloads. Notably, the proactive configuration eliminates micro-recovery delays, keeping the starvation ratio close to zero and reinforcing the baseline protocol through stable congestion window regulation. We conclude that shifting to application-driven signaling aligns network transmission with the receiver’s processing pace, preventing computational underutilization.
Anderson Henrique da Silva Marcondes, Enzo B. Boscatto, Guilherme Piêgas Koslovski· Anais do LIII Seminário Inte...· 0 citations
The proposed Multi-Path Multi-Level Feedback Queueing (MP-MLFQ) leverages the spatial diversity and regularity of DCNs to realize a scheduler with numerous logical priority levels while occupying as low as 2 physical priority queues within network switches.
Alessandro Cornacchia, Andrea Bianco, Paolo Giaccone et al.· 0 citations
The performance of datacenter congestion control algorithms (CCAs) is highly sensitive to bursty traffic patterns, yet a significant fidelity gap exists between evaluation workloads and production traffic. Current evaluations primarily rely on synthetic workloads constructed from flow-size CDFs with incast overlaid on top, an approach that, while intuitive, we show produces traffic that is dissimilar to production in its temporal burst clustering. As a result, these workloads fail to exercise the full range of conditions that CCAs encounter in production, and hence, protocols that demonstrate gains in simulation risk diminished performance or unexpected failure modes upon deployment. This motivates the need for a deeper understanding of burstiness for CCA evaluation. To this end, we decompose burstiness into four key dimensions, and use DCTCP as a case study to show distinct behavioral regimes in each dimension. Building on this, we envision a burst-centric evaluation stack: behavioral regime analysis across various CCA classes, and a burst generator to ensure regime coverage along these dimensions, enabling thorough and robust evaluations.
Pragna Mamidipaka, Srikanth Sundaresan, Theophilus A. Benson· Asia-Pacific Workshop on Net...· 0 citations
The experience in designing, deploying, and operating Pegasus, a data center network tailored for the AI cloud, along with operational lessons learned from its deployment are shared.
Xianneng Zou, Yadong Liu, Yiran Zhang et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.