Skip to content

Priority flow control-sensitive: Reducing tail latency with Priority flow control-sensitive in lossless data center networks

Aug 2026 · Journal of High Speed Networks · 0 citations · 16 references

TL;DR

Initial evaluations demonstrate that PFC-S can reduce the average flow completion time and effectively prevent congestion spreading, and experimental results show that PFC-S provides better protection for victim flows compared to standard PFC, BFC, and HPCC methods.

Abstract

Lossless fabrics are widely used in many production data centers, but they can give rise to issues such as head-of-line blocking and congestion spreading during network congestion, which can significantly degrade the performance of data center applications. Additionally, the latency of end-to-end solutions can lead to the buildup of switch queues. To address these challenges, this paper proposes a method called Priority Flow Control-Sensitive (PFC-S). PFC-S monitors buffer occupancy and traffic intensity, performs proactive rerouting, and prevents the impact of PFC congestion diffusion on victim flows. This approach helps maintain low buffer usage levels, thereby enabling control over tail latency. Initial evaluations demonstrate that PFC-S can reduce the average flow completion time and effectively prevent congestion spreading. Moreover, experimental results show that PFC-S provides better protection for victim flows compared to standard PFC, BFC, and HPCC methods.

View source

Similar papers

2026

FAFC: Fast and Accurate Flow Control in Data Center Networks

In data centers, large-scale many-to-one traffic can rapidly exhaust switch buffers and trigger priority-based flow control (PFC) pause, resulting in increased flow completion time (FCT) for uncongested flows. To address this issue, we propose an innovative switch-side fast and accurate flow control (FAFC) scheme. By differentially allocating pause time for each port during congestion, FAFC can minimize the performance loss for uncongested flows. Furthermore, FAFC is also coupled with an effective queue length prediction algorithm to enable proactive and reliable estimation of the congestion level. Extensive system-level simulations demonstrate that FAFC can flexibly allocate pause times across congested ports, which are not only compatible with existing PFC but also do not require per-flow states. We implemented FAFC in P4 programmable switches, showing it as lightweight flow control method that is portable for implementation in hardware. Remarkably, our large-scale simulations illustrate that compared to traditional PFC, FAFC improves the average FCT slowdown and 95% FCT slowdown by 10.6% and 23.3%, respectively, under Hadoop workload when performing HPCC congestion control.

Chengdi Lu, Yuang Chen, Fangyu Zhang et al. · 0 citations
Book Open access Aug 2026

ProLet: Proactive Multi-path Load Balancing for Lossless RDMA

ProLet is a load balancing scheme that enables proactive probing and reroutes elephant flows at flowlet granularity in lossless RDMA networks and reduces average and tail flow completion time slowdowns by 69% and 79%, respectively, compared to state-of-the-art load balancing schemes.

Hong Wang, Jin-Hao Luo, J. Tan et al. · 0 citations
Open access Jul 2026

Starvation ratio: letting applications drive datacenter congestion control

Datacenter performance is often limited by network-centric congestion controls relying on low-level metrics (e.g., packet loss, latency) that misinterpret applications needs. This work argues that applications should participate in congestion control decisions and introduces the starvation ratio (SR), a metric that detects when applications are truly limited by the network. Experimental evaluations within an 11-flow bottleneck scenario on a Linux-based prototype show that asynchronous applications can absorb network variations within a newly identified “silence zone” without degradation, proving conventional controls are overly restrictive. By deploying proactive and reactive mechanisms, our approach consistently reduces Flow Completion Time (FCT) for network-sensitive workloads. Notably, the proactive configuration eliminates micro-recovery delays, keeping the starvation ratio close to zero and reinforcing the baseline protocol through stable congestion window regulation. We conclude that shifting to application-driven signaling aligns network transmission with the receiver’s processing pace, preventing computational underutilization.

Anderson Henrique da Silva Marcondes, Enzo B. Boscatto, Guilherme Piêgas Koslovski · 0 citations
Open access Aug 2026

cdcPIM: a proactive congestion control scheme for cross-datacenter RDMA networks

Driven by the requirements of machine learning, cloud storage, and other network-intensive applications, remote direct memory access (RDMA) has been widely adopted in high-speed networks and is gradually being applied to geographically distributed datacenters. However, in cross-datacenter scenarios, long control loop latency and mixed traffic prevent existing RDMA congestion control schemes from perceiving and reacting to congestion in a timely and fair manner; this can lead to severe performance degradation and unfairness. To address these issues, we propose cdcPIM, a proactive congestion control scheme extended from datacenter parallel iterative matching (dcPIM) for cross-datacenter networks, which restructures the end-to-end control loop by introducing switch-coordinated control points, effectively transforming long-haul, RTT-bound feedback into localized control. Specifically, cdcPIM deploys a local control point by moving the token generation from the receiver to the sender side cross-datacenter switch, constraining the congestion control loop for inter-datacenter traffic within a single datacenter. Furthermore, cdcPIM introduces a remote control point to perform admission control for inter-datacenter traffic entering the receiver’s datacenter, thus avoiding intra-datacenter congestion caused by traffic bursts. Simulations demonstrate that when cdcPIM manages inter-datacenter traffic while cooperating with datacenter quantized congestion notification (DCQCN) for intra-datacenter traffic, long-haul congestion is effectively mitigated. Under mixed cross-datacenter workloads, DCQCN + cdcPIM reduces the overall average flow completion time (FCT) slowdown by up to 25.7% and the P99 FCT slowdown of intra-DC flows by up to 65.0% compared with the baseline scheme Themis.

Wenqiang Deng, Junyan Chen, Xuefeng Huang et al. · 0 citations
Book Open access Aug 2026

Congestion Quarantine in Lossless Ethernet

Lossless Ethernet uses hop-by-hop backpressure to prevent buffer overflow and has become the mainstream choice for running Remote Direct Memory Access (RDMA) in AI and cloud data centers. Despite preventing congestion-induced drops, lossless networks introduce congestion contagion, which causes head-of-line blocking, congestion spreading, and deadlocks. Congestion control schemes have been introduced to mitigate the drawbacks of lossless Ethernet. However, congestion control mechanisms struggle with bursty traffic, face a dilemma, and can be sidelined by backpressure. In this paper, we propose congestion quarantine (CQ) as a complementary congestion management mechanism for lossless Ethernet. CQ uses a separate queue to quarantine congested flows, preventing congestion contagion, resolving the CC dilemma, and preventing CC from being sidelined. Results show that congestion quarantine can eliminate head-of-line blocking in scenarios with multiple congestion trees. Large-scale simulations demonstrate that CQ reduces the average and 99th percentile FCT slowdown of normal flows by 25–86% and 33.4–90%, respectively, with negligible impact on bursty traffic.

Dongkang Hu, Ran Shu, Wenxue Cheng et al. · 0 citations
Book Open access Aug 2026

InfiniFlow: Decoupling Virtual Channel Scalability from Buffer Requirements in Lossless Datacenter Networks

InfiniFlow is presented, a credit-based hop-by-hop flow control method that supports massive VCs with a limited buffer budget via per-port buffer sharing, and introduces a paradigm shift in buffer management: Upstream Allocates Buffer for Downstream (UABD).

Zerui Tian, Sen Liu, Minkun Xue et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.