Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

Delphinus: Ultra-Fast Link Failure Detection and Recovery for AI Data Center Networks

High-performance artificial intelligence (AI) applications impose stringent reliability requirements on AI data center networks (DCNs), yet link failures are almost inevitable and can severely disrupt AI workloads such as large language model (LLM) training and inference. Existing deployed link failure detection and recovery mechanisms suffer from slow execution speed and limited failure coverage, failing to meet the demands of production AI DCNs. To address these issues, we propose Delphinus, an ultra-fast link failure detection and recovery solution built on the data-plane of programmable switches. It achieves ultra-fast failure detection via hardware-based port state monitoring, extends recoverable failure coverage through remote failure notification and relay, and enables fast recovery by path switchover. Delphinus can serve as a key generic function of switches, providing host-transparent link failure handling for Ethernet fabrics. We implement Delphinus on commercial hardware switches, and deploy it in large-scale production AI DCNs for over a year. Extensive evaluations demonstrate that Delphinus can complete link failure detection and recovery within sub-milliseconds, with negligible impact on application performance and imperceptible service interruption.

Junye Zhang, Jie Li, Zhigang Ji et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.