Stateless Temporal-Interleaved App-Aware RDMA
Abstract
Existing RDMA transports for AI fabrics compromise on scalability, robustness, or hardware overhead. Lossless protocols suffer from PFC-induced congestion spreading under extreme incast, while switch-assisted lossy transports collapse into catastrophic Retransmission Timeout (RTO) storms under silent physical link errors. Furthermore, supporting packet-level load balancing typically requires prohibitive SRAM for out-of-order tracking. We present STAR, a purely end-to-end, stateless, and receiver-driven RDMA architecture. STAR employs chunk-level token pacing to temporally interleave concurrent flows, intrinsically suppressing incast queues and eliminating PFC dependency. By extending Direct Data Placement (DDP) semantics for idempotent memory scatter, it absorbs out-of-order packets with $\mathcal{O}(1)$ hardware state. To handle rare physical drops, an app-aware software watchdog performs full-chunk retransmission, circumventing hardware RTO storms. Evaluations demonstrate that STAR scales linearly under extreme 63-to-1 incast, introduces negligible (~2.5%) collective completion time overhead in ideal lossless fabrics, and prevents catastrophic synchronization barrier stalling under physical link errors.