PLINK: A GPU-Initiated I/O Platform Exploiting NVMe Parallelism
Abstract
As LLM inference scales and KV-cache pressure forces data onto NVMe storage, the frequency and fine-grained nature of GPU-initiated storage accesses have grown substantially. Existing GPU-initiated I/O platforms do not fully exploit NVMe’s multi-queue parallelism because their serialized submit-then-poll path keeps only a single command outstanding per thread; moreover, non-uniform per-thread I/O completion latency induces warp divergence that wastes a substantial fraction of GPU clock cycles. Guided by a cycle-granularity breakdown of BaM’s I/O path that localizes over 93% of I/O turnaround time to three phases, we present PLINK (Parallel-Link), a GPU-initiated I/O platform that addresses this gap through three targeted optimizations: SQ/CQ decoupling to hide device latency, command-ID-indiscriminate polling to eliminate exhaustive CQ search overhead, and warp-level doorbell batching to minimize PCIe MMIO writes and reduce warp divergence. An evaluation on a state-of-the-art GPU paired with a PCIe Gen 6 NVMe SSD, in which each optimization is applied one at a time on top of the BaM baseline, shows that PLINK’s I/O backend reaches SSD IOPS saturation at 136 total warps—submission and completion combined—versus BaM’s 238.