Skip to content
Book Open access

PLINK: A GPU-Initiated I/O Platform Exploiting NVMe Parallelism

Sep 2026 · Proceedings of the 18th ACM Workshop on Hot Topics in Storage and File Systems · 0 citations · 4 references

Abstract

As LLM inference scales and KV-cache pressure forces data onto NVMe storage, the frequency and fine-grained nature of GPU-initiated storage accesses have grown substantially. Existing GPU-initiated I/O platforms do not fully exploit NVMe’s multi-queue parallelism because their serialized submit-then-poll path keeps only a single command outstanding per thread; moreover, non-uniform per-thread I/O completion latency induces warp divergence that wastes a substantial fraction of GPU clock cycles. Guided by a cycle-granularity breakdown of BaM’s I/O path that localizes over 93% of I/O turnaround time to three phases, we present PLINK (Parallel-Link), a GPU-initiated I/O platform that addresses this gap through three targeted optimizations: SQ/CQ decoupling to hide device latency, command-ID-indiscriminate polling to eliminate exhaustive CQ search overhead, and warp-level doorbell batching to minimize PCIe MMIO writes and reduce warp divergence. An evaluation on a state-of-the-art GPU paired with a PCIe Gen 6 NVMe SSD, in which each optimization is applied one at a time on top of the BaM baseline, shows that PLINK’s I/O backend reaches SSD IOPS saturation at 136 total warps—submission and completion combined—versus BaM’s 238.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.