Dissecting GPU Utilization for LLM Inference on Nvidia Hopper
This paper profiles vLLM with FlashAttention-3 and cuBLASLt on an H100 NVL across cold prefill, warm prefill, and decode, sweeping sequence length and batch size, and replaces the usual single utilization number with eight counter-validated views derived from raw Nsight Compute reports.
Mohammad Siavashi, Gerald Q. Maguire, Dejan Kostic et al.
· 0 citations