Characterizing Tensor-Parallel Crossover in vLLM Serving on Dual NVIDIA Tesla T4 GPUs
Adding a second GPU to a language-model server divides model work and memory but also introduces synchronization, collective communication, scheduling, and process overhead. We characterize the conditions under which two-way tensor parallelism (TP2) improves upstream vLLM serving on a pinned, PHB-connected pair of NVID...