Adaptive GPU Sharing for Real-time LLM Serving with Best-effort Workloads
Abstract
Large Language Models (LLMs) are increasingly deployed in latency-sensitive applications, where real-time serving must satisfy stringent service-level objectives (SLOs). However, request intensities fluctuate over time, and under low load LLM services leave a substantial fraction of GPU task idle. A promising approach to reclaim this capacity is co-locating best-effort workloads with LLM serving via GPU sharing. In this study, we propose AdMix, a scheduling system that enables adaptive GPU sharing between real-time LLM serving and best-effort non-LLM deep learning (DL) inference. AdMix dynamically regulates the best-effort workload’s resource allocation using NVIDIA Multi-Process Service (MPS), guided by real-time monitoring of LLM performance and a lightweight latency estimator. By real-world LLM serving traces, experiments on LLaMA-7B and Qwen-14B paired with representative DL workloads demonstrate that AdMix improves best-effort throughput by 1.3–5 × over MPS baseline while preserving LLM SLO attainment.