WAQ-LLM: Optimizing Multi-Instance LLM Deployment via Workload-Aware Queueing Model
Abstract
The deployment of Large Language Models (LLMs) on multi-instance GPU clusters has become essential to meet the surging demand for generative AI applications. While scaling out instances increases throughput, the distinct computational characteristics of prefill and decode phases introduce significant resource contention, making it challenging to satisfy stringent Service Level Objectives (SLOs) for responsiveness and generation speed. Existing serving solutions typically rely on static deployment strategies—either aggregating or disaggregating these phases—which often fail to adapt to the interference patterns caused by varying request arrival rates and sequence lengths. In this work, we propose WAQ-LLM, a performance optimization framework to find the optimal deployment configuration for multi-instance LLM serving. Specifically, we first establish an analytical workload-aware queueing model that captures the LLM computational characteristics and queuing behavior of both aggregation and disaggregation designs. We then formulate the deployment configuration problem as a constrained optimization problem and develop a polynomial-time algorithm to efficiently identify configurations that minimize Time-per-Output-Token (TPOT) while satisfying SLOs. Extensive experiments are conducted on a 16-GPU cluster across various LLMs and workload settings. WAQ-LLM consistently outperforms AIConfigurator, NVIDIA’s official product-level solution, reducing TPOT by 21.62% on average and up to 81.7%. Moreover, WAQ-LLM can provide adaptive and hybrid configurations to handle diverse workloads, surpassing static deployments limited to a single instance type (either PD aggregation or disaggregation).