Skip to content
Book Open access

WAQ-LLM: Optimizing Multi-Instance LLM Deployment via Workload-Aware Queueing Model

Sep 2026 · Proceedings of the International Conference on Parallel Processing · 0 citations · 31 references

Abstract

The deployment of Large Language Models (LLMs) on multi-instance GPU clusters has become essential to meet the surging demand for generative AI applications. While scaling out instances increases throughput, the distinct computational characteristics of prefill and decode phases introduce significant resource contention, making it challenging to satisfy stringent Service Level Objectives (SLOs) for responsiveness and generation speed. Existing serving solutions typically rely on static deployment strategies—either aggregating or disaggregating these phases—which often fail to adapt to the interference patterns caused by varying request arrival rates and sequence lengths. In this work, we propose WAQ-LLM, a performance optimization framework to find the optimal deployment configuration for multi-instance LLM serving. Specifically, we first establish an analytical workload-aware queueing model that captures the LLM computational characteristics and queuing behavior of both aggregation and disaggregation designs. We then formulate the deployment configuration problem as a constrained optimization problem and develop a polynomial-time algorithm to efficiently identify configurations that minimize Time-per-Output-Token (TPOT) while satisfying SLOs. Extensive experiments are conducted on a 16-GPU cluster across various LLMs and workload settings. WAQ-LLM consistently outperforms AIConfigurator, NVIDIA’s official product-level solution, reducing TPOT by 21.62% on average and up to 81.7%. Moreover, WAQ-LLM can provide adaptive and hybrid configurations to handle diverse workloads, surpassing static deployments limited to a single instance type (either PD aggregation or disaggregation).

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.