WAQ-LLM: Optimizing Multi-Instance LLM Deployment via Workload-Aware Queueing Model
The deployment of Large Language Models (LLMs) on multi-instance GPU clusters has become essential to meet the surging demand for generative AI applications. While scaling out instances increases throughput, the distinct computational characteristics of prefill and decode phases introduce significant resource contentio...