Skip to content

Author

Jia-Xin Lai

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

WAQ-LLM: Optimizing Multi-Instance LLM Deployment via Workload-Aware Queueing Model

The deployment of Large Language Models (LLMs) on multi-instance GPU clusters has become essential to meet the surging demand for generative AI applications. While scaling out instances increases throughput, the distinct computational characteristics of prefill and decode phases introduce significant resource contentio...

Jia-Xin Lai, Yi-Zhou Luo, Qiang Wang · 0 citations
#large language models Book Open access Sep 2026

WAQ-LLM: Optimizing Multi-Instance LLM Deployment via Workload-Aware Queueing Model

WAQ-LLM is proposed, a performance optimization framework to find the optimal deployment configuration for multi-instance LLM serving that can provide adaptive and hybrid configurations to handle diverse workloads, surpassing static deployments limited to a single instance type.

Jia-Xin Lai, Yi-Zhou Luo, Qiang Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.