Heterogeneous SLO Guaranteed Multi-Resource-Aware Batching in LLM Serving
Abstract
In this paper, we study a mixed-prompt scenario—where both short and long prompts coexist—in an LLM inference serving system that supports diverse applications with heterogeneous iteration-time SLOs. To improve throughput for long prompts, prior work divides them into chunks and batches requests or chunks to meet the token budget to fully utilize GPU compute resource. Our measurement analysis shows that it fails to meet heterogeneous iteration-time SLOs or fully utilize both GPU compute and memory resources. Also, delayed KV cache release from processing multiple long prompts may reduce throughput. To address the limitations, based on our findings, we propose a heterogeneous SLO guaranteed Multi-Resource-aware batching system (MuRa). MuRa orders queued requests by SLOs, enabling it to prioritize tight-SLO requests and batch those with similar SLOs to improve throughput. It selectively chooses queued requests to maximize both compute and memory utilizations, and processes all chunks from one prompt before moving to the next to enable earlier KV cache release. Trace-driven real experiments demonstrate that MuRa achieves 1.42-11.21 × higher throughput, 1.43-13.71 × higher goodput, 37-90% higher SLO attainment, and 1.61-12.22 × lower response latency compared to the state-of-the-art approaches. It achieves performance near Oracle, which optimally maximizes goodput.