Skip to content
Open access

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

Aug 2026 · ACM Transactions on Architecture and Code Optimization (TACO) · Vol 23, pp. 1 - 25 · 0 citations · 49 references

TL;DR

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Abstract

Large Language Models (LLMs) have enabled diverse AI applications. However, LLM inference imposes unprecedented computational and memory overhead, creating an inherent tradeoff between latency Service Level Objectives (SLOs) and resource constraints. Serverless computing, with on-demand provisioning and pay-as-you-go billing, is becoming a promising paradigm for LLM serving. But existing solutions fail to integrate state-of-the-art inference optimizations, resulting in suboptimal GPU utilization and prolonged latency. While Prefill-Decode (PD) disaggregation combined with continuous batching has resolved such inefficiencies in traditional cloud deployments, migrating these techniques to serverless makes two challenges particularly pronounced: (1) SLO-constrained resource provisioning for independently scaling prefill and decode phase functions, and (2) function lifespan management to mitigate resource waste from continuous batching-induced prolonged instance lifespans. To tackle these issues, we propose CELLServe, an SLO-aware and cost-efficient serverless LLM serving system that pioneers integrating PD disaggregation and continuous batching into serverless platforms. CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources. Comprehensive evaluations on five mainstream LLMs and real-world traces show that CELLServe achieves 1.71-1.85×higher request throughput than baselines under identical SLOs and GPU budgets, while sustaining high resource efficiency under dynamic workloads.

Read PDF

Similar papers

Preprint Sep 2026

SARA: SLO-Aware Resource Allocation for Disaggregated Agentic LLM Services

Recent advances in large language models (LLMs) are driving the emergence of multi-modal and agentic services for mobile users through cloud and edge infrastructures, where long-context workloads pose daunting challenges for inference latency. Existing disaggregated LLM serving systems largely rely on hardware profilin...

Shi-Cong Liu, Xiang-Hao Yu, Zheng-Run Gao et al. · 0 citations
Book Open access Sep 2026

Heterogeneous SLO Guaranteed Multi-Resource-Aware Batching in LLM Serving

In this paper, we study a mixed-prompt scenario—where both short and long prompts coexist—in an LLM inference serving system that supports diverse applications with heterogeneous iteration-time SLOs. To improve throughput for long prompts, prior work divides them into chunks and batches requests or chunks to meet the t...

Hai-Ying Shen, Tanmoy Sen, Yuxiong He · 0 citations
Preprint Aug 2026

Cascade: Exploiting SLO-Aware latency budget for fair and high goodput LLM inference serving

Cascade, an LLM serving system that estimates and continuously updates this per-request latency budget from request characteristics, KV-cache state, and current system load, and uses a single per-request budget to jointly coordinate request scheduling and KV-cache management across the memory hierarchy.

Muhammad Adnan, R. Mahapatra, Prashant J. Nair et al. · 0 citations
Preprint Aug 2026

OpScale: Operator-level Provisioning and Autoscaling for LLM Serving

OpScale is presented, a practical operator-level orchestration framework of profiling, provisioning, placement, and runtime serving that attains SLOs with up to 36.3% fewer GPUs and 28% less power, or achieves 44% higher throughput under fixed cost budgets.

Xingqi Cui, Chieh-Jan Mike Liang, Ziang T. Tang et al. · 0 citations
Preprint Sep 2026

PackServe: SLO-Aware Request Scheduling for Agentic LLM Serving at Scale

Request scheduling is a key challenge in large-scale clusters serving agentic large language model (LLM) workloads. An effective scheduler must preserve key-value cache (KVC) reuse across long, shared prefixes, meet token-level latency service-level objectives (SLOs), and minimize GPU resource footprint. Existing sched...

Zhi-Yuan Tan, De-Jiang Zhu, Jing-Zhe Jiang et al. · 0 citations
Book Open access Sep 2026

WAQ-LLM: Optimizing Multi-Instance LLM Deployment via Workload-Aware Queueing Model

The deployment of Large Language Models (LLMs) on multi-instance GPU clusters has become essential to meet the surging demand for generative AI applications. While scaling out instances increases throughput, the distinct computational characteristics of prefill and decode phases introduce significant resource contentio...

Jia-Xin Lai, Yi-Zhou Luo, Qiang Wang · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.