Jul 2026· ACM SIGOPS Operating Systems Review· Vol 60, pp. 30 - 40· 0 citations· 47 references
Computer Science
TL;DR
This paper proposes a deployment scheme for agentic workloads tailored for serverless, accompanied by pre-warming policies that minimize the idle resource footprint and startup latencies and outlines promising research directions for serverless agents.
Abstract
Accelerating generative AI adoption has driven the expansion of data centers, which amass GPUs, DRAM, and SSDs to feed emerging, resource-hungry AI workloads. The serverless cloud model offers a path to improve application resource efficiency by loading instances on demand. However, the suitability of emerging AI workloads for serverless remains insufficiently explored. We survey the state-of-the-art in serverless hosting for LLM applications and find that: (1) Despite advances in serverless LLM hosting, model loading and initialization processes still dominate startup latency. (2) Agentic AI workloads have not yet been characterized under the serverless context. We propose a deployment scheme for agentic workloads tailored for serverless, accompanied by pre-warming policies that minimize the idle resource footprint and startup latencies. This paper outlines promising research directions for serverless agents.
This work organizes agentic workflows in a taxonomy and presents its first architectural characterization with a production study at Microsoft Azure and a controlled study of open-source frameworks, showing that agentic execution is fragmented and heterogeneous.
Ji-Rong Yang, Peizhe Liu, Chaojie Zhang et al.· 2 citations
CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.
Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al.· ACM Transactions on Architec...· 0 citations
This work introduces WASP, a configurable framework that brings stateful serverless execution to the edge-cloud continuum by abandoning monolithic architectures in favor of strictly decoupled, pluggable components, and lets system administrators swap the WASM runtime and the datastore to fit available resources and app...
Text-to-image (T2I) workflows are increasingly deployed on serverless platforms because users often compose customized workflows and invoke them intermittently. Existing platforms typically deploy each workflow as an opaque GPU function, provisioning, placing, and scaling all constituent models in the workflow together...
Xiaoxiao Jiang, Suyi Li, Sheng Yao et al.· arXiv.org· 0 citations
Public LLM services serve diverse multi-agent applications with varying workflow dependencies and performance requirements. Requests generated by these applications often exhibit commonality and interdependence, yet current systems largely ignore such application-level structure. As a result, at the LLM engine cluster...
Uttam Rao, Ali Zafar Sadiq, Hai-Ying Shen et al.· International Conference on...· 0 citations
Agentic LLM workflows issue many dependent calls with unpredictable resource demand, causing queue buildup and latency degradation on shared serving backends when left unmanaged. In this paper, we propose CALM-MAS, a congestion-aware serving framework for LLM applications that treats LLM test-time computation as an ela...
Mouheb Ben Nasr, Muhammad Bilal, Alessandro Cornacchia et al.· Proceedings of the 17th ACM...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.