Skip to content

MAPS: Memory-Aware Predictive Scheduling Framework for Large Language Model Serving

Sep 2026 · 0 citations · 37 references
Computer Science

Abstract

The surge of large language model (LLM) applications on personal devices imposes massive, bursty workloads on cloud serving infrastructure. While prefill-decode disaggregation improves throughput and scalability, memory-bound decode instances often suffer from persistent load imbalance, as output lengths are unknown when requests arrive at the cloud. To address this, we propose MAPS, a Memory-Aware Predictive Scheduling framework tailored for disaggregated LLM serving. MAPS performs device-assisted speculative output length prediction overlapped with cloud-side prefilling, incurring negligible latency overhead. To handle generation uncertainty, MAPS applies uncertainty-aware calibration to derive output-length upper bounds with target coverage, enabling safe scheduling decisions. Building on these bounds, MAPS employs a hierarchical global-local scheduling strategy to mitigate inter-decoder queue buildup and intra-decoder head-of-line blocking. Extensive experiments on two real-world workloads and two LLMs show that MAPS significantly outperforms three state-of-the-art systems, reducing average end-to-end latency by 42.6 and tail latency by up to 84.8.

View source

Similar papers

Book Open access Aug 2026

Turbo: Efficiently Serving Long-Context Large Language Models with In-Network Aggregation

This work proposes Turbo, a first-of-its-kind in-network aggregation system that accelerates long-context inference by offloading query broadcast and attention aggregation to switches and introduces a rolling forward scheme that propagates states to enable cross-stage updates.

Ying Wan, Yuchen Xu, Chuwen Zhang et al. · 0 citations
#machine learning Preprint Sep 2026

Mira: Memory-Efficient MoE Inference Using Adaptive Caching and Predictive Expert Staging

Mixture-of-Experts (MoE) models are a compelling architecture for scaling model capacity, making them especially attractive for deployment on resource-constrained, single-GPU systems. However, this benefit is difficult to realize because expert parameters dominate memory, and token-level routing is dynamic, unpredictab...

Sanjali Yadav, Bahar Asgari · 0 citations
#machine learning Preprint Sep 2026

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

Large language models (LLMs) are moving onto mobile devices for increasingly diverse workloads over text, images, video, and audio. These applications often require long contexts, making the Key-Value (KV) cache a dominant memory bottleneck because it grows linearly with sequence length and is accessed at every decodin...

Zhi-Hao Shu, Md Musfiqur Rahman Sanim, Jie Hu et al. · 0 citations
Open access Aug 2026

CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments

CELLServe formalizes SLO-constrained joint resource provisioning as an optimization problem with a dedicated algorithm, and introduces an opportunistic instance merging strategy for decode phase functions to reclaim fragmented resources.

Ze-Jian Wang, Nan Lin, Zi-Nuo Cai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SeqMoE: Toward Full-Load Performance via Predictive and Graph-Compatible MoE Offloading

This work is the first to recast expert activation prediction as sequence modeling, enabling accurate multi-step, multi-layer forecasts that provide a long and reliable window for downstream decisions, and derives general runtime principles encompassing compute-transparent expert placement and synchronization-free orch...

Zi-Han Wang, Yu-Qi Wang, Lei Gong et al. · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.