Skip to content
Book Open access

FaaSlim: Partial Caching of Snapshot-based VMs for Serverless Computing

Jul 2026 · International Conference on Supercomputing · pp. 54-66 · 0 citations · 42 references
Computer Science

TL;DR

FaaSlim is proposed, a snapshot-based serverless computing system that integrates both in-memory caching and snapshot-based in-storage caching to reduce cold start latency, and devise a partial caching-aware eviction policy, GDSF-CE, which selects VMs and page subsets to evict based on cache efficiency.

Abstract

To mitigate the cold-start overhead of serverless functions, two orthogonal approaches have been studied: in-memory caching and snapshot-based in-storage caching. In this work, we propose FaaSlim, a snapshot-based serverless computing system that integrates both approaches to reduce cold start latency. FaaSlim classifies VM pages into three categories, read, write, and permission update, based on page fault types, enabling partial caching of only a selected subset of pages rather than the entire VM while reclaiming the rest. When the VM is reused, only reclaimed pages are fetched from disk, reducing page fault overhead and function latency, while requiring less memory during idle periods. We also devise a partial caching-aware eviction policy, GDSF-CE, which selects VMs and page subsets to evict based on cache efficiency, a metric that quantifies the benefit of partial caching by relating latency reduction to memory consumption. Our evaluation using real-world traces shows that FaaSlim reduces the total overhead of snapshot-based cold starts by 23.0–27.8% compared to the state-of-the-art combination of in-storage and in-memory caching, FaaSnap with CIDRE’s GDSF-C, across diverse serverless workloads.

Read PDF

Similar papers

Preprint Sep 2026

Gutenberg: Taming Latency-Critical Cloud Services with Near-Data-Processing

Latency-critical cloud services place growing pressure on memory while requiring isolation, fairness, and predictable QoS. Near-data processing (NDP) reduces data movement by executing requests close to memory, and prior systems further improve locality through caching and replication. However, writes make replica main...

Qi Lin, Phillip B. Gibbons, Jovan Stojkovic et al. · 0 citations
Conference Aug 2026

Poster: AutoThermKV: An Efficient User-Transparent In-Memory Management of Hot Data for Key-Value Stores

Log-structured merge-tree (LSM-tree) key-value stores rely on caching to mitigate multi-component lookups and long tails, yet block and KP caches are prone to compactioninduced expiry, and all three cache types suffer from scan pollution under LRU eviction. We present AutoThermKV, a two-level in-memory architecture tha...

Yunfan Chi, E. Sha, Longshan Xu et al. · 0 citations
Conference Aug 2026

Characterizing Predictability–Latency Trade-offs of KV-Cache SSD Offloading in LMCache for LLM Serving Systems

KV-cache offload is widely used to stretch GPU memory for LLM serving, but its storage behavior has not been characterized at the block-device level. In this paper, we study LMCache through realworld multi-session workloads that span same/different context $\times$ same/different prompt, using over 100 stateless reques...

Ying He, Dingsen Shi, Yanbo Dai et al. · 0 citations
Conference Aug 2026

Paging-Resilient Prefetching in Flash-Based CXL SSDs

CXL SSDs extend system memory using NAND flash, providing a scalable solution to the capacity and bandwidth limits of memory-intensive, multi-tenant cloud services. Since CXL SSDs are directly addressable by the CPU, SSD-internal prefetching is crucial for hiding flash-grade latency from the host. However, host-side pa...

Chung-Min Yu, Chih-Kang Yeh, Ying-Shuo Lin et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.