Skip to content
#edge computing Preprint

HeatCache: Thermal-aware Energy-efficient LLM Inference Scheduling for Chassis-level Liquid Cooling in Sustainable Edge Server Rooms

Sep 2026 · 0 citations · 38 references
Computer Science Engineering

TL;DR

This paper implements HeatCache atop vLLM and shows that it reduces computing energy by up to 18.0%, decreases thermal-throttle exposure by 81.7%, and maintains SLO violation rates below 0.9% even up to $48~^{\circ}\mathrm{C}$.

Abstract

LLM inference is increasingly deployed at institution-scale edges to meet service requirements. However, multi-GPU inference consumes a large amount of electricity and produces substantial heat. To improve sustainability, operators and regulations often demand raising the ambient setpoint to reduce cooling electricity. This can increase thermal throttling and hardware aging, leading to Service-Level Objective violations. In this paper, we present HeatCache, a thermal-aware, energy-efficient LLM inference scheduler for commercial chassis-level AIO liquid-cooled GPUs at sustainable ambient temperatures. HeatCache treats AIO loops as a temporary heat buffer, measured by heat budget and schedules requests to minimize energy subject to thermal safety and SLO constraints, based on an electrical-informed heat-demand estimation from HeatiTS. We implement HeatCache atop vLLM and show that it reduces computing energy by up to 18.0%, decreases thermal-throttle exposure by 81.7%, and maintains SLO violation rates below 0.9% even up to $48~^{\circ}\mathrm{C}$.

View source

Similar papers

Preprint Sep 2026

ETCInfer: An Energy-efficient Thermal-aware Cooling-joint Scheduler for LLM Inference in AI Datacenters

Large language model (LLM) inference in AI datacenters creates a coupled control problem between GPU serving and facility cooling. Raising ambient temperature setpoints can reduce cooling energy and carbon, but also shrinks thermal headroom, induces GPU throttling, and leads to Service-Level-Objective (SLO) violations....

Rui Lu, Rui Ge, Huang-Huang Liang et al. · 0 citations
Preprint Sep 2026

ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs

Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path. Prefill and decode therefore consume shared, time-varying thermal headroom, yet vendor governors react only near hardware throttling thresholds without know...

Rui Lu, Yu-Heng Wang, Bo-Zheng-Hong-Fan Liu · 0 citations
Aug 2026

Experimental Evaluation of Cooling Loop Configurations for Server-Level Direct-to-Chip Liquid Cooling Systems

Effective heat removal has become a critical challenge in modern data centers as increasing computational demands from artificial intelligence and machine learning push GPUs, CPUs, and network switches toward unprecedented power densities. These heat loads often exceed the capabilities of traditional air-cooling solu...

Ali Heydari, Abdallah Soud, Mohammad I. Tradat et al. · 0 citations
Preprint Sep 2026

Job Class Thermal Intent Aware Liquid Cooling Allocation for AI Data Centers

GPU-dense AI data centers need to run on liquid cooling as air simply cannot shed the heat at these power densities. Yet the cooling loops themselves are blind to what workloads are about to run; they crank up flow only after a sensor catches a temperature climb, which can take 30 to 50 seconds. We built Job-Class Ther...

Krishna Chaitanya Sunkara · 0 citations
Preprint Aug 2026

A Smallest-Need-First Job Scheduling Framework with Adaptive Optimization of Idle Node Counts for Energy-Efficient HPC Systems

Power-state management in high-performance computing (HPC) clusters must reduce idle energy without excessive wake-up delays for rigid parallel jobs. This paper presents SNF-ICON, an event-driven controller combining smallest-need-first (SNF) gang scheduling, predictive wake timing, and adaptive warm-spare control. At...

Reza Pulungan, Raka Satya Prasasta, Santana Yuda Pradata et al. · 0 citations
Conference Aug 2026

Thermal-Aware GPU Energy Optimization for On-Device LLM Split Fine-Tuning

Split fine-tuning divides a large language model (LLM) between a device and an edge server at the base station, with the front layers on the device and the remaining layers on the server. However, the device's limited thermal dissipation can easily cause graphics processing unit (GPU) overheating, which in turn increas...

Zu-Guang Li, Wen Wu, Shao-Hua Wu et al. · 0 citations

Related blog posts

Microsoft Research Blog Sep 29, 2026

Introducing Quine: An AI research system designed for the complexity of biology

Biology doesn't operate in silos, and neither should the AI representation of it. Quine is an early-stage research effort to create a multimodal world model of biology. By connecting insights across biological scales and modalities, Quine helps scientists computationally search a space far larger than intuition allows and prioritize hypotheses before they reach the lab. Experimental results provide important feedback, helping researchers sharpen future research directions. The post Introducing Q…

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.