Skip to content
Preprint

ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs

Sep 2026 · 0 citations · 31 references
Engineering Computer Science

Abstract

Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path. Prefill and decode therefore consume shared, time-varying thermal headroom, yet vendor governors react only near hardware throttling thresholds without knowledge of request state or upcoming work. A control decision that improves current performance can thus consume headroom too quickly and degrade subsequent service. In this paper, we present ThermE, a runtime system that predicts and jointly manages shared thermal headroom for sustained LLM inference on edge SoCs. Its Fast LLM-to-Heat Compiler maps the model, requests, runtime state, and candidate actions to domain heats without executing LLMs. A partial differential equation (PDE)-constrained Headroom Predictor uses ThermPINN for offline thermal identification and a Reduced Headroom Predictor (RHP) for low-overhead online uncertainty-calibrated headroom forecasts. An Uncertainty-Aware Action Scheduler then selects actions that balance serving quality and future headroom. We implement ThermE atop vLLM and evaluate it across four LLM inference workloads. The results show that ThermE reduces TTFT and TPOT by 40.55% and 12.17%, respectively, relative to vLLM. It achieves a 5.70% SLO violation rate, compared with 12.30% for the strongest baseline, while its predictor obtains a 1.94 $^\circ$C MAE with 11.50 ms overhead.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.