ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs
Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path. Prefill and decode therefore consume shared, time-varying thermal headroom, yet vendor governors react only near hardware throttling thresholds without know...