Skip to content
Book Open access

Phase-aware Peak Power Reduction for Minimizing the Capital Expense of LLM Inference

Jul 2026 · International Conference on Supercomputing · pp. 1128-1140 · 0 citations · 43 references
Computer Science

TL;DR

PPPR is proposed, a Phase-aware Peak Power Reduction framework for minimizing the CapEx of LLM inference that features two novel LLM-specific designs and phase-aware timing that shaves more peak power and leads to lower CapEx.

Abstract

The rapid growth of large language model (LLM) inference services has been driving data centers to quickly increase their hosted GPUs and servers, causing data centers to approach their peak power capacities. This is a serious challenge because 1) expensive upgrades of power facilities must be done to existing data centers or 2) new data centers must be constructed, both of which require a significant increase of capital expense (CapEx). Existing solutions for peak power reduction include frequency throttling and power capping, which both rely on GPU frequency throttling. Hence, when applied directly to LLM inference, they can degrade performance and increase latency. While energy storage can be exploited to reduce peak power draws without hurting performance, prior solutions are not designed for LLM inference, which has a unique and recurring two-phase power profile with short-duration, high-power prefill phases followed by longer, low-power decode phases. In this paper, we propose PPPR, a Phase-aware Peak Power Reduction framework for minimizing the CapEx of LLM inference. PPPR features two novel LLM-specific designs. First, in contrast to prior work that usually relies on preset thresholds to decide when storage should discharge or recharge, PPPR leverages LLM phase information to make better decisions. This phase-aware timing shaves more peak power and leads to lower CapEx. Second, based on real-world LLM traces, a hybrid storage architecture is designed with server-level supercapacitors to handle short prefill spikes and rack-level batteries to smooth longer, aggregated power changes. This novel placement and sizing strategy can lead to more CapEx savings for LLM inference. Our hardware evaluation with representative LLM workloads shows that PPPR achieves 1.29 × more peak reduction, on average, than three state-of-the-art baselines. Extensive trace-driven simulations also show that PPPR achieves up to 3.53 × higher CapEx savings and up to 1.52 × longer UPS lifetime than the baselines.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

Characterizing Job Power Elasticity for Power-Flexible AI Training

Large language model (LLM) training is among the fastest-growing sources of electricity demand in modern data centers, and power availability is a primary bottleneck to continued AI infrastructure growth. Making the power consumption of these workloads flexible could unlock additional power for AI growth, limit increas...

Philip Colangelo, Charles Dawson, Shayan Sengupta et al. · 0 citations
Preprint Aug 2026

LLM-Powered Predictive Decision-Making for Sustainable Data Center Operations

This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.

Hanzhao Wang, Jingxuan Wu, Yumeng Li et al. · 0 citations
Book Open access Jul 2026

Phase-Wise Analysis of LLM Inference Acceleration on GPU, CPU, and Edge Device

This study presents a cross-platform, multi-model empirical study, where several important observations are brought, including the contrastive effect of quantization under different hardware bottlenecks, along with a quantification of runtime delays caused by the lack of parallelism in the ARM architecture.

Subhransu Das, Jiaming Cheng, Swathi Vallabhajosyula et al. · 0 citations
#machine learning Preprint Sep 2026

Phase-Decoupled, Model-Calibrated Power Control for Disaggregated LLM Serving

Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end...

Jae Gon Kim, Donghoon Yoo, Hanyul Ryu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.