Jul 2026· International Conference on Supercomputing· pp. 1128-1140· 0 citations· 43 references
Computer Science
TL;DR
PPPR is proposed, a Phase-aware Peak Power Reduction framework for minimizing the CapEx of LLM inference that features two novel LLM-specific designs and phase-aware timing that shaves more peak power and leads to lower CapEx.
Abstract
The rapid growth of large language model (LLM) inference services has been driving data centers to quickly increase their hosted GPUs and servers, causing data centers to approach their peak power capacities. This is a serious challenge because 1) expensive upgrades of power facilities must be done to existing data centers or 2) new data centers must be constructed, both of which require a significant increase of capital expense (CapEx). Existing solutions for peak power reduction include frequency throttling and power capping, which both rely on GPU frequency throttling. Hence, when applied directly to LLM inference, they can degrade performance and increase latency. While energy storage can be exploited to reduce peak power draws without hurting performance, prior solutions are not designed for LLM inference, which has a unique and recurring two-phase power profile with short-duration, high-power prefill phases followed by longer, low-power decode phases. In this paper, we propose PPPR, a Phase-aware Peak Power Reduction framework for minimizing the CapEx of LLM inference. PPPR features two novel LLM-specific designs. First, in contrast to prior work that usually relies on preset thresholds to decide when storage should discharge or recharge, PPPR leverages LLM phase information to make better decisions. This phase-aware timing shaves more peak power and leads to lower CapEx. Second, based on real-world LLM traces, a hybrid storage architecture is designed with server-level supercapacitors to handle short prefill spikes and rack-level batteries to smooth longer, aggregated power changes. This novel placement and sizing strategy can lead to more CapEx savings for LLM inference. Our hardware evaluation with representative LLM workloads shows that PPPR achieves 1.29 × more peak reduction, on average, than three state-of-the-art baselines. Extensive trace-driven simulations also show that PPPR achieves up to 3.53 × higher CapEx savings and up to 1.52 × longer UPS lifetime than the baselines.
Results show that energy-aware serving should jointly optimize both request energy and token energy, rather than only reducing per-token energy cost, and substantially narrows the dense-vs-MoE token-energy gap.
P. Vellaisamy, Vanessa Lam, Shawn Blanton et al.· 2 citations
Hydra is presented, a common-schema, phase-aware workload characterization framework for LLM inference on edge SoCs that enables reproducible, phase-aware characterization of edge LLM inference.
Amir Taherin, Sana Taghipour Anvari, Charles Amante et al.· 2 citations
Large language model (LLM) training is among the fastest-growing sources of electricity demand in modern data centers, and power availability is a primary bottleneck to continued AI infrastructure growth. Making the power consumption of these workloads flexible could unlock additional power for AI growth, limit increas...
Philip Colangelo, Charles Dawson, Shayan Sengupta et al.· 0 citations
This work introduces a novel LLM-based predictive scheduling system designed to enhance operational efficiency while reducing the environmental impact of data centers, using an LLM to predict key metrics such as execution time and energy consumption from source code.
Hanzhao Wang, Jingxuan Wu, Yumeng Li et al.· 0 citations
This study presents a cross-platform, multi-model empirical study, where several important observations are brought, including the contrastive effect of quantization under different hardware bottlenecks, along with a quantification of runtime delays caused by the lack of parallelism in the ARM architecture.
Subhransu Das, Jiaming Cheng, Swathi Vallabhajosyula et al.· Practice and Experience in A...· 0 citations
Datacenter GPU power is the binding constraint on LLM serving capacity, and production serving has shifted to prefill/decode (PD) disaggregation. Deploying NVIDIA's Max-Q inference profile on a disaggregated B200 system, we found its realized gain modest (+8.6% tokens/J), model-dependent, and carrying a mean end-to-end...
Jae Gon Kim, Donghoon Yoo, Hanyul Ryu et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.