Skip to content

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

Power-Centric Observability for HPC Systems: Design, Deployment, and Evaluation on REPACSS

Energy efficiency is a critical concern in modern heterogeneous HPC systems, where CPUs, GPUs, and dense cooling infrastructure increase both idle and peak energy demand. This paper presents a power-centric observability framework deployed on the NSF REPACSS system that enables non-intrusive and consistent energy analysis across facility, rack, node, and job scopes. The framework collects instantaneous power telemetry exclusively from out-of-band hardware sources, including in-row cooling units, rack power distribution units, and compute nodes accessed via iDRAC. These measurements are stored in a time series database and processed through a power-centric energy model that aligns heterogeneous telemetry in time and space. A scalable query engine and web API provide a uniform interface for deriving energy metrics across infrastructure and job contexts, while a lightweight Slurm epilog integration enables optional job level energy reporting without application modification or runtime instrumentation. Using this unified telemetry and query framework, we conduct a multi-scope energy characterization of the NSF REPACSS data center. The results show stable facility efficiency, efficient in-row cooling behavior, distinct differences in idle and peak power between CPU and GPU nodes, and diverse non-CPU energy contributions in representative workloads. Together, these findings demonstrate that unified out-of-band telemetry combined with a power-centric energy model, enables practical cross-scope energy observability in production HPC systems.

Yongjia Zhao, Jie Li, Chenxu Niu et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.