Power-Centric Observability for HPC Systems: Design, Deployment, and Evaluation on REPACSS
Energy efficiency is a critical concern in modern heterogeneous HPC systems, where CPUs, GPUs, and dense cooling infrastructure increase both idle and peak energy demand. This paper presents a power-centric observability framework deployed on the NSF REPACSS system that enables non-intrusive and consistent energy analysis across facility, rack, node, and job scopes. The framework collects instantaneous power telemetry exclusively from out-of-band hardware sources, including in-row cooling units, rack power distribution units, and compute nodes accessed via iDRAC. These measurements are stored in a time series database and processed through a power-centric energy model that aligns heterogeneous telemetry in time and space. A scalable query engine and web API provide a uniform interface for deriving energy metrics across infrastructure and job contexts, while a lightweight Slurm epilog integration enables optional job level energy reporting without application modification or runtime instrumentation. Using this unified telemetry and query framework, we conduct a multi-scope energy characterization of the NSF REPACSS data center. The results show stable facility efficiency, efficient in-row cooling behavior, distinct differences in idle and peak power between CPU and GPU nodes, and diverse non-CPU energy contributions in representative workloads. Together, these findings demonstrate that unified out-of-band telemetry combined with a power-centric energy model, enables practical cross-scope energy observability in production HPC systems.