Monitoring High-Performance Computing Infrastructure at George Washington University Using Open-Source Solutions
Abstract
Monitoring High-Performance Computing (HPC) environments is a challenging task due to the scale, diverse hardware, and specialized workloads these systems support. At the George Washington University (GW), supporting a reliable research computing environment with a small administrative team required moving away from siloed, component-specific monitoring to an integrated, unified approach. This paper describes GW’s experience deploying Zabbix [11], an open-source enterprise monitoring platform, to oversee its HPC ecosystem, including compute nodes, GPU resources, storage systems, Infiniband fabric and the Slurm workload manager. We detail our distributed architecture, custom integrations, and lessons learned in minimizing monitoring overhead while maximizing actionable visibility. Our experience demonstrates that open-source tooling can meet the rigorous telemetry demands of an academic HPC center without incurring the licensing costs of commercial alternatives.