Skip to content
Book Open access

Monitoring High-Performance Computing Infrastructure at George Washington University Using Open-Source Solutions

Jul 2026 · Practice and Experience in Advanced Research Computing · 0 citations · 7 references
Computer Science

Abstract

Monitoring High-Performance Computing (HPC) environments is a challenging task due to the scale, diverse hardware, and specialized workloads these systems support. At the George Washington University (GW), supporting a reliable research computing environment with a small administrative team required moving away from siloed, component-specific monitoring to an integrated, unified approach. This paper describes GW’s experience deploying Zabbix [11], an open-source enterprise monitoring platform, to oversee its HPC ecosystem, including compute nodes, GPU resources, storage systems, Infiniband fabric and the Slurm workload manager. We detail our distributed architecture, custom integrations, and lessons learned in minimizing monitoring overhead while maximizing actionable visibility. Our experience demonstrates that open-source tooling can meet the rigorous telemetry demands of an academic HPC center without incurring the licensing costs of commercial alternatives.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.