Skip to content

Author

Clark Gaylord

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

Monitoring High-Performance Computing Infrastructure at George Washington University Using Open-Source Solutions

Monitoring High-Performance Computing (HPC) environments is a challenging task due to the scale, diverse hardware, and specialized workloads these systems support. At the George Washington University (GW), supporting a reliable research computing environment with a small administrative team required moving away from siloed, component-specific monitoring to an integrated, unified approach. This paper describes GW’s experience deploying Zabbix [11], an open-source enterprise monitoring platform, to oversee its HPC ecosystem, including compute nodes, GPU resources, storage systems, Infiniband fabric and the Slurm workload manager. We detail our distributed architecture, custom integrations, and lessons learned in minimizing monitoring overhead while maximizing actionable visibility. Our experience demonstrates that open-source tooling can meet the rigorous telemetry demands of an academic HPC center without incurring the licensing costs of commercial alternatives.

Rubeel Muhammad Iqbal, Glen MacLachlan, Clark Gaylord · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.