Jul 2026· Practice and Experience in Advanced Research Computing· pp. 1-3· 0 citations· 2 references
Computer Science
TL;DR
A Python toolkit that aggregates raw job data into four cluster-level views, designed to answer a specific question a researcher or facilitator would ask when triaging a workload, transforming thousands of individual job records into a concise, interpretable summary.
Abstract
HTCondor users for high-throughput computing often struggle to quickly understand how their computational workloads are performing. Current interfaces expose large volumes of raw job data, making it difficult to diagnose common issues such as jobs stuck on hold, poor resource utilization, or unexpected failures. These issues further snowball when dealing with large clusters. We present a Python toolkit, developed at the Center for High Throughput Computing (CHTC) at the University of Wisconsin–Madison, that bridges this gap. To be included as a part of the HTCondor suite, given a single cluster ID the toolkit aggregates raw job data into four cluster-level views: a status dashboard showing the distribution of job states, a runtime histogram revealing duration variance and flagging anomalously short runs, a hold classifier that groups held jobs by reason code with plain-language explanations, and a resource utilization report comparing requested versus actual CPU, memory, and disk usage. Each view is designed to answer a specific question a researcher or facilitator would ask when triaging a workload, transforming thousands of individual job records into a concise, interpretable summary.
Performance improvements in data-intensive Python applications have become more critical due to the increasing computational needs of modern analytics, machine learning, and large-scale data processing systems. Although the Python environment is enormous as well as flexible in development, frequent delay in execution,...
Madhurima Kommuru, Appala Nooka Kumar Doodala· International Journal of App...· 0 citations
Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run HPC jobs. In this work, we present a reinforcement learning (RL) recommender VERA that formulates vertical memory scali...
Key-value stores are widely adopted as the storage engine for modern applications, as they offer high throughput for writes, support for heterogeneous workloads, and are easily tunable. Given the large number of key-value stores available and their performance variability with workload shifts, finding the suitable data...
Abhishek Chanda, Shubham Kaushik, A. Lavrov et al.· Proceedings of the VLDB Endo...· 0 citations
Data centers need tooling that validates an entire installation rather than individual nodes, at acceptance and at regular intervals thereafter. This requires dispatching identical benchmarks to every node in a single submission, and therefore cluster-aware scheduling. This paper presents ClusterBench, a framework for...
A. Ujeniya, Jan Eitzinger, Thomas Gruber et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.