Skip to content
Book Open access

From Per-Job Data to an Aggregated Workload Insight: A Toolkit for Profiling HTCondor Workload

Jul 2026 · Practice and Experience in Advanced Research Computing · pp. 1-3 · 0 citations · 2 references
Computer Science

TL;DR

A Python toolkit that aggregates raw job data into four cluster-level views, designed to answer a specific question a researcher or facilitator would ask when triaging a workload, transforming thousands of individual job records into a concise, interpretable summary.

Abstract

HTCondor users for high-throughput computing often struggle to quickly understand how their computational workloads are performing. Current interfaces expose large volumes of raw job data, making it difficult to diagnose common issues such as jobs stuck on hold, poor resource utilization, or unexpected failures. These issues further snowball when dealing with large clusters. We present a Python toolkit, developed at the Center for High Throughput Computing (CHTC) at the University of Wisconsin–Madison, that bridges this gap. To be included as a part of the HTCondor suite, given a single cluster ID the toolkit aggregates raw job data into four cluster-level views: a status dashboard showing the distribution of job states, a runtime histogram revealing duration variance and flagging anomalously short runs, a hold classifier that groups held jobs by reason code with plain-language explanations, and a resource utilization report comparing requested versus actual CPU, memory, and disk usage. Each view is designed to answer a specific question a researcher or facilitator would ask when triaging a workload, transforming thousands of individual job records into a concise, interpretable summary.

Read PDF

Similar papers

Open access 2023

Performance Bottlenecks Necks in Data Heavy Python Applications

Performance improvements in data-intensive Python applications have become more critical due to the increasing computational needs of modern analytics, machine learning, and large-scale data processing systems. Although the Python environment is enormous as well as flexible in development, frequent delay in execution,...

Madhurima Kommuru, Appala Nooka Kumar Doodala · 0 citations
Preprint Sep 2026

VERA: Reinforcement Learning for Dynamic Memory Scaling of HPC Workloads in Kubernetes

Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run HPC jobs. In this work, we present a reinforcement learning (RL) recommender VERA that formulates vertical memory scali...

A. Pramono, Jie Ren, Ivy Bo Peng · 0 citations
Aug 2026

TexBench : A Unified Benchmarking Suite for Shifting Workloads

Key-value stores are widely adopted as the storage engine for modern applications, as they offer high throughput for writes, support for heterogeneous workloads, and are easily tunable. Given the large number of key-value stores available and their performance variability with workload shifts, finding the suitable data...

Abhishek Chanda, Shubham Kaushik, A. Lavrov et al. · 0 citations
Preprint Aug 2026

ClusterBench: A Framework for Cluster-Wide Continuous Benchmarking and Regression Testing

Data centers need tooling that validates an entire installation rather than individual nodes, at acceptance and at regular intervals thereafter. This requires dispatching identical benchmarks to every node in a single submission, and therefore cluster-aware scheduling. This paper presents ClusterBench, a framework for...

A. Ujeniya, Jan Eitzinger, Thomas Gruber et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.