This work offers data-driven validation of SPEC's methodology, showing that the fidelity gap is not a flaw but a quantifiable consequence of enforcing portability, determinism, and CPU-centric measurement.
Abstract
Standardized benchmarks are often criticized for not being"real workloads,"but this critique is rarely backed by data. This paper provides the first systematic, quantitative analysis of the"fidelity gap"between the SPEC CPU2026 suite and its original, upstream open-source counterparts. We compile both the SPEC benchmarks and their upstream applications and execute them with official input workloads under two scenarios: a single-copy latency run and a 192-copy throughput run. Our findings show that most benchmarks exhibit high fidelity in single-copy runs, while a few outliers reveal the impact of SPEC's adaptation process. The multi-copy results further highlight the necessity of this adaptation: several benchmarks become significantly more efficient than their upstream versions under heavy load, underscoring the importance of I/O reduction. This work offers data-driven validation of SPEC's methodology, showing that the fidelity gap is not a flaw but a quantifiable consequence of enforcing portability, determinism, and CPU-centric measurement.
SPEC CPU 2026 is the first major update to the industry-standard CPU benchmark suite since 2017. This paper presents the first microarchitecture based performance characterization of the new suite, conducted on AMD EPYC"Zen 5", also the first SPEC CPU characterization study on this microarchitecture. Using a multi-lens...
Kunal Kashyap, Rajiv Ramanathan, S. Bhattacharya· 0 citations
PerfReasoning is introduced, a benchmark that evaluates LLMs both as direct performance reasoners and as generators of analytical performance-model code, exposing the gap between plausible architectural reasoning and reliable performance-model construction.
Da Zhao, K. Sankaralingam, Christos Kozyrakis et al.· 0 citations
CC-Bench is presented, a lightweight, extensible, and application-oriented benchmark suite for evaluating communication compression under realistic execution conditions, and uses declarative application-environment modeling to decouple profiling logic from communication libraries, datasets, and fidelity metrics, enabli...
Hao Fan, Wei Wang, Xing-Chen Liu et al.· 0 citations
Parallel performance depends not only on programming language and runtime design, but also on how the dominant execution bottleneck changes as parallelism increases. We present a controlled cross-language study of Rust, Julia, Haskell, and Python using Merge Sort, Closest Pair of Points, and Numerical Sum in a multicor...
Muhammad Hassam Aslam Khan, Daniel Stapleton, Medha Kulkarni et al.· Software· 0 citations
Data centers need tooling that validates an entire installation rather than individual nodes, at acceptance and at regular intervals thereafter. This requires dispatching identical benchmarks to every node in a single submission, and therefore cluster-aware scheduling. This paper presents ClusterBench, a framework for...
A. Ujeniya, Jan Eitzinger, Thomas Gruber et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.