Jul 2026· Practice and Experience in Advanced Research Computing· pp. 1-3· 0 citations· 10 references
Computer Science
TL;DR
The results reveal that x86 instances outperform ARM-based instances in raw HPL throughput, that latest CPU generations offer substantially better price-performance, and that spot pricing offers additional opportunities for reducing execution costs.
Abstract
This poster presents performance results and cost analysis of the High Performance Computing Challenge (HPCC) benchmark suite across diverse commercial cloud compute instances. We evaluate HPL (High Performance Linpack) performance and spot instance pricing on AMD EPYC, ARM, and Intel Xeon architectures under Amazon Web Services, Google Cloud, and Microsoft Azure. Our results reveal that x86 instances outperform ARM-based instances in raw HPL throughput, that latest CPU generations offer substantially better price-performance, and that spot pricing offers additional opportunities for reducing execution costs. These findings offer practical recommendations for researchers considering commercial cloud resources for small-scale computationally intensive workloads.
This paper conducts a comprehensive evaluation of general-purpose compute instances offered by leading cloud service providers, focusing on the interplay between processor architecture, cost, and performance metrics. Utilizing standardized configurations and benchmarking methodologies across Intel, AMD, and ARM architectures, this research delineates the tradeoffs inherent in cloud infrastructure selection. Results demonstrate ARM-based instances provide superior cost-efficiency for scale-out and cloud-native workloads, while Intel architectures maintain dominance in legacy-sensitive, performance-critical environments. Insights derived aim to empower computing professionals in optimizing cloud resource allocation, maximizing computational throughput per expense, and guiding strategic deployments within evolving distributed systems paradigms.
Rahul Sadhwani· International Conference on...· 0 citations
Results show that selecting instances based on the second PI achieves at least 97% of the best achievable execution time in most cases, while highlighting cases where additional PIs improve selection accuracy.
J. R. Brunetta, J. F. Borin, E. Borin· Concurrency and Computation· 0 citations
The findings indicate that integrating a provider-independent resource selection strategy with a structured cloud-native deployment approach can enhance the efficiency of scientific workload execution in HPC cloud environments.
This work investigates the feasibility of reproducing benchmarks originally run on datacenter GPUs such as the NVIDIA A100 and RTX 8000 using consumer-grade graphics cards, focusing on the NVIDIA GeForce RTX 3050 and GTX 1060 with CUDA Graphs support. Seven NAS Parallel Benchmarks (BT, LU, SP, EP, IS, MG, and CG) are evaluated across problem classes W, A, B, and C. Results show that the RTX 3050 delivers stable performance, typically 6×–12× slower than the A100, while VRAM limitations severely constrain the GTX 1060 for larger-scale problems. Although enterprise GPUs remain essential for massive, memory-bound workloads, modern consumer hardware combined with CUDA Graphs enables economical reproduction of moderate scientific experiments, supporting the democratization of high-performance computing research.
Leandro L. Retzlaff, Calebe C. Pereira, Helena P. Veltri et al.· Anais do LIII Seminário Inte...· 0 citations
The new EMC+ proposal is an OS‐driven elasticity manager for container‐based environments that continuously estimates idle core cycles left by regular (inelastic) applications, and reallocates idle cores to elastic ones, even during short time intervals, and has minimal impact on the performance and QoS of colocated inelastic applications.
J. C. Saez, Carlos Bilbao, Manuel Prieto-Matías· Concurrency and Computation· 0 citations
Distributed HPC and LLM workloads increasingly require efficient communication for scalability, yet growing data movement has become a major performance bottleneck. Communication compression can reduce this overhead and complement execution-level optimizations, but its benefits remain difficult to assess because existing benchmarks lack support for diverse backends, realistic datasets, application-specific accuracy metrics, and overlap-induced resource contention. We present CC-Bench, a lightweight, extensible, and application-oriented benchmark suite for evaluating communication compression under realistic execution conditions. CC-Bench uses declarative application-environment modeling to decouple profiling logic from communication libraries, datasets, and fidelity metrics, enabling portable cross-library evaluation. It further combines function-level interception and hardware counter monitoring to characterize per-phase latency, hardware utilization, numerical fidelity, and computation interference. With representative datasets from HPC and LLM workloads, CC-Bench evaluates three compression-enabled communication libraries on CPU and GPU clusters, revealing accuracy-performance trade-offs and bottlenecks to guide practical deployment and optimization.
Unknown authors· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.