Skip to content
Book Open access

Scaling the Extended OpenDwarfs: Evaluating Cross-Vendor CUDA Performance Portability with SCALE

Sep 2026 · Workshop Proceedings of the 55th International Conference on Parallel Processing · 0 citations · 6 references

Abstract

CUDA is the dominant GPU programming model in HPC and industrial accelerator software, and a large body of production code is written directly in it. Deploying that code on non-NVIDIA accelerators has traditionally required source translation, backend-specific rewrites, or a full rewrite in a new programming model. This paper evaluates SCALE, a CUDA compiler that preserves the CUDA programming model while targeting GPUs from multiple vendors through an nvcc-compatible interface, rather than treating portability as a source-to-source translation problem. We evaluate SCALE using Extended OpenDwarfs, a modernised version of the OpenDwarfs benchmark suite with new CUDA and native HIP implementations, an updated Spectral Methods dwarf built around a new Continuous Wavelet Transform benchmark, new CFD problem-size generation tooling, and consistent timing regions across the suite. Across 16 benchmarks and ten systems spanning five NVIDIA and seven AMD GPU models, SCALE passes checksum-validated correctness against native NVCC and HIP builds in every case but one. Performance tracks native toolchains closely: SCALE is 5–6% faster than HIP at small problem sizes and close to parity with HIP at large sizes, carries a modest (under 5%) overhead against NVCC at large sizes, and shows one isolated AMD-side regression rather than a systemic slowdown on either platform. Existing CUDA applications can target NVIDIA and AMD hardware through SCALE without being rewritten in a second programming model.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.