Scaling the Extended OpenDwarfs: Evaluating Cross-Vendor CUDA Performance Portability with SCALE
CUDA is the dominant GPU programming model in HPC and industrial accelerator software, and a large body of production code is written directly in it. Deploying that code on non-NVIDIA accelerators has traditionally required source translation, backend-specific rewrites, or a full rewrite in a new programming model. Thi...