Chapel for Parallel AND Distributed GPU Computing: A Case Study with Jaccard Similarity
Abstract
This work explores the efficacy of the Chapel programming language’s GPU support for implementing irregular distributed GPU graph applications. Chapel’s partitioned global address space (PGAS) model provides a cohesive way to target distributed nodes, CPU concurrency and task parallelism, and both CPU and GPU single instruction, multiple data/thread (SIMD/SIMT) kernels. In contrast, dominant high-performance computing approaches often require interoperation of multiple programming models for inter- and intra-node operations. This work considers developer and performance impacts of using Chapel over such traditional approaches. To facilitate, we first create a novel tiled partitioning of the edge-connected Jaccard similarity graph workload in both Chapel and MPI+OpenMP+CUDA (i.e. MPI+X). We then evaluate Chapel’s first-class distributed GPU capability relative to a traditional one-sided MPI, OpenMP task, and CUDA approach. We contrast how the programming models express GPU, I/O, tasking, and remote data accesses, as well as the relative code bulk required. Finally, we evaluate performance attained by our Chapel implementation on real-world datasets, both relative to the MPI+X solution and a single-GPU pure CUDA baseline. Overall, when evaluating both Chapel and MPI+X implementations of partitioned Jaccard similarity on up to 16 A100-80GB GPUs across four nodes, Chapel achieves comparable performance to MPI+X while also delivering significantly better programmer productivity and agility.