Architectural diversity has turned accelerator performance portability into a compiler/runtime problem: portable source code is useful only if the surrounding ecosystem can also coordinate devices, backends, and data movement. This paper evaluates three SYCL ecosystems—Intel oneAPI DPC++, AdaptiveCpp, and UniSYCL—acros...
Nabayan Chaudhury, Norihisa Fujita, Beau Johnston et al.· Proceedings of the Internati...· 0 citations
This work explores the efficacy of the Chapel programming language’s GPU support for implementing irregular distributed GPU graph applications. Chapel’s partitioned global address space (PGAS) model provides a cohesive way to target distributed nodes, CPU concurrency and task parallelism, and both CPU and GPU single in...
Paul Sathre, Wu-Chun Feng· Proceedings of the Internati...· 0 citations
A fine-grained GPU parallelization approach that adapts to workload characteristics by assigning set intersection computations within a graph to the most suitable parallelization strategies, and predicts the optimal parallelization approach between this fine-grained and other state-of-the-art GPU kernels with 88% accur...
Atharva Gondhalekar, Wu-Chun Feng· Proceedings of the Internati...· 0 citations
Overall, when evaluating both Chapel and MPI+X implementations of partitioned Jaccard similarity on up to 16 A100-80GB GPUs across four nodes, Chapel achieves comparable performance to MPI+X while also delivering significantly better programmer productivity and agility.
Paul Sathre, Wu-Chun Feng· Proceedings of the Internati...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.