A SYCLic Investigation of SYCL Ecosystems: UniSYCL for Heterogeneous Scaling
Abstract
Architectural diversity has turned accelerator performance portability into a compiler/runtime problem: portable source code is useful only if the surrounding ecosystem can also coordinate devices, backends, and data movement. This paper evaluates three SYCL ecosystems—Intel oneAPI DPC++, AdaptiveCpp, and UniSYCL—across single-device execution, homogeneous multi-GPU scaling, and mixed CUDA+HIP execution. DPC++ and AdaptiveCpp represent mature conventional SYCL stacks that provide portable compilation but leave heterogeneous orchestration largely explicit in application code. UniSYCL targets a different design point: native integration with an intelligent runtime system (IRIS), which allows SYCL applications to expose task structure while delegating placement and mixed-backend coordination to a runtime. We conduct the study on five matched applications: XSBench, K-means, Hotspot3D, N-body, and continuous wavelet transform (CWT). On a single Tesla V100, UniSYCL is competitive with DPC++, AdaptiveCpp, and CUDA, establishing a credible single-device baseline. Under weak scaling, UniSYCL+IRIS reaches 5.29 × and 4.05 × wall-clock throughput speedup on CWT and N-body on six V100s, respectively, (compared to a single V100) and leads the manual paths on XSBench, while Hotspot3D and K-means show limited weak scaling. On a representative mixed RTX 3090 + RX 6900 XT node, UniSYCL+IRIS executes one portable source over CUDA and HIP devices with a small heterogeneous wrapper, achieving a 1.11 × geometric-mean wall speedup over the per-application best manual mixed-backend path across five applications. Profiling confirms copy-time reductions of 1,267 × for CWT and 8.44 × for Hotspot3D for 6 V100s, and the mixed NVIDIA+AMD case achieves a 1.11 × geometric-mean kernel speedup relative to the selected manual paths.