Preprint
Aug 2026
CommBench: Can LLMs Write Correct and Efficient GPU Communication Code?
Evaluating leading frontier and open-source code generation models on both intra-node NVLink and inter-node RDMA platforms reveals that even the strongest model, GPT-5.5, correctly implements and achieves competitive performance on only 30.7\% of the benchmark tasks.
Shuang Ma, Yu-Yi Li, Yihan Zhang et al.
· 0 citations