Large language model training involves massive computation on GPU streaming multiprocessors (SMs), the primary compute units of GPUs. Since SMs host specialized accelerators such as Tensor Cores, their efficient utilization is critical to training efficiency. Unfortunately, existing collective communication systems com...
Yao Fei, Gong-Ming Zhao, Hong-Li Xu et al.· 0 citations
VarioPath, an efficient AlltoAllv scheduling framework for PCIe GPU systems that combines an offline topology-aware analyzer with an online demand-aware scheduler, and shows average AlltoAllv speedups of 5.88x over FAST and 1.72x over DeepEP.
Yao Fei, Jin Fang, Si-Ze Zheng et al.· 0 citations
A GPU-native, topology-aware load-balancing system for large-scale MoE training that converts the current routing result into hot-expert replication and token-rerouting decisions and executes the resulting plan without data-dependent host synchronization, reducing critical-path overhead.
Jia-Cheng Zhu, Xie Zhao, Gong-Ming Zhao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.