Jul 2026· International Conference on Supercomputing· pp. 328-341· 0 citations· 56 references
Computer Science
TL;DR
By enabling scalable GPU memory expansion and introducing an alternative inter-node data path, COMETS reduces reliance on NIC-based inter-node networking and supports efficient training across diverse cluster environments and proposes a heterogeneous-aware training strategy to identify optimal training configurations.
Abstract
Rapid growth of large language models (LLMs) presents major challenges for distributed multi-node training, due to insufficient inter-node bandwidth and lack of scalability in memory expansion. To address both challenges, we propose COMETS, a cost-effective multi-node efficient training system with memory pooling and sharing. By enabling scalable GPU memory expansion and introducing an alternative inter-node data path, COMETS reduces reliance on NIC-based inter-node networking and supports efficient training across diverse cluster environments. We also propose a heterogeneous-aware training strategy to identify optimal training configurations. Experiments show that COMETS improves training throughput by up to 2.11 × over ZeRO-Infinity across eight hardware setups, and boosts performance-per-dollar by 1.98 × and 1.37 × on homogeneous and heterogeneous clusters, respectively. Code is available at: https://github.com/sharc-lab/COMETS.
Mixture-of-Experts (MoE) has become the de facto architecture for scaling large language models, offering expanded capacity with manageable compute. Pipeline parallelism (PP) is indispensable for distributed MoE training, but state-of-the-art PP schemes face three major limitations: large pipeline bubbles, high per-sta...
Jun Li, Zhi Ma, Shigang Li· Proceedings of the Internati...· 0 citations
Experimental evaluations demonstrate the effectiveness of the framework ASTRA, which achieves lower time-to-accuracy than a resource-heterogeneity-aware baseline and several compression-based frameworks, while preserving convergence quality and robustness across heterogeneous hybrid cloud environments.
Tuan Anh Vuong, Thanh Loi Hoang, Huan Le et al.· IEEE Access· 0 citations
HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation...
Hao-Chen Huang, Shu-Zhang Zhong, Sheng-Xuan Qiu et al.· IEEE Transactions on Compute...· 0 citations
This work proposes a BP-free algorithm, called ZeroLock, that decouples the model updates into independent chunk updates by local objective construction and provides the first theoretical framework for such local objective construction-based approaches under general model chunk division by mapping local objectives to t...
Wen-Tao Dai, Xuan-Ran Li, Yu-Xiang Zhang et al.· 0 citations
Preliminary results from a geo-distributed LLM training prototype that treats networking constraints as first-order design concerns are presented, motivating adaptive networking support for synchronization, compression, placement, telemetry, and recovery in geo-distributed LLM training.
Zi-Yue Luo, Jiaxuan Cai, Cedric Le Denmat et al.· Conference on Applications,...· 0 citations
DynamoServe is presented, a multi-tenant LLM serving framework that addresses challenges through three key innovations: leveraging stranded GPU memory to offload model weights and KV caches, mitigating resource fragmentation in multi-workload environments, and improving memory locality through coordinated data placemen...
Diman Zad Tootaghaj, Khaled Diab, Bob Lantz et al.· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.