Skip to content
Book Open access

COMETS: Cost-effective Multi-node Efficient Training System with Memory Pooling and Sharing

Jul 2026 · International Conference on Supercomputing · pp. 328-341 · 0 citations · 56 references
Computer Science

TL;DR

By enabling scalable GPU memory expansion and introducing an alternative inter-node data path, COMETS reduces reliance on NIC-based inter-node networking and supports efficient training across diverse cluster environments and proposes a heterogeneous-aware training strategy to identify optimal training configurations.

Abstract

Rapid growth of large language models (LLMs) presents major challenges for distributed multi-node training, due to insufficient inter-node bandwidth and lack of scalability in memory expansion. To address both challenges, we propose COMETS, a cost-effective multi-node efficient training system with memory pooling and sharing. By enabling scalable GPU memory expansion and introducing an alternative inter-node data path, COMETS reduces reliance on NIC-based inter-node networking and supports efficient training across diverse cluster environments. We also propose a heterogeneous-aware training strategy to identify optimal training configurations. Experiments show that COMETS improves training throughput by up to 2.11 × over ZeRO-Infinity across eight hardware setups, and boosts performance-per-dollar by 1.98 × and 1.37 × on homogeneous and heterogeneous clusters, respectively. Code is available at: https://github.com/sharc-lab/COMETS.

Read PDF

Similar papers

#large language models Book Open access Sep 2026

OmniPipe: Efficient, Flexible and Scalable Pipeline Parallelism for Large Model Training

Mixture-of-Experts (MoE) has become the de facto architecture for scaling large language models, offering expanded capacity with manageable compute. Pipeline parallelism (PP) is indispensable for distributed MoE training, but state-of-the-art PP schemes face three major limitations: large pipeline bubbles, high per-sta...

Jun Li, Zhi Ma, Shigang Li · 0 citations
Open access 2026

Synchronous Distributed Training With Runtime-Adaptive Mechanisms in Hybrid Cloud Environments

Experimental evaluations demonstrate the effectiveness of the framework ASTRA, which achieves lower time-to-accuracy than a resource-heterogeneity-aware baseline and several compression-based frameworks, while preserving convergence quality and robustness across heterogeneous hybrid cloud environments.

Tuan Anh Vuong, Thanh Loi Hoang, Huan Le et al. · 0 citations
Open access Sep 2026

HDA-MoE: Hybrid Parallelism and Dynamic, Adaptive Scheduling for Mixture-of-Experts with 3D Near-Memory Processing

HDA-MoE is presented, a framework that optimizes MoE execution on NMP architectures through hybrid parallel deployment and runtime scheduling and integrates an offline hybrid parallel mapping algorithm with an online dynamic and adaptive scheduling mechanism to reduce communication overhead while improving computation...

Hao-Chen Huang, Shu-Zhang Zhong, Sheng-Xuan Qiu et al. · 0 citations
Preprint Aug 2026

ZeroLock: Concurrent Memory-Efficient LLM Training via Modular Update Decoupling

This work proposes a BP-free algorithm, called ZeroLock, that decouples the model updates into independent chunk updates by local objective construction and provides the first theoretical framework for such local objective construction-based approaches under general model chunk division by mapping local objectives to t...

Wen-Tao Dai, Xuan-Ran Li, Yu-Xiang Zhang et al. · 0 citations
Book Open access Aug 2026

Toward WAN-Aware LLM Training Across Heterogeneous, Geo-Distributed Sites

Preliminary results from a geo-distributed LLM training prototype that treats networking constraints as first-order design concerns are presented, motivating adaptive networking support for synchronization, compression, placement, telemetry, and recovery in geo-distributed LLM training.

Zi-Yue Luo, Jiaxuan Cai, Cedric Le Denmat et al. · 0 citations
Book Open access Aug 2026

DynamoServe: A Distributed Tiered Memory System for Multi-tenant LLM Serving

DynamoServe is presented, a multi-tenant LLM serving framework that addresses challenges through three key innovations: leveraging stranded GPU memory to offload model weights and KV caches, mitigating resource fragmentation in multi-workload environments, and improving memory locality through coordinated data placemen...

Diman Zad Tootaghaj, Khaled Diab, Bob Lantz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.