Skip to content

Witcher: Training Large Language Models With Heterogeneous Spot GPUs in Data Centers

2026 · IEEE Transactions on Networking · Vol 34, pp. 7271-7284 · 0 citations · 46 references

Abstract

As the size of large language models (LLMs) continues to grow, training these models typically relies on data centers equipped with high-performance GPUs. However, in modern data centers, acquiring large-scale homogeneous GPU resources often results in long queueing delays, and training with on-demand instances incurs prohibitively high financial costs, limiting the efficiency of large-scale model training. In this paper, we present Witcher, a system for efficiently training large language models with heterogeneous spot GPUs in data centers. It utilizes more easily obtained heterogeneous GPUs and low-cost, preemptible spot instances, while still achieving stable and efficient training performance. Witcher introduces a heterogeneous training framework that supports asymmetric 3D parallelism, allowing flexible and efficient parallel configurations across heterogeneous GPUs. To address frequent preemption and allocation events of spot instances, Witcher employs a fast recovery mechanism that reconstructs interrupted pipelines by copying from intact model replicas. Witcher further develops an efficient automatic parallelism algorithm that includes two dynamic programming approaches, pipeline construction and micro-batch distribution, which minimize the training iteration time. Evaluation on multiple LLMs using two real-world traces of heterogeneous spot GPUs shows that Witcher achieves up to $14.8\times $ higher training throughput compared to state-of-the-art frameworks.

View source

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.