Low-bit and Sparsified Gradient Communication for Accelerating Distributed Deep Learning with Convergence Guarantees
Abstract
Communication poses a dominant bottleneck in distributed data parallel training with synchronous stochastic gradient descent, whereas the traditional AllReduce collective used for gradient synchronization limits the efficient utilization of communication compression strategies. In this paper, we propose a low-bit and sparsified gradient communication algorithm with convergence guarantees, called QTopKA2A, which enables effectively overlapping communications with both feed-forward and backpropagation computations through a decomposed communication framework. Specifically, top-k sparsified gradients are communicated via AlltoAll in the backward phase, while low-bit quantized gradients are exchanged via AllGather in the forward phase. Residual-based error feedback mechanisms are employed in both sparsification and quantization stages to compensate for compression errors and guarantee convergence with significantly higher compression ratios. In addition, we leverage an effective tensor fusion to reduce communication startup overhead. Experimental results demonstrate that QTopKA2A preserves dense-level accuracy while achieving up to 4.74 × end-to-end training speedup over existing dense and compressed baselines on a 32-GPU cluster.