SOLO: Pretraining Billion-Parameter Language Models with Shared-Output Local Learning
SOLO is the first local learning method to show such memory and throughput gains in billion-parameter language-model pretraining, and becomes a practical alternative to backpropagation for large-scale pretraining.
Bo-Jian Yin, Shu-Rong Wang, Yu-Qi Pan et al.
· 0 citations