Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration
Communication-efficient MoE models (CE-MoE), in which a heterogeneous layer pattern that decouples token-mixing and channel-mixing depth is adopted, consistently reduce training cost while matching validation loss and downstream benchmarks with full-MoE baselines.
Simeng Sun, R. Waleffe
· 0 citations