Book
Open access
Jul 2026
Dynamo-MoE: Accelerating Sparse Large Model Inference with Dynamic Parallelization
Dynamo-MoE, an out-of-box MoE inference framework to bridge the performance gap by dynamic parallelization strategies, integrates a novel load balancing approach based on token sorting and on-demand expert loading to solve the workload imbalance issue in the scenario of high workload.
Jiahao Chen, Shigang Li, Rongtian Fu et al.
· IEEE International Symposium... · 0 citations