Preprint
Aug 2026
LLaDA MoE v2: Scaling Mixture-of-Experts Diffusion Language Models
This work systematically characterize how optimization hyperparameters, compute allocation, and architecture scale for MoE dLLMs, identifying quantitative differences from scaling trends previously reported for AR models.
Fengqi Zhu, Shaoxuan Xu, Jingyang Ou et al.
· 0 citations