Skip to content

Author

Simeng Sun

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Aug 2026

Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration

Communication-efficient MoE models (CE-MoE), in which a heterogeneous layer pattern that decouples token-mixing and channel-mixing depth is adopted, consistently reduce training cost while matching validation loss and downstream benchmarks with full-MoE baselines.

Simeng Sun, R. Waleffe · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.