As the size of large language models (LLMs) continues to grow, training these models typically relies on data centers equipped with high-performance GPUs. However, in modern data centers, acquiring large-scale homogeneous GPU resources often results in long queueing delays, and training with on-demand instances incurs...
Chen-Hao Wang, Jun Wu, Ying-Hao Yu· IEEE Transactions on Network...· 0 citations
Weave is presented, to the authors' knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling - deciding per layer and per GPU by routing results at runtime, and achieves a 2.89x geometric-mean MoE-layer speedup and a 1.33x geometric-mean end-to-end speedup over five state-of-the-art basel...
Ziyu Huang, Yangjie Zhou, Chen-Hao Zhu et al.· 0 citations
Atrex-Bench is presented, a benchmark whose 30 operators and 440 shapes are sampled directly from full-cluster production inference traces of compute-limited, memory-rich GPUs, and a profile-driven kernel-optimization agent that combines iterative measure-revise search, optimization dropout for escaping stalled search...