Skip to content

Author

Jiaqi Zheng

We have 3 of 107 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Jul 2026

UBEP: Re-architecting Expert Parallelism Communication Library for Production Superpods

The deployment of Mixture-of-Experts (MoE) models on production high-bandwidth superpods, such as NVIDIA's NVL72/576 and Huawei's CloudMatrix384, introduces critical challenges beyond raw interconnect bandwidth. While these systems provide unified global address spaces and high-bandwidth fabrics, their full potential for sparse MoE communication is hindered by three fundamental bottlenecks: (1) Strict execution serialization imposed by coarse-grained Bulk Synchronous Parallel (BSP) orchestration of interdependent communication phases; (2) Prohibitive synchronization overhead that fails to scale alongside high interconnect bandwidth; and (3) Severe load imbalance resulting from distance-agnostic scheduling of irregular token traffic. To eliminate these bottlenecks, we introduce UBEP (Unified-Bus Expert Parallelism), a production-ready communication library that rethinks MoE's All-to-All primitives for modern superpod architectures. Through large-scale experiments, UBEP reduces All-to-All latency by up to 52.4% and MoE inference Time Per Output Token (TPOT) by up to 11.1%.

Yipeng Liu, Chang Liu, Sitan Shen et al. · 0 citations
Book Open access Aug 2026

CubeTrace: Microscopic Network Tracing for Heterogeneous Cloud Gateways

CubeTrace is presented, a unified, function-level flow tracing system that enables microscopic tracing inside heterogeneous cloud gateways and introduces minimal overhead, consuming less than 1% of memory resources and adding less than 1% to forwarding latency.

Yunming Xiao, Yinchao Yang, Jiaqi Zheng et al. · 1 citation
Preprint Aug 2026

CoRun: Padding is Simple and Efficient for Deterministic LLM Inference

CoRun is presented, a scheduling-based system that achieves deterministic inference without requiring batch invariance, and employs isolated prefill and fixed-shape batched decode to handle the two stages of LLM inference, respectively, leveraging CUDA graphs for efficient execution and simplified implementation.

Shiju Zhao, Jiacheng Yang, Qihang Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.