Wafer-scale accelerators offer a new scaling point for AI infrastructure, but they also create a new compilation regime: communication cost varies sharply with location, and the space of possible placements and execution schedules is enormous. Existing GPU, distributed, and vendor compilation systems largely retain a s...
Ye-Qi Huang, Cong-Jie He, Hao-Cheng Xiao et al.· Proceedings of the ACM SIGOP...· 0 citations
This paper proposes closed-loop feature probing (CLFP), a generic and systematic framework for constructing bit-accurate arithmetic behavior models of MMA operations that explain previously observed cross-platform numerical discrepancies and accuracy issues, enable white-box numerical error analysis, and inform softwar...
Peichen Xie, Shuotao Xu, Yang Wang et al.· 3 citations
Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (...
Pei-Ran Wang, An-Qi Wang, Jia-Ying Zhao et al.· 0 citations
This paper presents FastTPS, a high performance and low-precision loss method for accelerating the token-phase in LLM inference on general AI accelerators which includes three key components: AI accelerator-enabled reloading-free KV Cache concatenation which decreases memory access overhead as well as enables full fusi...
Wenzong Yang, Danyang Zhang, Kunteng Cao et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.