Skip to content

Author

Fan Yang

We have 4 of 8 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Sep 2026

Wavel: A Fast and Efficient Compilation System for Wafer-Scale Accelerators

Wafer-scale accelerators offer a new scaling point for AI infrastructure, but they also create a new compilation regime: communication cost varies sharply with location, and the space of possible placements and execution schedules is enormous. Existing GPU, distributed, and vendor compilation systems largely retain a s...

Ye-Qi Huang, Cong-Jie He, Hao-Cheng Xiao et al. · 0 citations
#machine learning Preprint Nov 2025

Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy

This paper proposes closed-loop feature probing (CLFP), a generic and systematic framework for constructing bit-accurate arithmetic behavior models of MMA operations that explain previously observed cross-platform numerical discrepancies and accuracy issues, enable white-box numerical error analysis, and inform softwar...

Peichen Xie, Shuotao Xu, Yang Wang et al. · 3 citations
#machine learning Preprint Aug 2026

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

Low-bit quantization of linear layers is often dominated by a small number of outlier channels. Existing smoothing, rotation, and residual-based methods can mitigate this issue, but may shift the quantization bottleneck to weights or introduce costly online operations. To address these limitations, we propose PRQuant (...

Pei-Ran Wang, An-Qi Wang, Jia-Ying Zhao et al. · 0 citations
Jul 2026

FastTPS: An Optimized Method for LLM Token Phase for AI accelerators

This paper presents FastTPS, a high performance and low-precision loss method for accelerating the token-phase in LLM inference on general AI accelerators which includes three key components: AI accelerator-enabled reloading-free KV Cache concatenation which decreases memory access overhead as well as enables full fusi...

Wenzong Yang, Danyang Zhang, Kunteng Cao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.