Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack
This work integrates SmoothQuant into TorchAO and optimize the resulting inference path for Intel Xeon CPUs through graph-level fusion in TorchInductor and efficient INT8 GEMM kernel selection across oneDNN-, AVX512_VNNI-, and AMX-based implementations.
Weiwen Xia, Yuxin Cui, E. Cao
· 0 citations