Skip to content

Author

Kaiqiang Xu

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

TurboBus: Pooling PCIe Bandwidth for LLM Workloads via Scale-Up Fabrics

TurboBus is presented, which pools PCIe bandwidth across co-located jobs via emerging scale-up fabrics and reduces first-token latency by up to 40% for on-demand model loading, achieves up to 1.6x throughput for KV-cache-offloaded inference, and accelerates training by up to 7%, while imposing less than 1% overhead on co-located workloads.

Xinyu Yang, Kaiqiang Xu, Kai Chen · 0 citations
#edge computing Preprint Aug 2026

WiCi: Wireless GPU Computing Infrastructure

The proposed Wireless GPU Computing Infrastructure (WiCi) can reduce time to first token by up to 90%, improve the token rate by approximately 39x compared to local inference on mobile devices for the same model, and support much larger models.

Yibin Shen, Wei Li, Kaiqiang Xu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.