TurboBus is presented, which pools PCIe bandwidth across co-located jobs via emerging scale-up fabrics and reduces first-token latency by up to 40% for on-demand model loading, achieves up to 1.6x throughput for KV-cache-offloaded inference, and accelerates training by up to 7%, while imposing less than 1% overhead on co-located workloads.
Xinyu Yang, Kaiqiang Xu, Kai Chen· Conference on Applications,...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.