Skip to content

Author

Xin Wang

We have 1 of 32 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer

SmartGen is designed, a KV cache transfer engine that allows seamless disaggregated LLM inference with three data transfer paths that reduces time-to-second-token by up to 4.3x compared with the typical full KV cache transfer approach while offering comparable subsequent decoding performance and accuracy.

Xuchuan Luo, Jiacheng Shen, Xin Wang et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.