Skip to content

Author

Can Xiao

We have 3 of 10 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

When Does Disaggregation Pay? Simulating Prefill--Decode--Attention--FFN Specialization for Agentic LLM Inference

Agentic inference now dominates the LLM inference landscape, requiring LLMs to actively engage in multi-turn interactions with tool-calling capabilities. This introduces a more complex workload for the underlying inference system: serving stages such as prefill and decode exhibit substantially different behaviors and d...

Przemyslaw Forys, Haoran Wu, Can Xiao et al. · 0 citations
#machine learning Preprint Sep 2026

AgentKV: Phase-Aware KV Eviction for Agentic LLMs

This work proposes AGENTKV, which maintains a small query buffer per phase and scores cached keys against their union and implements AGENTKV in a persistent multi-turn serving path that carries compressed KV state across turns and compacts retained KV pages online.

T. Liu, Jeffrey T. H. Wong, Can Xiao et al. · 1 citation
Preprint Aug 2026

OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse Prefetching

OasisKV is presented, a memory-centric LLM inference system design that alleviates HBM capacity pressure by decoupling full KV-cache storage from HBM during LLM decoding and observes that future important tokens can be predicted accurately in advance using lookahead tokens drafted by speculative decoding (SD).

Can Xiao, Sukmin Cho, J. We et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.