Skip to content

Author

Kunming Shao

We have 3 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

LLM agents resend their whole conversation on every turn, and most of it was already processed on the previous turn. Serving systems avoid recomputing it by caching its key-value (KV) state and, when GPU memory runs out, by offloading that state to host memory. For agents, offloading gives inconsistent results: on the...

Kun-Ming Shao, Jie-Run Chen, Jiang-Nan Yu et al. · 0 citations
#machine learning Preprint Sep 2026

PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention

PQ-HSA (hybrid sparse-approximate attention) attends the selected tokens with their original keys and values, and the unselected tokens, the background, enter the same softmax through those scores, summed per inverted list and multiplied by the list's mean value.

Kun-Ming Shao, Jie-Run Chen, Yan-Li Wang et al. · 0 citations
Preprint Aug 2026

MoE Expert Execution in Disaggregated LLM Serving with a High-Bandwidth ReRAM Near-Memory Architecture

A ReRAM near-memory architecture that keeps expert weights resident behind high-bandwidth local reads and recovers occupancy with bounded core-local multicast pooling, coactivation-aware placement, and load-aware fetch, and sizes each communication level from induced demand is presented.

Kun-Ming Shao, Ming Zeng, Xin Yuan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.