Skip to content

Author

Wenfeng Wang

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#large language models Open access Oct 2026

PCR: A Prefetch-Enhanced Cache Reuse System for Low-Latency RAG Serving

Retrieval-Augmented Generation (RAG) significantly improves Large Language Models (LLMs) but introduces massive input sequences that severely bottleneck the prefill stage. While KV-cache reuse reduces redundant computation for shared document prefixes, the reusable KV working set in RAG serving can exceed GPU memory ca...

Wenfeng Wang, Xiaofeng Hou, Peng Tang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.