Preprint
Aug 2026
Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RAN
This work presents Pallas, a \textit{proactive} KV-cache migration framework that prepares the inference state at the predicted target before handover, in parallel with ongoing source-side inference and token delivery.
Tianhang Ding, Jianchun Liu, Hong-Li Xu
· 0 citations