Book
Open access
Jul 2026
Scaling Attention Beyond GPUs for LLM Inference
Beyond is presented, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference.
Weishu Deng, Yujie Yang, Peiran Du et al.
· IEEE International Symposium... · 1 citation