#artificial intelligence
May 2026
How to Compress KV Cache in RL Post-Training? Shadow Mask Distillation for Memory-Efficient Alignment
This work states that the sampler generates responses under a sparse context, whereas the learner updates parameters using the full, dense context, whereas the sampler updates parameters using the full, dense context of the RL framework.
Rui Zhu, Wei-Heng Bai, Qiu-Shi Wu et al.
· arXiv.org · 2 citations