PReM: Learning What to Preserve and When to Refresh for Context Compression
PReM (Preserve and Refresh Memory), a context-compression framework that maintains the long context as the model's internal layer-wise KV memory and learns what to preserve and when to refresh it, is introduced.