Beyond Sparse Weights: When Is Attention Compressible?
Results motivate CertKV, a training-free compressor that reserves one tail-summary slot per head and allocates the rest by value dispersion, which is top-two in seven of nine LongBench-v2 settings, remains in the leading compressed tier on 128K RULER, and realizes a ten-fold cache budget in a packed Llama prototype.