#artificial intelligence
May 2026
Locality-Aware Redundancy Pruning for LLM Depth Compression
It is shown that inter-layer redundancy can be either localized or globally distributed depending on the LLM architecture, and Representation Locality Score (RLS) is introduced, derived from global inter-layer hidden-state similarity.
Vincent-Daniel Yun, Youngrae Kim, Woosang Lim et al.
· arXiv.org · 1 citation
· ⚡1