Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression
Structured pruning of attention heads provides a hardware-friendly way to compress Transformer language models. However, existing methods for measuring head-level importance require calibration data, gradient computation, or Hessian estimation. These requirements add extra overhead and make the methods depend on the da...