Preprint
Aug 2026
HAP: Head-Adaptive Visual Token Pruning via Cross-Modal Alignment
PAQ (Prompt-Grounded Attention Quality), a metric quantifying how well each head aligns the prompt with image regions, is proposed and built on, which delivers state-of-the-art trade-offs on LLaVA-1.5-7B.
Yuan Sun, Huawei Ji, Yuanhao Jin et al.
· 0 citations