TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning
TReVS is proposed, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM.
Jing Wang, Zhi-Ping Wu, Dong-Dong Ren et al.
· 0 citations