TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning
TReVS is proposed, a training-free framework that combines textual relevance with vision-encoder saliency for pre-LLM pruning and leverages high-variance attention heads to remove task-irrelevant tokens at shallow-to-intermediate layers of the LLM.