We have 2 of 21 papers
We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.
Not the right person? Other researchers publish under this name.
When Vision Becomes Text: Visual Token Pruning via Cross-Modal Residual Guidance in VLMs
This paper revisit VLM inference and presents a new efficient guidance scheme that complements similarity-based guidance, and proposes Cross Modal Residual (CMR), a training-free visual token compression method that combines CMR, text-attention relevance, and residual-space diversity to retain task-relevant and complementary tokens.