Large Vision-Language Models (LVLMs) face significant computational inefficiencies caused by the large number of visual tokens. Existing visual token pruning methods mainly focus on either retaining individually important tokens or selecting mutually diverse ones. In this work, we revisit visual token pruning from a co...
Few-shot visual anomaly detection is fundamentally a visual comparison task, requiring fine-grained inspection of a query against normal references. Many recent methods based on large vision-language models (LVLMs) emphasize comparative reasoning through language chain-of-thought. Yet discrete, abstract descriptions ma...
Meng-Yang Zhao, Zhuo-Lin He, Hai-Yang Yu et al.· 0 citations
CRISP is proposed, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context that serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.
Xu Li, Yi Zheng, Mengyang Zhao et al.· arXiv.org· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.