EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail, is introduced, a lightweight attachment that learns where a pretrained vision transformer should acquire detail without refitting.
Abstract
Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.
Experiments on document parsing benchmarks show that SCVER improves robustness under reduced input resolution and achieves a better accuracy-efficiency trade-off, demonstrating the effectiveness of on-demand visual evidence retrieval for fine-grained perception.
Ming-Xu Chai, Chen-Yu Liu, Zi-Yu Shen et al.· 0 citations
A spatial novelty constraint is introduced that promotes coverage of distinct image regions and prevents the retained tokens from concentrating in a few locally salient areas and prevents the retained tokens from concentrating in a few locally salient areas in E2S-Pruner.
FAVE (Foveated Adaptive Visual Encoding), a lightweight variable-resolution ViT that encodes externally selected regions at high acuity while preserving native geometry, is introduced and integrated as a complementary local branch in FastVLM.
Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or requi...
Kai-Xuan Fan, Kai-Tuo Feng, Tian-Shuo Peng et al.· 0 citations
Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have d...
Yuheng Shi, Xiao-Huan Pei, Min-Jing Dong et al.· 0 citations
The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.
Yiwei Sun, Chuan-Bin Liu, Shancheng Fang et al.· International Journal of Com...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.