Skip to content
Preprint

EviViT: Evidence-Adaptive Vision Transformers for Fine-Grained Perception

Sep 2026 · 0 citations · 33 references
Computer Science

TL;DR

EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail, is introduced, a lightweight attachment that learns where a pretrained vision transformer should acquire detail without refitting.

Abstract

Fine-grained visual perception enables vision-language models to distinguish subtle attributes and ground their answers in visual evidence. In high-resolution scenes, processing the whole image at greater resolution spends visual tokens on irrelevant content, while isolated crops can lose the context needed to interpret the selected evidence. We introduce EviViT, a lightweight attachment that learns where a pretrained vision transformer should acquire detail. Human visual-search traces supervise a question-conditioned evidence density, which guides regional re-reading from the original pixels and the allocation of visual tokens. A sparse, coordinate-aware bridge then connects the regional features to the global scene, allowing the host to interpret precise evidence in context. Learned with the host backbone frozen, the attachment serves both the base model and compatible post-trained descendants without refitting. Experiments across nine hosts show consistent gains in average fine-grained accuracy. Matched-budget comparisons further show that EviViT outperforms global-only processing at every tested token ceiling while using fewer visual tokens.

View source

Similar papers

Preprint Aug 2026

State-Conditioned Visual Evidence Retrieval for Fine-Grained Perception in Document Vision-Language Models

Experiments on document parsing benchmarks show that SCVER improves robustness under reduced input resolution and achieves a better accuracy-efficiency trade-off, demonstrating the effectiveness of on-demand visual evidence retrieval for fine-grained perception.

Ming-Xu Chai, Chen-Yu Liu, Zi-Yu Shen et al. · 0 citations
Preprint Aug 2026

E2S-Pruner: Progressive Two-Stage Evidence Fusion for Visual Token Pruning in Vision-Language Models

A spatial novelty constraint is introduced that promotes coverage of distinct image regions and prevents the retained tokens from concentrating in a few locally salient areas and prevents the retained tokens from concentrating in a few locally salient areas in E2S-Pruner.

Taoyu Qian, Qi Wang, D. Shi et al. · 1 citation
#artificial intelligence Preprint Sep 2026

EviRover: Reinforcing Agentic Perception Beyond a Glance

Visual perception is conventionally formulated as a one-shot prediction from a single glance at the image, under the assumption that the image content and the model's parametric knowledge suffice to resolve the query. This assumption often fails in real-world scenarios that hinge on fine-grained visual details or requi...

Kai-Xuan Fan, Kai-Tuo Feng, Tian-Shuo Peng et al. · 0 citations
Preprint Sep 2026

Region-Level Policy Optimization for Fine-grained MLLM Perception

Fine-grained visual perception in MLLMs is commonly improved by raising the resolution, but the added visual tokens inflate vision-encoding and language-model prefilling costs. We show that the two operations underlying fine-grained perception, localizing the region of interest (RoI) and recognizing its content, have d...

Yuheng Shi, Xiao-Huan Pei, Min-Jing Dong et al. · 0 citations
Open access Aug 2026

RoFLIP: Robust and Fine-Grained Alignment for Vision-Language Compositional Reasoning

The Robust and Fine-grained training framework for CLIP-based vision-language models (RoFLIP) is proposed, enhancing both the robustness and granularity of vision-language alignment and underscore RoFLIP’s compositional reasoning and generalization abilities.

Yiwei Sun, Chuan-Bin Liu, Shancheng Fang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.