Preprint
Aug 2026
Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
A fine-grained dynamic visual selection framework named EviSelect, grounded in the target MLLM internal attention evidence, which efficiently approximate attention maps of the target MLLM using highly compressed visual inputs and sparse attention, well-aligned to the full counterpart.
Bo Zhang, Wenxin Wang, Feng Chen et al.
· 0 citations