Preprint
Aug 2026
DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models
DIVE (Dynamic Iterative Visual Evidence Construction), a training-free framework that recasts visual-token pruning as dynamic evidence construction and builds a retained set of complementary, prompt-relevant evidence.
Cheng Zhong, Xiao An, Zijie Wang et al.
· 0 citations