Conference
Open access
2026
Vista-LLM: Decoupled Query-Guided Visual Token Pruning for Efficient Long-Video Large Language Models
Vista-LLM is introduced, a decoupled framework for query-guided visual token pruning that reduces visual tokens by 90% and accelerates inference while retaining over 98% of baseline performance on average, effectively filtering visual noise.
Zhenyu Li, Zuchao Li, Ping Wang et al.
· Annual Meeting of the Associ... · 0 citations