Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression, demonstrating its practicality for resource-constrained LVLM deployment.
Abstract
Visual token pruning is a promising approach to reducing the inference cost of large vision-language models (LVLMs), yet aggressive token reduction often causes substantial performance degradation. We identify three key factors behind this degradation: text-guided selection bias, information loss from discarded tokens, and positional distortion caused by sequence compaction. Based on these observations, we propose \textbf{VPRune}, a training-free pre-LLM pruning framework consisting of visual-only diversity selection, similarity-guided token recycling, and position-preserving restoration. Experiments on FastVLM-1.5B across multiple vision-language benchmarks demonstrate that VPRune achieves a favorable accuracy--compression trade-off, with particularly pronounced advantages under aggressive compression. Furthermore, evaluations on edge-device show that VPRune effectively reduces end-to-end inference latency while maintaining superior task performance, demonstrating its practicality for resource-constrained LVLM deployment.
CoverPruner is proposed, a training-free pruner that asks the complementary demand-side question: after a token is removed, which surviving original token represents it for the target VLM?
Qin Zhu, Wei-Hang You, Han-Qi Jiang et al.· 3 citations
Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We...
Large vision-language models (LVLMs) achieve strong multimodal understanding, but the hundreds to thousands of visual tokens they process impose substantial computational overhead, motivating training-free visual token pruning. In this work, we conduct two complementary analyses of visual token pruning. First, we measu...
Yi-Chen Guo, Ting-Hao Wang, Qi-Zhe Zhang et al.· 1 citation
Vision language models have achieved strong performance across a wide range of multimodal applications, yet their substantial computational and memory costs hinder efficient deployment. Visual token pruning and post-training quantization reduce inference overhead along two complementary dimensions, namely sequence leng...
Hai-Zhao Jing, Zhen-Hao Shang, Hao-Kui Zhang et al.· 0 citations
Braco is a lightweight four-step coder that combines transform-basis truncation, input-independent basis-coordinate embeddings, budget-dependent orthogonal re-parameterization, and learned spatial residual tokens from lightweight pooling that forms the favorable empirical accuracy-efficiency frontier under compression.
Rui-Lian Zhong, Yu Li, Zhe-Yu Yan et al.· 0 citations
VisCache is proposed, a plug-and-play framework for coarse-to-fine KV pruning without training that consistently outperforms existing baselines, establishing a new Pareto frontier between efficiency and performance for long-context VLLM inference.
Known for his clear and elegant writing style, Bertsekas shaped fields from control and optimization to large-scale computation and artificial intelligence.