Jul 2026· 2026 IEEE 9th International Conference on Big Data and Artificial Intelligence (BDAI)· pp. 96-102· 0 citations· 34 references
Abstract
The inference performance of Multimodal Large Language Models (MLLMs) typically relies on a massive number of visual block tokens, leading to substantial computational overhead. However, existing training-free compression methods often sacrifice semantic relevance, disrupt spatial coherence, or require extensive architectural modifications. Consequently, we propose Text-Conditioned Spectral Prototyping for Multimodal Token Compression (TCSP). We construct a joint affinity matrix based on both text relevance and visual similarity to derive boundary-continuous and semantically aligned segments via spectral embedding, subsequently replacing redundant information within these segments with weighted prototypes to reduce visual tokens for downstream inference. This method is entirely training-free and plug-and-play, allowing it to be inserted into intermediate layers of the vision encoder or before/after the projector module as needed. Evaluations on LLaVA-1.5-7B demonstrate that after a approximately $3 \times$ reduction in visual tokens, the overall performance remains baseline levels, even achieving marginal denoising-like improvements in several general QA benchmarks. More aggressive compression results in a graceful performance degradation, whereas tasks emphasizing compositional reasoning exhibit higher sensitivity to the compression scale. TCSP integrates text-awareness, spectral partitioning, and prototype replacement within a minimally invasive design. It achieves high, near-lossless compression ratios and a robust accuracy-efficiency trade-off without training, providing a concise and reliable pathway for general reasoning and resource-constrained deployment.
This work proposes Single-Forward Pruner (SFPruner), a structural reformulation of visual token pruning that embeds redundancy control directly into the scoring space, bypassing the need for iterative combinatorial optimization and achieves redundancy-aware importance selection in a single forward pass.
J. Song, Woohyeong Kim, Kyeongbo Kong· arXiv.org· 0 citations
Experiments on multiple representative video benchmarks show that CRAFT consistently outperforms prior state-of-the-art token-compression methods and shows significant efficiency improvement.
Yu Chen, Xiao-Hong Li, Xiaole Wang et al.· 0 citations
It is observed that attention scores from both vision and text tokens peak at modality separator tokens, suggesting that these separators bridge the two modalities and proposes SepPrune, an efficient, training-free, plug-and-play pruning method that uses the separator token as a unified query to rank and select informative vision tokens.
Yucheng Wang, Qihui Zhu, Yang Liu et al.· arXiv.org· 0 citations
VLZip is introduced, a framework that unifies visual and textual compression for high-fidelity reasoning within a pure Transformer, and establishes an efficient and powerful new standard for long-context multimodal AI.
Yuqi Zhang, Cheng Chen, Yuyu Guo et al.· 0 citations
Large Vision-Language Models (LVLMs) typically require processing hundreds to thousands of visual tokens, leading to substantial inference overhead. Existing visual token pruning methods either operate before the LLM using text-agnostic heuristics or prune inside the LLM at the cost of efficiency and noisy cross-modal attention. To address these limitations, we propose CRISP, a pre-LLM yet text-driven visual token pruning framework that preserves both instruction-relevant evidence and essential scene context. CRISP works in a two-stage pipeline: Stage 1 first identifies text-aligned visual tokens, and Stage 2 enhances contextual completeness through semantic diversity. Extensive experiments on LLaVA-1.5 and LLaVA-NeXT demonstrate that CRISP achieves superior performance retention under aggressive pruning ratios, maintaining up to 99.5% accuracy while reducing inference cost and latency by more than 2 times. CRISP serves as a practical solution for efficient LVLM inference, especially in resource-constrained scenarios.
Xu Li, Yi Zheng, Mengyang Zhao et al.· arXiv.org· 0 citations
Compared to prior CLIP-enhancement methods, MLLMCLIP achieves state-of-the-art compositional accuracy while delivering consistent gains on standard zero-shot classification and image-text retrieval, showing that feature-level distillation strengthens both compositional and general vision-language representation capability.
Jongsuk Kim, Qiyu Wu, Zhuoyuan Mao et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.