Skip to content

Hierarchical Compression of Vision-Language Model Benchmarks

Sep 2026 · 0 citations · 84 references
Computer Science

TL;DR

The analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.

Abstract

Thorough evaluation of vision-language models (VLMs) has become prohibitively expensive, as benchmarks span an ever-broader spectrum of capabilities and new models arrive at a relentless pace. Benchmark compression methods that preserve model rankings at a fraction of the cost are well studied for language models, but for VLMs the question remains under-explored. We present PRIMEBench (Pruning Redundant Items for Multimodal Evaluation), a vision-aware hierarchical benchmark compression framework that substantially reduces evaluation cost while preserving model rankings. This hierarchical framework operates in four stages: data cleaning to remove items answerable without the image and all-correct items, category representative selection to pick one benchmark per capability category, item pruning with Vision-Aware Variance (VAW), and category-count pruning. VAW combines inter-model variance with a vision-dependence score computed from multimodal embeddings alone, while encouraging coverage of diverse items within each benchmark. On models held out from item selection, it has the highest mean fidelity at the released 5% retention. The hierarchical design lets practitioners stop at any stage to match their compute budget; the released suite removes over 97% of items while preserving model rankings. Beyond compression, our analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.

View source

Similar papers

#natural language process... Preprint Sep 2026

Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models

This work presents ZipBench, a simple and low-cost BCM with theoretical error and rank-consistency guarantees that reduces the cost of both LLM evaluation and compact benchmark construction, lowering the barrier to broad LLM research for compute-constrained researchers.

Zhongzhan Huang, Jun-Xin Li, Guo-Ming Ling et al. · 0 citations
#natural language process... Preprint Sep 2026

PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models

PRISM-VLM is introduced, a multi-axis discriminative benchmark that scores every item along seven axes covering the recurring failure modes (task quality, behavioral robustness, and capability bottlenecks) and combines them into a single PScore, with items recycled from fifteen public benchmarks.

Sanghee Park, Kee-Eung Kim · 0 citations
Preprint Aug 2026

Prune Once: Retraining-Free Task-Agnostic Pruning for Vision-Language Models

A retraining-free VLM pruning framework called PORTA is introduced that derives a task- and modality-agnostic importance formulation based on activation variation, estimated from generic calibration data, which reliably captures feature-level representation utility across modalities.

Minseok Kang, Hyunwoo J. Kim, Chanyoung Kim et al. · 1 citation
Review Aug 2026

Which Source Wins? Task-Dependent Reliance in Vision-Language Models

Modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, models, and evaluation settings, which shows that modality reliance in VLMs is not fixed, but varies across tasks, evidence structures, and evaluation settings.

Rodela Ghosh, Aviral Gupta, Guang-Jing Wang · 0 citations
Review Aug 2026

Not All Attention Is Equal: A Quantitative Survey of the EEI Trade-off

This survey traces attention from Bahdanau-Luong alignment through the Transformer and into vision architectures, and reviews fixed and learned sparse attention, linear attention, IO-aware exact algorithms including FlashAttention, and state-space alternatives including Mamba.

Aditya Singh · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.