Hierarchical Compression of Vision-Language Model Benchmarks
The analyses show how VLM evaluation behaves as model panels grow and evolve, providing guidance for designing future benchmarks that are more efficient, robust to model turnover, and explicit about the limits of evaluation-side pruning.
Hyunjong Ok, Seung-Gu Kang, Jaeho Lee
· 0 citations