This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.
Hongbo Liu, Peixian Chen, Siyuan Liu et al.· 0 citations
Modality Subspace Activation (MSA) is proposed, a training-free inference-time framework that uses Singular Value Decomposition (SVD) to estimate modal activation strengths and dynamically balances modal projections in the last hidden state, effectively restoring CMS across benchmarks.
Hongbo Jiang, Jie Li, Yunhang Shen et al.· 0 citations
This work introduces A 2 -Judger, a novel MLLM-based A gentic instantiation of A uto Judger equipped with semantic-aware retrieval and dynamic memory that significantly improves sample efficiency while maintaining reliable evaluation results.
Xuanwen Ding, Chengjun Pan, Zejun Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.