This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.
Hongbo Liu, Peixian Chen, Siyuan Liu et al.· 0 citations
This work proposes leveraging the model’s intrinsic self-evaluation to guide its optimization, and designs two novel reward functions: Sequential Confidence Rigorous Evaluation (SCRE) for challenging problems that demand strict logical reasoning, and intra-group Score Re-ranking (IGSR) for general-purpose, open-ended scenarios.
Xing Xi, Yushu Qiu, Ronghua Luo et al.· 0 citations
This work introduces A 2 -Judger, a novel MLLM-based A gentic instantiation of A uto Judger equipped with semantic-aware retrieval and dynamic memory that significantly improves sample efficiency while maintaining reliable evaluation results.
Xuanwen Ding, Chengjun Pan, Zejun Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.