Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
This work conducts a large-scale evaluation of more than 20 recent MLLMs and shows that video instruction following remains challenging for current models, especially for instructions with many constraints, semantic constraints, or complex conditional structures that require selecting the correct branch or path based on video content.