Review
Open access
Jul 2026
Multimodal Video Understanding: A Capability-Based Survey of Alignment, Expression, and Reasoning
A structured, comprehensive survey of the latest MVU progress is presented, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning.
Rongyong Zhao, Da Pu, Cuiling Li et al.
· Applied Sciences · 0 citations