A structured, comprehensive survey of the latest MVU progress is presented, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning.
Abstract
Multimodal video understanding (MVU) has emerged as a fast-growing research frontier, driven by major advances in video-language pre-training and large multimodal models over the past decade. MVU aims to synergistically integrate visual, audio and textual modalities to interpret complex video semantics, supporting widespread downstream tasks including cross-modal retrieval, dense captioning, video question answering, event analysis and intelligent assistance. Despite the rapid proliferation of specialized MVU models, the community still lacks a unified capability-centric framework to systematically clarify the hierarchical competency architecture and evolutionary trajectory of state-of-the-art approaches. To address this issue, this paper presents a structured, comprehensive survey of the latest MVU progress, establishing a novel three-tier taxonomy that categorizes existing studies into cross-modal alignment, multi-granularity semantic expression and multimodal reasoning. Along this pipeline, we further systematically synthesize core modality fusion strategies, mainstream benchmark datasets and standardized evaluation protocols. Through a fine-grained analysis of representative published results, we highlight the critical impact of inconsistent evaluation settings, cross-experiment comparability bottlenecks and inherent methodological trade-offs between performance and efficiency. Finally, we identify and dissect three key open challenges: ultra-long video scalability, performance degradation from modality noise and missing data, and factual reliability risks in generative MVU systems. This capability-oriented systematic reference clarifies the methodological evolution logic of MVU, and provides actionable guidance for developing next-generation robust, high-performance multimodal video understanding systems.
This survey investigates the current research landscape of multimodality modeling from three perspectives: the first group of multimodal models adopts a heterogeneous architecture to bridge different modality data, the second leverages LLM for multimodality modeling via a unified language modeling objective, and the third represents multimodal data entirely within a single visual representation.
Zhongfen Deng, Yibo Wang, Yueqing Liang et al.· 2 citations
A new prototype-based hierarchical alignment network (PHA-Net) to align individual/local/global level representations across modalities and introduces multiple modality-shared prototypes as the bridge to efficiently optimize text and video representations for cross-modal alignment.
Xiaolun Jing, Kezhao Yin, Xinxing Yang et al.· Neurocomputing· 0 citations
Video summarization aims to produce concise representations of videos by selecting the most informative frames or shots. Existing methods predominantly rely on unimodal visual features extracted from convolutional neural networks, overlooking the rich semantic information that natural language can provide. In this paper, we present an integrated multimodal video summarization framework that makes two tightly coupled contributions. Initially, we construct a semantically enriched multimodal dataset by augmenting standard video summarization benchmarks (SumMe, TVSum, OVP, and YouTube) with BLIP‐2‐generated frame‐level captions and CLIP‐based aligned visual–textual embeddings, producing four feature variants—Image‐only, Text‐only, Averaged fusion, and concatenated fusion—for systematic analysis of modality contributions. We validate the dataset through cross‐modal retrieval, similarity gap analysis, Spearman complementarity correlation, and caption quality diagnostics, confirming that the generated captions are semantically meaningful and that the two modalities provide complementary information (average Spearman ). Finally, we develop a multi‐scale temporal summarization architecture comprising a U‐Net‐based temporal backbone and a hierarchical transformer‐style scoring head. Our best configuration achieves
F
1‐scores of 59.27% on SumMe, competitive with recent state‐of‐the‐art methods while maintaining only 28.27M parameters with linear computational scaling.
Saadman Sakib, Kaushik Deb· Applied AI Letters· 0 citations
This paper presents a multimodal pipeline that combines vision-language captioning models and openvocabulary object detectors to investigate the impact of automatically generated textual prompts on semantic image understanding and reveals complementary behaviors between Grounding DINO and OWLv2.
Xin Gao, Madjid Maidi, B. Daachi· NLP & Big Data· 0 citations
This work proposes Video-GAR, a novel framework that reframes the conventional retrieval task from superficial matching to generative understanding, positing that the capability for query reconstruction evidences deep semantic comprehension.
Mingjin Kuai, Qianyin Xiao, Juncheng Li et al.· Annual International ACM SIG...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.