Skip to content

Author

Kaushik Deb

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Aug 2026

Multimodal Video Summarization Using Vision‐Language Embeddings and Hierarchical Temporal Modeling

Video summarization aims to produce concise representations of videos by selecting the most informative frames or shots. Existing methods predominantly rely on unimodal visual features extracted from convolutional neural networks, overlooking the rich semantic information that natural language can provide. In this paper, we present an integrated multimodal video summarization framework that makes two tightly coupled contributions. Initially, we construct a semantically enriched multimodal dataset by augmenting standard video summarization benchmarks (SumMe, TVSum, OVP, and YouTube) with BLIP‐2‐generated frame‐level captions and CLIP‐based aligned visual–textual embeddings, producing four feature variants—Image‐only, Text‐only, Averaged fusion, and concatenated fusion—for systematic analysis of modality contributions. We validate the dataset through cross‐modal retrieval, similarity gap analysis, Spearman complementarity correlation, and caption quality diagnostics, confirming that the generated captions are semantically meaningful and that the two modalities provide complementary information (average Spearman ). Finally, we develop a multi‐scale temporal summarization architecture comprising a U‐Net‐based temporal backbone and a hierarchical transformer‐style scoring head. Our best configuration achieves F 1‐scores of 59.27% on SumMe, competitive with recent state‐of‐the‐art methods while maintaining only 28.27M parameters with linear computational scaling.

Saadman Sakib, Kaushik Deb · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.