Skip to content

Author

Waseem Ullah

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

Cross-modal Fusion Transformer for Integrating Retrieved Knowledge into Video Caption Generation

In recent years, long-form video captioning has been an important task for indexing, accessibility, and downstream video analytics. However, it becomes challenging when videos are untrimmed and extend over minutes or hours. In these settings, visual evidence is often incomplete: objects are occluded, viewpoints shift, and crucial semantics may be implied rather than directly observable. Although recent hierarchical captioning models improve the temporal coverage, they still rely primarily on the available visual stream and can produce captions that omit key entities or relations. External knowledge retrieved from training data or knowledge bases can improve the visual signal, but existing integrations are often shallow and are either added late at the decoder input or applied without accounting for retrieval reliability, making them sensitive to noisy evidence and limiting cross-modal interaction. To solve this issue, we propose a video-aware knowledge fusion approach that integrates retrieved knowledge at the feature level prior to generation. The method combines (i) a contextual gating mechanism that modulates knowledge strength using retrieval confidence signals together with pooled visual context, and (ii) a cross-modal fusion transformer that refines the joint sequence of video and gated-knowledge query tokens through self-attention before decoding. We evaluate on challenging datasets such as YouCook2 and Ego4D-HCap, where our full model improves CIDEr by up to +11.1% on YouCook2 and by +4.2% / +5.8% on Ego4D-HCap segment descriptions and video summaries, compared with a strong baseline. This indicates that feature-level knowledge fusion can enhance semantic coverage in long-form caption generation.

K. Abubakirova, Waseem Ullah, Latif U. Khan et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.