This work proposes a novel zero-shot video captioning framework (WSV) consisting of two training stages, which first generates corresponding synthetic video latent representations via a pretrained text-to-video generation model, and designs a polisher capable of bridging the gap between real and synthetic video distributions.
Abstract
Text-only training is a popular paradigm in zero-shot video captioning, where the video distribution is not available to the model during training, leading to a cross-modal gap between the training (text-only) and the inference (video-only). Previous works attempt to bridge the gap through simple linear transformations. However, the inherent gap between text and video makes cross-modal representation space alignment insufficient, resulting in inaccurate sentences. To address this issue, we propose a novel zero-shot video captioning framework (WSV) consisting of two training stages, which first generates corresponding synthetic video latent representations via a pretrained text-to-video generation model. To strengthen the fidelity of the latent representations, we propose a polisher capable of bridging the gap between real and synthetic video distributions. Subsequently, we design a prompter that conditions GPT-2 on the polished latent representations to generate the captions in the second training stage. During inference, an input video is encoded by a pretrained 3D Causal VAE and then fed directly into the prompter, which in turn guides GPT-2 to produce the final caption. Experimental results conducted on MSVD, MSR-VTT, and VATEX datasets demonstrate that our proposed method achieves scores of 52 and 95.7 on the B@4 and CIDEr metrics, respectively.
OmniVAE is presented, a jointly trained audio-video VAE that learns fine-grained semantic alignment between audio and video latent representations that translates into higher generation quality and more accurate cross-modal synchronization in downstream text-to-audio-video generation.
Jun Zhan, Chenchen Yang, Yitian Gong et al.· arXiv.org· 0 citations
A novel framework that performs selective cross-modal alignment through a learnable masking mechanism, enabling the model to isolate and align only the shared latent components relevant to both modalities is proposed.
Jiyang Zheng, Siqi Pan, Yu Yao et al.· Neural Information Processin...· 6 citations
Visual content is the dominant medium of communication, yet without audio descriptions (ADs), it remains inaccessible to blind and low-vision people. ADs narrate context-relevant visual events during natural audio pauses. Manually creating ADs is expensive, limiting coverage to a small fraction of available content. Most existing automatic AD generation methods frame the task as video clip captioning, requiring ground-truth timestamps and additional context cues such as character databases. Current benchmarks reinforce this framing, consisting of short video segments paired with automatic or task-mismatched annotations. We introduce StrAD, a benchmark for long-form AD generation on full-length videos spanning diverse genres such as movies, documentaries, short films, performances, and video games. We reformulate AD generation as streaming dense video captioning. Our approach processes full-length videos with a sliding window, inserting ADs into existing transcripts without ground-truth timestamps, and supports both fine-tuned models and zero-shot prompting of vision-language models. On the segment-level task with given timestamps, our fine-tuned StrAD-FT sets the state of the art on CMD-AD with 36.3 CIDEr (+10.0 over Shot-by-shot), establishes a reference point on StrAD (51.0 CIDEr), and remains competitive on MAD-Eval at 24.9 CIDEr. On the full-video streaming task, StrAD-FT reaches a SODA score of 2.4 against 1.1 for our zero-shot baseline StrAD-Zero, though both exhibit limitations in temporal localization and narrative coherence. While prior work has tackled full-video AD generation in an offline, multi-stage fashion, ours is the first streaming approach, generating ADs on the fly without ground-truth timestamps. StrAD makes progress on full-video AD generation measurable, a prerequisite for scaling accessibility.
Julian Spravil, Sebastian Houben, Sven Behnke· 0 citations
An uncertainty-aware variational latent anchoring module that dynamically selects informative frames and compresses cross-frame latents into a compact set of semantic anchors that achieves state-of-the-art inversion fidelity, editing quality, and temporal consistency, while reducing memory and runtime compared with prior backbone training-freebaselines.
Zhangkai Wu, Xuhui Fan, Zhongyuan Xie et al.· Proceedings of the 32nd ACM...· 0 citations
Recently, the zero-shot image captioning (zero-shot IC) method based on pre-trained visual language models (VLMs) and large language models (LLMs) has made significant progress. However, how to adapt it to the zero-shot video captioning (zero-shot VC) scenario (without video-text paired supervision) has not been well explored. Inspired by various recent test-time strategies (sacrificing additional test time to improve performance), we try to introduce a new paradigm of Test-time Reinforcement Polishing in zero-shot VC scenario. We take temporal dependency modeling as the starting point and propose a novel framework for Refinable Zero-shot VC, called RefZVC. RefZVC can greatly cover the long-term context of the video and continuously polish and refine the generated captions in a reward-feedback manner. We first design an Adaptive Frame Skipping module (AdaSkip) to skip redundant frames and select diverse keyframe sequences. Subsequently, we propose a Multi-granularity Reinforcement Polishing (MRP) mechanism, which iteratively polishes captions by leveraging Gaussian Kernel Cache (GKC) to capture temporal dynamics, store and reuse relevant historical context. In addition, MRP calculates rewards for generated captions at both the sentence-level and entity-level to achieve test-time polishing. With the MRP mechanism, RefZVC achieves superior zero-shot generalization performance, outperforming previous zero-shot VC methods on benchmarks such as MSVD, MSR-VTT, and VATEX.
Qianyue Bao, Fang Liu, Licheng Jiao et al.· IEEE Transactions on Image P...· 0 citations
This work proposes Self-SiMS, a self-similarity-based Moment Proposal and Scoring that exploits intrinsic relationships within videos, enabling robust span generation and scoring and introduces a query-aware MLLM-based reasoning stage to further sharpen alignment between text and video.
Jihyun Lee, Cheol-Ho Cho, Woojin Jun et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.