This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.
Abstract
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.
A Motion-Aware Frame Sampling (MAFS) strategy for constructing motion-aware captions enables more effective utilization of video frames and improves captioning quality for large-motion videos.
Ke-Pan Nan, Tie-Han Fan, Rui Xie et al.· International Journal of Com...· 1 citation
Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos often contain motion blur, compression artifacts, noise, and low-resoluti...
Jinlong Li, Jiaming Ding, Ding-Fu Lu et al.· 0 citations
Video Object Segmentation (VOS) aims to accurately segment target objects from consecutive video frames and track the changes of the objects in each frame of the video. Conventional VOS methods typically demand substantial quantities of pixel-level labeled video sequences for fully supervised learning, which limits the...
Zheng Zhu, Jia-Qing Fan, Han-Wen Qian et al.· 0 citations
Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending...
Cong Wei, Xuan-Chi Ren, Bryan Chu et al.· 0 citations
This work introduces StrAD, a benchmark for long-form AD generation on full-length videos spanning diverse genres such as movies, documentaries, short films, performances, and video games, and reformulates AD generation as streaming dense video captioning, generating ADs on the fly without ground-truth timestamps.
Julian Spravil, Sebastian Houben, Sven Behnke· 0 citations
Now a days identification of convincing synthetic videos created with the help of deepfake is a big challenge. Unfortunately, deepfakes represent a serious threat to the integrity of media, as they can cause individuals to lose trust in the digital content they see. Among all types of deepfakes, face-swap videos are ex...
Pratik Bhosale, Rujul Rajarapollu, Kartikya Durgesh Gawali et al.· International Conference on...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.