Skip to content
Preprint

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Aug 2026 · 0 citations
Computer Science

TL;DR

This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.

Abstract

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.

View source

Similar papers

Sep 2026

OpenVid++: A Large-scale High-quality Dataset for Text-to-video Generation

A Motion-Aware Frame Sampling (MAFS) strategy for constructing motion-aware captions enables more effective utilization of video frames and improves captioning quality for large-motion videos.

Ke-Pan Nan, Tie-Han Fan, Rui Xie et al. · 1 citation
Preprint Aug 2026

ClearText-Video: A Large-Scale Text-Centric Video Dataset Bridging Video Restoration and Scene-Text Enhancement

Multimodal Large Language Models (MLLMs) have recently made strong progress in visual--linguistic understanding. However, their performance on text-centric video reasoning remains highly sensitive to input quality. Real-world user-provided videos often contain motion blur, compression artifacts, noise, and low-resoluti...

Jinlong Li, Jiaming Ding, Ding-Fu Lu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation

Video Object Segmentation (VOS) aims to accurately segment target objects from consecutive video frames and track the changes of the objects in each frame of the video. Conventional VOS methods typically demand substantial quantities of pixel-level labeled video sequences for fully supervised learning, which limits the...

Zheng Zhu, Jia-Qing Fan, Han-Wen Qian et al. · 0 citations
Preprint Sep 2026

PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending...

Cong Wei, Xuan-Chi Ren, Bryan Chu et al. · 0 citations
Preprint Aug 2026

StrAD: A Streaming Method and Benchmark for Audio Description Generation for Long-form Videos

This work introduces StrAD, a benchmark for long-form AD generation on full-length videos spanning diverse genres such as movies, documentaries, short films, performances, and video games, and reformulates AD generation as streaming dense video captioning, generating ADs on the fly without ground-truth timestamps.

Julian Spravil, Sebastian Houben, Sven Behnke · 0 citations
Conference Aug 2026

A Two-Stage CNN-LSTM Framework for Spatiotemporal Deepfake Video Detection

Now a days identification of convincing synthetic videos created with the help of deepfake is a big challenge. Unfortunately, deepfakes represent a serious threat to the integrity of media, as they can cause individuals to lose trust in the digital content they see. Among all types of deepfakes, face-swap videos are ex...

Pratik Bhosale, Rujul Rajarapollu, Kartikya Durgesh Gawali et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.