LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
This work presents LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl, and explores video frames as an alternative source of image-text data by extracting scene-changing frames.