Research
LAION Releases BVD, 80M Videos and 10 Million Hours of Open Multimodal Pre-training Data From CommonCrawl
LAION-BVD starts from 1.3B platform-specific video URLs collected from CommonCrawl, from which 80M videos totaling 10 million hours were downloaded, then uses content-aware scene detection to cut clips and synthetically generates both video and audio captions. Models trained on it reach competitive results on standard video-text and audio-text benchmarks with consistent gains as training or model scale increases. The secondary result is arguably more interesting: extracting scene-changing frames yields image-text data with a visual distribution distinct from standard web image corpora, and models trained on it achieve strong image-text retrieval.
↳ Follow the thread