Non-profit research organization LAION has released the Big Video Dataset (BVD), a open dataset designed for training multimodal AI models. Sourced from 1.3 billion video URLs identified in CommonCrawl, the dataset comprises 10 million hours of downloaded video across 80 million files. From this collection, LAION extracted 55 million curated clips featuring automated audio and visual descriptions, alongside 300 million still frames.
According to the published research paper, AI models trained on BVD demonstrated improvements over existing public baselines, outperforming comparable models trained on InternVid by up to 2.1 percentage points on video-to-text benchmarks. The dataset aligns visual data, native audio, and text annotations to advance research in multimodal representation learning. LAION released the dataset and associated codebase exclusively for non-commercial research purposes.
Why it matters
Provides an open dataset containing 10 million hours of video data for multimodal and video foundation model research.
Demonstrates benchmark performance improvements over existing open baselines like InternVid for video-to-text tasks.
Offers early-stage startups and academic researchers open training data to compete with proprietary lab video models.
Source: the-decoder.com



