# LAION-BVD：面向多模态预训练的千万小时级开放视频数据集

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-08-25 08:00
- AIHOT 分数：54
- AIHOT 链接：https://aihot.virxact.com/items/cmt9kopdm0n37rolywapr8ohz
- 原文链接：https://arxiv.org/abs/2608.24845

## AI 摘要

LAION 团队发布 LAION-BVD，一个包含 1.3B 条平台视频 URL、下载 80M 段视频、总时长 1000 万小时的大规模开放视频数据集，面向视频、音频和图像多模态预训练。基于该数据训练的模型在视频-文本和音频-文本基准上表现具竞争力，且性能随训练或模型规模提升而持续增长。数据集已向研究社区开放。

## 正文

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.
