메뉴
HN
Hacker News • 30일 전

LAION, 1천만 시간 규모 오픈 비디오 데이터셋 공개

IMP
7/10
핵심 요약

LAION이 CommonCrawl에서 수집한 13억 개의 비디오 URL을 기반으로 총 1천만 시간 분량의 8천만 개 비디오를 포함한 대규모 오픈 데이터셋 'LAION-BVD'를 연구 커뮤니티에 공개했습니다. 이 데이터셋은 비디오·오디오·이미지-텍스트 멀티모달 사전학습용으로 설계되었으며, 장면 감지로 추출한 클립에 합성 캡션을 생성했습니다. 이 데이터로 학습된 모델들은 표준 비디오-텍스트·오디오-텍스트 벤치마크에서 경쟁력 있는 성능을 보였고, 특히 InternVid 기반 모델 대비 최대 2.1% 향상된 성과를 달성했습니다.

번역된 본문

오픈 리서치 데이터셋 → LAION 빅 비디오 데이터셋: 비디오, 오디오, 이미지-텍스트 사전학습을 위한 1천만 시간 규모의 오픈 웹 비디오 코퍼스

저자: Andreas Hochlehnert*, Marianna Nezhurina*, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke◇, Jenia Jitsev◇, Matthias Bethge◇ (*공동 제1저자, ◇공동 교신저자)

소속: 튀빙겐 AI 센터(튀빙겐 대학교), LAION, JSC/FZJ, GRASS, 지능시스템 막스플랑크연구소(ELLIS 튀빙겐), MCML(뮌헨공대)

[개요 및 초록]

우리는 멀티모달 학습을 위한 대규모 오픈 비디오 데이터셋인 LAION-BVD(LAION Big Video Dataset)를 발표합니다. 이 데이터셋은 CommonCrawl에서 수집한 13억 개의 플랫폼별 비디오 URL을 포함하며, 이 중에서 총 재생 시간 1천만 시간에 달하는 8천만 개의 비디오를 다운로드했습니다. 데이터셋은 비디오, 오디오, 이미지 멀티모달 사전학습을 위해 설계되었습니다.

콘텐츠 인식 장면 감지(content-aware scene detection)를 사용해 클립을 추출하고, 이 클립에 대해 합성적으로 비디오 및 오디오 캡션을 생성했습니다. 이 데이터로 학습된 모델은 표준 비디오-텍스트 및 오디오-텍스트 벤치마크에서 경쟁력 있는 성능을 달성했으며, 학습 규모나 모델 규모가 커질수록 일관되게 성능이 향상되었습니다.

또한 장면이 전환되는 프레임을 추출하여 비디오 프레임을 이미지-텍스트 데이터의 대체 출처로 활용하는 방안을 탐구했습니다. 이러한 프레임은 일반적인 웹 이미지 코퍼스와는 다른 시각적 분포를 보이며, 이 데이터셋으로 학습된 모델은 우수한 이미지-텍스트 검색 성능을 달성했습니다.

우리는 LAION-BVD를 연구 커뮤니티에 공개합니다. 이는 전례 없는 규모로 멀티모달 비디오에 대한 오픈 액세스를 크게 확장합니다.

[주요 수치]

  • 비디오 URL: Common Crawl에서 수집한 플랫폼별 URL 13억 개
  • 다운로드된 비디오: 성공적으로 다운로드 및 처리된 비디오 8천만 개
  • 총 재생 시간: 모든 다운로드 비디오를 합친 1천만 시간
  • 어노테이션된 클립: 비디오 캡션이 생성된 클립 (5,500만 개)
  • 추출된 프레임: 이미지-텍스트 사전학습용 비디오 프레임 (3억 개)

[벤치마크 실험 결과]

LAION-BVD로 학습된 모델은 비디오, 오디오, 이미지-텍스트 벤치마크 전반에서 경쟁력 있는 성능을 달성했습니다.

ViCLIP 비디오-언어 벤치마크: LAION-BVD로 학습된 ViCLIP 모델은 표준 비디오-텍스트 벤치마크에서 InternVid로 학습된 모델과 동등하거나 최대 2.1% 높은 성능을 보였으며, 학습 규모가 1천만 클립에서 5천만 클립으로 증가할수록 일관된 향상을 보였습니다.

CLAP 오디오-언어 벤치마크: LAION-BVD로 학습된 CLAP 모델은 비디오에서 직접 추출한 풍부한 실제 환경 사운드스케이프를 활용하여 다른 대규모 비선별 오디오 데이터셋과 경쟁력 있는 성능을 달성했습니다. 모델과 데이터 규모를 늘릴 때 좋은 확장 추세를 보였습니다.

CLIP 이미지-텍스트 벤치마크: 프레임 기반 CLIP 모델은 표준 벤치마크에서 우수한 이미지-텍스트 검색 성능을 달성했습니다. 비디오 프레임은 일반적인 웹 코퍼스와 다른 시각적 분포를 보여 기존 이미지 사전학습 출처를 보완합니다.

[책임 있는 사용]

윤리 및 공개 성명: LAION-BVD는 대규모의 개방적이고 재현 가능한 멀티모달 연구를 지원하기 위해 공개되었습니다.

원문 보기
원문 보기 (영어)
Open Research Dataset --> LAION Big Video Dataset A 10-Million-Hour Open Web Video Corpus for Video, Audio, and Image-Text Pre-training --> Andreas Hochlehnert 1,* · Marianna Nezhurina 2,3,* · Mehdi Cherti 2,3 · Andrej Radonjic 4 · Thaddäus Wiedemer 5,1 · Christoph Schuhmann 2 · Romain Beaumont 2 · Wieland Brendel 5 · Bernhard Schölkopf 5 · A. Sophia Koepke 1,6,◇ · Jenia Jitsev 2,3,◇ · Matthias Bethge 1,◇ * Shared first authors  •  ◇ Shared last authors 1 Tübingen AI Center, University of Tübingen 2 LAION 3 JSC, FZJ 4 GRASS 5 MPI for Intelligent Systems, ELLIS Institute Tübingen 6 MCML, Technical University Munich --> A dataset for models that see --> Read Paper Download Dataset GitHub * Shared first authors  •  ◇ Shared last authors 1 Tübingen AI Center, University of Tübingen 2 LAION 3 JSC, FZJ 4 Wynd Labs 5 MPI for Intelligent Systems, ELLIS Institute Tübingen 6 MCML, Technical University Munich Affiliated Institutions --> --> Overview Abstract We present LAION-BVD (LAION — Big Video Dataset), a large-scale open video dataset for multimodal learning, containing 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours . The dataset is designed for multimodal pre-training across video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale. By the Numbers Unprecedented Scale The largest openly accessible video corpus for multimodal learning research 🔗 0 Video URLs Platform-specific URLs collected from Common Crawl 🎬 0 Downloaded Videos Successfully downloaded and processed videos ⏱️ 0 Total Duration Combined video content across all downloads ✂️ 0 Annotated Clips Clips with generated video captions 🖼️ 0 Extracted Frames Video frames for image-text pre-training Dataset Dataset Statistics --> Benchmarks Experimental Results Models trained on LAION-BVD achieve competitive performance across video, audio, and image-text benchmarks ViCLIP Video-Language Benchmarks ViCLIP models trained on LAION-BVD match or exceed InternVid-trained models by up to 2.1% on standard video-text benchmarks, with consistent improvements as training scale grows from 10M to 50M clips. Up to +2.1% over InternVid (FLT) baseline Consistent gains across 10M-50M clips 55M clips with synthetic video captions CLAP Audio-Language Benchmarks CLAP models trained on LAION-BVD achieve competitive performance against other large-scale uncurated audio datasets, leveraging rich in-the-wild soundscapes extracted directly from video. Competitive with uncurated audio datasets Audio-text pairs from diverse video Good scaling trends on when increasing model and data scale CLIP Image-Text Benchmarks Frame-based CLIP models achieve strong image-text retrieval performance on standard benchmarks. Video frames exhibit a visual distribution distinct from typical web corpora, complementing existing image pre-training sources. Strong retrieval on standard benchmarks 300M frames with unique visual distribution Complements standard web image datasets Responsible Use Ethics & Release Statement LAION-BVD is released to support open and reproducible multimodal research at scale. Large-scale video datasets and the models trained on them are increasingly concentrated within a small number of predominantly proprietary technology companies, limiting independent scientific investigation and reproducibility. By providing an open resource for academic research, we aim to broaden access to multimodal training data and enable more transparent evaluation of large-scale video, audio, and image models. LAION-BVD is released exclusively for research purposes and not for commercial use . The dataset is intended to support scientific research, reproducibility, safety analysis, and the study of multimodal foundation models and related systems. We encourage users to respect the rights and copyright of content creators and to use the dataset responsibly and in accordance with applicable laws and platform terms. Like other large-scale web datasets, LAION-BVD may contain biases, stereotypes, and uneven representation across languages, regions, and topics. Models trained on this data may inherit such biases. Researchers using the dataset should be aware of these limitations and, where relevant, evaluate and report them alongside model capabilities. Reference Citation BibTeX Copy @misc{laionbvd2026, title={LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training}, author={Andreas Hochlehnert and Marianna Nezhurina and Mehdi Cherti and Andrej Radonjic and Thaddäus Wiedemer and Christoph Schuhmann and Romain Beaumont and Wieland Brendel and Bernhard Schölkopf and A. Sophia Koepke and Jenia Jitsev and Matthias Bethge}, year={2026}, eprint={2608.24845}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2608.24845}, }