
Fable 5.1 reportedly decoded the Cyphral Distich in 44 minutes with no human help.
LAION’s Big Video Dataset is an open research dataset built from CommonCrawl video URLs, with 80 million downloaded videos, 10 million hours of runtime, 55 million auto-described clips, and 300 million still images.


LAION has released the Big Video Dataset, or BVD, described as one of the largest open video datasets for AI research. The dataset draws from 1.3 billion video URLs found in CommonCrawl, with 80 million videos downloaded for the release.
The collection totals 10 million hours of footage and includes 55 million clips with auto-generated video and audio descriptions. It also includes 300 million still images, giving researchers a large multimodal resource for video, audio, image, and text training.
According to the paper cited in the article, models trained on BVD outperformed comparable models trained on InternVid by up to 2.1 percentage points on common video-to-text benchmarks. That makes the dataset relevant not just for scale, but for measurable training performance.
The training setup connects video, audio, and text, helping models learn which visual content matches specific descriptions or sounds. For researchers, that can support work on video understanding, captioning, retrieval, and multimodal alignment.
The article reports that most videos in the dataset come from YouTube and that the majority are in English. BVD’s scale is built around clips, descriptions, and still images rather than raw video alone.
That structure matters because AI systems often need aligned examples: what is shown, what is heard, and how that content is described. BVD is positioned as a research resource for building and evaluating those connections at large scale.
LAION is releasing the dataset for research only, and the article notes that the organization asks users to respect the rights of original content creators. The dataset and code are described as freely available.
On the legal side, The Decoder reports that LAION can likely point to a 2024 Hamburg Regional Court ruling that allowed it to collect copyrighted content for non-commercial research. The key takeaway for users is clear: the dataset is aimed at research use, not unrestricted commercial deployment.

Fable 5.1 reportedly decoded the Cyphral Distich in 44 minutes with no human help.

Claude Code and Codex overestimate task times and rate their own work too highly, raising oversight concerns for long-running AI tasks.

A study argues AI time savings may push scientists toward more projects, not deeper work.

AI may boost efficiency now while weakening the talent pipeline professions need later.