Common Crawl
Also called Common Crawl
A free, public archive of web-crawl data that underpins a large share of AI training corpora.
Think of it like
A public reservoir everyone draws from — you still have to filter the water before you drink it.
Example
Many open datasets start from Common Crawl snapshots, then apply their own cleaning to produce a usable subset.
How it actually works
Common Crawl is a nonprofit that has crawled the web for years and releases petabyte-scale snapshots for free. It is the raw feedstock behind well-known derived datasets, which apply heavy filtering to it. Because so many models trace back to it, its coverage and gaps — languages, domains, time periods — quietly shape the field. It is raw material, not a ready-to-use dataset.
For product teams
A shared starting point that makes strong base data accessible, but never usable straight out of the box.
For engineers
Public multi-petabyte web-crawl snapshots used as the base feedstock for derived pretraining datasets.
Related
- Web crawl — The kind of data it provides.
- Pretraining corpus — What it ultimately feeds.
Read anything AI without the jargon
Look up any term in plain English, or save terms as you read with the free Chrome extension.
Open DecoderAdd to Chrome