Decoder. plain-English AI glossary

Common Crawl

● Core

Also called Common Crawl

A free, public archive of web-crawl data that underpins a large share of AI training corpora.

Think of it like

A public reservoir everyone draws from — you still have to filter the water before you drink it.

Example

Many open datasets start from Common Crawl snapshots, then apply their own cleaning to produce a usable subset.

How it actually works

Common Crawl is a nonprofit that has crawled the web for years and releases petabyte-scale snapshots for free. It is the raw feedstock behind well-known derived datasets, which apply heavy filtering to it. Because so many models trace back to it, its coverage and gaps — languages, domains, time periods — quietly shape the field. It is raw material, not a ready-to-use dataset.

For product teams

A shared starting point that makes strong base data accessible, but never usable straight out of the box.

For engineers

Public multi-petabyte web-crawl snapshots used as the base feedstock for derived pretraining datasets.

Related

Read anything AI without the jargon

Look up any term in plain English, or save terms as you read with the free Chrome extension.

Open DecoderAdd to Chrome