Civic Tech Field Guide
Donate

https://commoncrawl.org/

Common Crawl

Common Crawl maintains a free, open repository of web crawl data that can be used by anyone.

Active
Since 2007
Open source: Yes

Over 250 billion pages spanning 17 years. Free and open corpus since 2007. Cited in over 10,000 research papers. 3–5 billion new pages added each month.

Project type
Resource
Founded
2007
Language(s)
English
Added
2024-04-29
Last modified
2026-07-12T16:44:50.000Z

Additional details

Number of integrations
0
Geographic focus
Global

More in Civic data