White House Briefing on Open Data’s Role in Technology
October 8, 2024
We recently had the honor of briefing the White House Office of Science and Technology Policy (OSTP) on the role of The Common Crawl Foundation as critical infrastructure in the artificial intelligence ecosystem and how we can support U.S. federal efforts in advancing responsible AI use and research.
Read More...IAB Workshop on AI-CONTROL
September 30, 2024
Earlier this month, the Common Crawl Foundation had the privilege of participating in a groundbreaking workshop hosted by the Internet Architecture Board (IAB) in Washington DC.
Read More...Host- and Domain-Level Web Graphs July, August, and September 2024
September 26, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of July, August, and September 2024. The crawls used to generate the graphs were CC-MAIN-2024-30, CC-MAIN-2024-33, and CC-MAIN-2024-38.
Read More...September 2024 Crawl Archive Now Available
September 24, 2024
The crawl archive for September 2024 is now available. The data was crawled between September 7th and September 21st 2024, and contains 2.8 billion web pages (or 410 TiB of uncompressed content).
Read More...August/September 2024 Newsletter
September 10, 2024
We're pleased to announce our newsletter for August and September 2024.
Read More...Host- and Domain-Level Web Graphs June, July, and August 2024
August 21, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of June, July, August 2024. The crawls used to generate the graphs were CC-MAIN-2024-33, CC-MAIN-2024-30, and CC-MAIN-2024-26.
Read More...August 2024 Crawl Archive Now Available
August 18, 2024
The crawl archive for August 2024 is now available. The data was crawled between August 3rd and August 16th, and contains 2.3 billion web pages (or 327.4 TiB of uncompressed content).
Read More...The Increase of Common Crawl Citations in Academic Research
August 6, 2024
Common Crawl's impact on research has grown substantially since its beginning. Our crawls have become a vital resource for researchers in various fields, from natural language processing to red teaming.
Read More...Host- and Domain-Level Web Graphs May, June, and July 2024
July 30, 2024
We are pleased to announce a new release of host-level and domain-level Web Graphs based on the crawls of May, June, and July 2024.
Read More...July 2024 Crawl Archive Now Available
July 28, 2024
We are pleased to announce that the crawl archive for July 2024 is now available, containing 2.5 billion web pages, or 360 TiB of uncompressed content.
Read More...Common Crawl Statistics Now Available on Hugging Face
July 22, 2024
We're excited to announce that Common Crawl’s statistics are now available on Hugging Face!
Read More...The Environmental Impact of the Cloud - the Common Crawl Case Study
July 16, 2024
Looking at tools (Green Software) and methodologies to evaluate the environmental impact of the cloud (a nascent activity coined GreenOps).
Read More...Host- and Domain-Level Web Graphs April, May, and June 2024
June 30, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of April, May, June 2024. The crawls used to generate the graphs were CC-MAIN-2024-18, CC-MAIN-2024-22, and CC-MAIN-2024-26.
Read More...Dialog and Discovery at AI_dev 2024
June 28, 2024
This month members from the Common Crawl Foundation attended the AI_dev: Open Source GenAI & ML Summit in Paris, where discussions focused on AI advancements, ethics, and Open Source solutions.
Read More...June 2024 Crawl Archive Now Available
June 28, 2024
The crawl archive for June 2024 is now available. The data was crawled between June 12th and June 26th, and contains 2.7 billion web pages (or 382 TiB of uncompressed content). Page captures are from 52.7 million hosts or 41.4 million registered domains and include 945 million new URLs, not visited in any of our prior crawls.
Read More...May/June 2024 Newsletter
June 25, 2024
We’re pleased to share our newsletter for May/June 2024, featuring the latest updates, events, and highlights from our community.
Read More...Host- and Domain-Level Web Graphs February/March, April, and May 2024
June 4, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of February, April, and May 2024.
Read More...May 2024 Crawl Archive Now Available
June 3, 2024
The crawl archive for May 2024 is now available. The data was crawled between May 18th and May 31st, and contains 2.7 billion web pages (or 377 TiB of uncompressed content). This is our 100th crawl!
Read More...Host- and Domain-Level Web Graphs November/December 2023, February/March 2024, and April 2024
May 5, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of November, February, April 2024.
Read More...April 2024 Crawl Archive Now Available
May 1, 2024
We are pleased to announce that the crawl archive for April 2024 is now available. The data was crawled between April 12th and April 25th, and contains 2.7 billion web pages (or 386 TiB of uncompressed content). Page captures are from 47.24 million hosts or 37.65 million registered domains and include 0.98 billion new URLs not visited in any of our prior crawls.
Read More...March/April 2024 Newsletter
March 26, 2024
We're excited to share an update on some of our recent projects and initiatives in this newsletter!
Read More...Host- and Domain-Level Web Graphs September/October, November/December 2023 and February/March 2024
March 14, 2024
We are pleased to announce a new release of host-level and domain-level web graphs based on the crawls of September, November, February 2023-24.
Read More...February/March 2024 Crawl Archive Now Available
March 11, 2024
The crawl archive for February/March 2024 is now available. The data was crawled between February 20th and March 5th, and contains 3.16 billion web pages (or 424.7 TiB of uncompressed content).
Read More...Web Archiving File Formats Explained
March 1, 2024
In the ever–evolving landscape of digital archiving and data analysis, it is helpful to understand the various file formats used for web crawling. From the early ARC format to the more advanced WARC, and the specialised WET and WAT files, each plays an important role in the field of web archiving. In this post, we explain these formats, exploring their unique features, applications, and the enhancements they offer.
Read More...A Further Look Into the Prevalence of Various ML Opt–Out Protocols
February 22, 2024
This post details some experiments that we have done regarding Machine Learning Opt–Out protocols. We decided to investigate the prevalence of some of these protocols, by taking a deeper look at our WARC files, and finding which proportions of domains are using which opt–out protocols.
Read More...