Prompt
What are the best web datasets for AI training?
Latest observation
Sep 1, 2026
Perplexity
For AI training—especially large language models (LLMs) and foundation models—the best web datasets are large‑scale, high‑quality crawls and curated corpora derived from the web. In 2026, the leading options are:
Core web‑scale pretraining datasets
1. Common Crawl
- What it is: Free, open repository of raw web crawl data (hundreds of billions of pages, multi‑petabyte scale, 19+ years of history).
- Best for: Base corpus for LLM pretraining; source for building custom filtered datasets.
- Notes: Raw and noisy; almost always filtered/deduped before use (e.g., into C4, FineWeb, Dolma, DCLM, etc.).
2. C4 (Colossal Clean Crawled Corpus)
- What it is: Filtered English (and multilingual) subset of Common Crawl, created for T5.
- Size: ~750 GB, ~172B tokens (English).
- Best for: Classic LLM pretraining baseline; widely used in research.
- License: ODC‑By (check variant).
3. FineWeb / FineWeb‑Edu
- What it is: High‑quality, heavily filtered Common Crawl–derived datasets (Hugging Face).
- Size: FineWeb ~15T tokens; FineWeb‑Edu is a smaller, education‑focused subset with very high signal.
- Best for: Compute‑efficient pretraining; often outperforms larger, noisier corpora per token.
- License: Open (check specific release).
4. The Pile
- What it is: 825 GB English corpus combining 22 diverse sources (web, books, code, academic papers, etc.).
- Best for: General‑purpose LLM pretraining with broad topical coverage; strong on academic and technical tasks.
- License: MIT.
5. Dolma / RedPajama / DCLM / Common Pile
- What they are: Large, open pretraining corpora built from Common Crawl and other sources with various filtering/dedup strategies.
- Sizes: Ranging from several trillion to tens of trillions of tokens depending on version.
- Best for: Open‑source LLM pretraining with permissive licenses; alternatives to FineWeb/C4.
6. Nemotron‑CC (NVIDIA)
- What it is: Synthetic‑enhanced Common Crawl–based corpus (e.g., Nemotron‑CC‑v2) with billions of samples and trillions of tokens.
- Best for: General web pretraining for NVIDIA’s open models; available on Hugging Face.
Specialized web‑derived datasets
7. The Stack / Stack v2
- What it is: Massive code corpus from GitHub and other sources.
- Best for: Code LLMs and multi‑modal models with strong coding ability.
8. OpenWebMath / FineMath
- What it is: Web‑scraped and filtered math‑heavy corpora.
- Best for: Improving mathematical reasoning in LLMs.
9. Wikipedia + Stack Exchange + Books corpora
- What they are: High‑quality, relatively clean text from specific domains.
- Best for: Mixing into pretraining data to boost factual and structured knowledge.
How teams typically use these
- Base pretraining: Start from Common Crawl → apply quality filters → get something like FineWeb, Dolma, or DCLM.
- Efficient training: Use FineWeb‑Edu or similar high‑signal subsets to get more performance per token.
- Diversity: Mix in The Pile, Wikipedia, Stack Exchange, books, and code (The Stack) for breadth and reasoning.
- Domain tuning: Add specialized corpora (legal, math, code, multilingual) depending on target capabilities.
Where to get them
- Hugging Face Datasets: Most of these (C4, FineWeb, FineWeb‑Edu, The Pile, Dolma, The Stack, etc.) are available with one‑line loaders.
- Common Crawl: Direct downloads from commoncrawl.org (raw WARC/processed text).
- GitHub / project pages: For specific versions of Dolma, RedPajama, DCLM, Nemotron, etc.
If you tell me your goal (e.g., “pretrain a 1–7B model,” “fine‑tune for reasoning,” “build a code model”), I can suggest a concrete mix of datasets and approximate sizes.