Prompt

What are the best web datasets for AI training?

Technology / Data Infrastructure5 observationsLast seen Sep 1, 2026

Latest observation

Sep 1, 2026Perplexity

For AI training—especially large language models (LLMs) and foundation models—the best web datasets are large‑scale, high‑quality crawls and curated corpora derived from the web. In 2026, the leading options are:

Core web‑scale pretraining datasets

1. Common Crawl

  • What it is: Free, open repository of raw web crawl data (hundreds of billions of pages, multi‑petabyte scale, 19+ years of history).
  • Best for: Base corpus for LLM pretraining; source for building custom filtered datasets.
  • Notes: Raw and noisy; almost always filtered/deduped before use (e.g., into C4, FineWeb, Dolma, DCLM, etc.).

2. C4 (Colossal Clean Crawled Corpus)

  • What it is: Filtered English (and multilingual) subset of Common Crawl, created for T5.
  • Size: ~750 GB, ~172B tokens (English).
  • Best for: Classic LLM pretraining baseline; widely used in research.
  • License: ODC‑By (check variant).

3. FineWeb / FineWeb‑Edu

  • What it is: High‑quality, heavily filtered Common Crawl–derived datasets (Hugging Face).
  • Size: FineWeb ~15T tokens; FineWeb‑Edu is a smaller, education‑focused subset with very high signal.
  • Best for: Compute‑efficient pretraining; often outperforms larger, noisier corpora per token.
  • License: Open (check specific release).

4. The Pile

  • What it is: 825 GB English corpus combining 22 diverse sources (web, books, code, academic papers, etc.).
  • Best for: General‑purpose LLM pretraining with broad topical coverage; strong on academic and technical tasks.
  • License: MIT.

5. Dolma / RedPajama / DCLM / Common Pile

  • What they are: Large, open pretraining corpora built from Common Crawl and other sources with various filtering/dedup strategies.
  • Sizes: Ranging from several trillion to tens of trillions of tokens depending on version.
  • Best for: Open‑source LLM pretraining with permissive licenses; alternatives to FineWeb/C4.

6. Nemotron‑CC (NVIDIA)

  • What it is: Synthetic‑enhanced Common Crawl–based corpus (e.g., Nemotron‑CC‑v2) with billions of samples and trillions of tokens.
  • Best for: General web pretraining for NVIDIA’s open models; available on Hugging Face.

Specialized web‑derived datasets

7. The Stack / Stack v2

  • What it is: Massive code corpus from GitHub and other sources.
  • Best for: Code LLMs and multi‑modal models with strong coding ability.

8. OpenWebMath / FineMath

  • What it is: Web‑scraped and filtered math‑heavy corpora.
  • Best for: Improving mathematical reasoning in LLMs.

9. Wikipedia + Stack Exchange + Books corpora

  • What they are: High‑quality, relatively clean text from specific domains.
  • Best for: Mixing into pretraining data to boost factual and structured knowledge.

How teams typically use these

  • Base pretraining: Start from Common Crawl → apply quality filters → get something like FineWeb, Dolma, or DCLM.
  • Efficient training: Use FineWeb‑Edu or similar high‑signal subsets to get more performance per token.
  • Diversity: Mix in The Pile, Wikipedia, Stack Exchange, books, and code (The Stack) for breadth and reasoning.
  • Domain tuning: Add specialized corpora (legal, math, code, multilingual) depending on target capabilities.

Where to get them

  • Hugging Face Datasets: Most of these (C4, FineWeb, FineWeb‑Edu, The Pile, Dolma, The Stack, etc.) are available with one‑line loaders.
  • Common Crawl: Direct downloads from commoncrawl.org (raw WARC/processed text).
  • GitHub / project pages: For specific versions of Dolma, RedPajama, DCLM, Nemotron, etc.

If you tell me your goal (e.g., “pretrain a 1–7B model,” “fine‑tune for reasoning,” “build a code model”), I can suggest a concrete mix of datasets and approximate sizes.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (5 observations for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.