Prompt

How do AI companies collect training data?

Technology / Data Infrastructure6 observationsLast seen Aug 31, 2026

Latest observation

Aug 31, 2026Perplexity

AI companies collect training data through a mix of large-scale web crawling, licensed feeds, curated open datasets, human-generated data, and synthetic data. The exact mix depends on the model type (foundation LLM, domain-specific model, vision, speech, etc.) and how much emphasis they place on quality, diversity, and legal compliance.


1. Large-scale web crawls

This is the backbone for most foundation models.

Open web crawls

  • Common Crawl

    • Nonprofit that continuously crawls the public web and releases raw and processed data.
    • Used by almost all major labs as a base layer.
    • Contains petabytes of text, HTML, and metadata.
  • Lab-operated crawlers

    • Companies run their own crawlers (e.g., GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, OAI-SearchBot).
    • These crawl the public web under specific user-agents and policies.
    • Often focused on:
      • Fresh content not yet in Common Crawl.
      • Specific domains or languages.
      • Content that opts in to AI training.

Post-crawl processing

Raw crawl data is heavily filtered before training:

  • Quality filtering: Remove spam, gibberish, low-quality pages.
  • Deduplication: MinHash, LSH, and other techniques to remove near-duplicates.
  • Language identification: Separate by language for multilingual models.
  • PII and sensitive content removal: Strip or mask personal data where possible.
  • License and policy filtering: Exclude sources that explicitly opt out or have restrictive terms.

2. Curated open datasets

These are pre-cleaned, research-grade datasets that complement raw crawls.

  • FineWeb, FineWeb-Edu, RefinedWeb, Dolma, RedPajama-V2

    • Quality-filtered subsets of web data designed for LLM pretraining.
    • Often used as “high-quality anchors” in the training mix.
  • The Stack v2, CodeParrot, other code corpora

    • Large, permissively licensed code datasets for code-capable models.
  • Wikipedia, Project Gutenberg, Open Library, arXiv, PubMed Central

    • Encyclopedic, book, and scientific text for knowledge and reasoning.
  • Multilingual corpora

    • OSCAR, HPLT, FineWeb-2, MADLAD-400, etc.
    • Used to ensure coverage across many languages.

These are often combined in carefully tuned ratios (e.g., X% web, Y% books, Z% code) to shape model behavior.


3. Licensed publisher and platform feeds

To improve quality, recency, and legal clarity, many labs license data directly.

News and media

  • Deals with publishers like News Corp, Associated Press, Financial Times, Vox, Axel Springer, The Atlantic, etc.
  • Data: articles, headlines, metadata, sometimes archives.
  • Use: high-quality, fact-checked text; domain-specific knowledge.

Social and community platforms

  • Agreements with platforms like Reddit, Stack Overflow, Quora, sometimes X/Twitter (via API or special deals).
  • Data: posts, comments, Q&A threads, with varying levels of filtering and anonymization.
  • Use: conversational data, domain expertise, real-world Q&A.

Image, audio, and video

  • Licenses with stock media providers (Shutterstock, Getty, AP Images, etc.) and media companies.
  • Data: images, captions, audio, video clips with metadata.
  • Use: multimodal models (text+image, text+video, etc.).

Licensed data is typically:

  • Higher quality and better documented.
  • More legally defensible for commercial use.
  • More expensive and sometimes limited in scale.

4. Code and developer data

Critical for coding assistants and code-capable LLMs.

  • GitHub and other code hosts

    • Public repositories filtered by license (e.g., permissive licenses like MIT, Apache, BSD).
    • Often processed into datasets like The Stack v2.
  • StackExchange, Stack Overflow

    • Q&A threads with code snippets and explanations.
  • Internal code archives

    • Some companies also use their own internal codebases (where they have rights) for domain adaptation.

Code data is usually:

  • Deduplicated and filtered for license compliance.
  • Balanced against natural language to avoid overfitting to code.

5. Human-generated and contractor data

This is the “human layer” that makes models helpful, safe, and aligned.

Instruction and SFT data

  • Created by:
    • In-house annotators.
    • Contractors via platforms like Scale AI, Surge AI, Toloka, Outlier, Invisible, Mercor, etc.
  • Data:
    • Prompt–response pairs for specific tasks (summarization, coding, reasoning, customer support, etc.).
    • Domain-specific instructions (legal, medical, finance, etc., with expert reviewers).

Preference and RLHF data

  • Human raters compare multiple model outputs and choose preferred responses.
  • Datasets like Anthropic HH-RLHF, UltraFeedback, SHP, plus proprietary internal data.
  • Used for:
    • Reward modeling.
    • Direct Preference Optimization (DPO) and similar alignment techniques.

Red-teaming and safety data

  • Human-generated adversarial prompts and failure cases.
  • Used to train safety classifiers and improve robustness.

This human data is relatively small in volume but extremely high in impact.


6. Synthetic and model-generated data

Increasingly used to augment or specialize training.

  • Self-generated data

    • Larger models generate:
      • Rephrased instructions.
      • Reasoning traces (chain-of-thought style).
      • Q&A pairs from existing text.
  • Distillation data

    • Outputs from big models used to train smaller models.
  • Synthetic multimodal data

    • Generated image–text pairs, dialogues, or code for specific tasks.

Benefits:

  • Can target specific capabilities or rare scenarios.
  • Helps when real data is scarce or sensitive.

Risks:

  • Over-reliance can lead to “model collapse” or degraded diversity.
  • Must be carefully filtered and balanced with real data.

7. User interaction data (with varying consent models)

Many AI companies also use product usage data to improve models.

  • Chat logs and conversations

    • User interactions with chatbots (e.g., ChatGPT, Gemini, Copilot).
    • Often used for:
      • Fine-tuning.
      • RLHF and safety improvements.
      • Evaluating model behavior.
  • Telemetry and feedback

    • Thumbs up/down, edits, reported issues, usage patterns.

Consent and privacy vary:

  • Some companies default to using this data unless users opt out.
  • Others require explicit opt-in, especially for enterprise or regulated contexts.
  • Regulations like GDPR and the EU AI Act are pushing clearer consent and data-minimization practices.

8. Data curation pipelines

Collecting data is only the first step; the real work is in curation.

Typical pipeline stages:

  1. Ingestion

    • Pull from crawls, APIs, licenses, contractors, and synthetic generators.
  2. Filtering

    • Quality classifiers, language ID, spam detection, PII removal.
  3. Deduplication

    • Document-level and near-duplicate removal.
  4. Balancing

    • Adjust ratios of domains, languages, modalities, and tasks.
  5. Documentation and provenance

    • Track sources, licenses, and processing steps for each subset.
  6. Evaluation

    • Test subsets on benchmarks and internal evals before including them in training.

9. Legal and compliance considerations

Increasingly, data collection is shaped by regulation and policy:

  • Copyright and licensing

    • Preference for permissively licensed or public-domain content.
    • Growing number of explicit licensing deals with publishers and platforms.
  • Privacy and data protection

    • GDPR, CCPA, and similar laws constrain use of personal data.
    • EU AI Act and EDPB guidelines require:
      • Legitimate interest assessments for scraping.
      • Data minimization and accuracy checks.
      • Mechanisms to honor opt-outs and robots.txt-like signals.
  • Opt-out mechanisms

    • Some labs respect “no AI training” signals in robots.txt or via dedicated opt-out pages.
    • Publishers and creators increasingly negotiate explicit terms.

Summary

In practice, a frontier AI company’s training data stack looks like:

  1. Open web crawls (Common Crawl + own crawlers) as the base.
  2. Curated open datasets for quality and diversity.
  3. Licensed publisher/platform feeds for high-quality, rights-cleared text and media.
  4. Code corpora from public repos and Q&A sites.
  5. Human-curated instruction and preference data from contractors and experts.
  6. Synthetic data to fill gaps and specialize capabilities.
  7. User interaction data (with varying consent) for fine-tuning and safety.

If you tell me what kind of AI system you’re interested in (e.g., general LLM, coding assistant, vertical-specific model, multimodal), I can outline a more specific data collection strategy and typical sources.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (6 observations for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.