Prompt
How do AI companies collect training data?
Latest observation
AI companies collect training data through a mix of large-scale web crawling, licensed feeds, curated open datasets, human-generated data, and synthetic data. The exact mix depends on the model type (foundation LLM, domain-specific model, vision, speech, etc.) and how much emphasis they place on quality, diversity, and legal compliance.
1. Large-scale web crawls
This is the backbone for most foundation models.
Open web crawls
-
Common Crawl
- Nonprofit that continuously crawls the public web and releases raw and processed data.
- Used by almost all major labs as a base layer.
- Contains petabytes of text, HTML, and metadata.
-
Lab-operated crawlers
- Companies run their own crawlers (e.g., GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, Meta-ExternalAgent, OAI-SearchBot).
- These crawl the public web under specific user-agents and policies.
- Often focused on:
- Fresh content not yet in Common Crawl.
- Specific domains or languages.
- Content that opts in to AI training.
Post-crawl processing
Raw crawl data is heavily filtered before training:
- Quality filtering: Remove spam, gibberish, low-quality pages.
- Deduplication: MinHash, LSH, and other techniques to remove near-duplicates.
- Language identification: Separate by language for multilingual models.
- PII and sensitive content removal: Strip or mask personal data where possible.
- License and policy filtering: Exclude sources that explicitly opt out or have restrictive terms.
2. Curated open datasets
These are pre-cleaned, research-grade datasets that complement raw crawls.
-
FineWeb, FineWeb-Edu, RefinedWeb, Dolma, RedPajama-V2
- Quality-filtered subsets of web data designed for LLM pretraining.
- Often used as “high-quality anchors” in the training mix.
-
The Stack v2, CodeParrot, other code corpora
- Large, permissively licensed code datasets for code-capable models.
-
Wikipedia, Project Gutenberg, Open Library, arXiv, PubMed Central
- Encyclopedic, book, and scientific text for knowledge and reasoning.
-
Multilingual corpora
- OSCAR, HPLT, FineWeb-2, MADLAD-400, etc.
- Used to ensure coverage across many languages.
These are often combined in carefully tuned ratios (e.g., X% web, Y% books, Z% code) to shape model behavior.
3. Licensed publisher and platform feeds
To improve quality, recency, and legal clarity, many labs license data directly.
News and media
- Deals with publishers like News Corp, Associated Press, Financial Times, Vox, Axel Springer, The Atlantic, etc.
- Data: articles, headlines, metadata, sometimes archives.
- Use: high-quality, fact-checked text; domain-specific knowledge.
Social and community platforms
- Agreements with platforms like Reddit, Stack Overflow, Quora, sometimes X/Twitter (via API or special deals).
- Data: posts, comments, Q&A threads, with varying levels of filtering and anonymization.
- Use: conversational data, domain expertise, real-world Q&A.
Image, audio, and video
- Licenses with stock media providers (Shutterstock, Getty, AP Images, etc.) and media companies.
- Data: images, captions, audio, video clips with metadata.
- Use: multimodal models (text+image, text+video, etc.).
Licensed data is typically:
- Higher quality and better documented.
- More legally defensible for commercial use.
- More expensive and sometimes limited in scale.
4. Code and developer data
Critical for coding assistants and code-capable LLMs.
-
GitHub and other code hosts
- Public repositories filtered by license (e.g., permissive licenses like MIT, Apache, BSD).
- Often processed into datasets like The Stack v2.
-
StackExchange, Stack Overflow
- Q&A threads with code snippets and explanations.
-
Internal code archives
- Some companies also use their own internal codebases (where they have rights) for domain adaptation.
Code data is usually:
- Deduplicated and filtered for license compliance.
- Balanced against natural language to avoid overfitting to code.
5. Human-generated and contractor data
This is the “human layer” that makes models helpful, safe, and aligned.
Instruction and SFT data
- Created by:
- In-house annotators.
- Contractors via platforms like Scale AI, Surge AI, Toloka, Outlier, Invisible, Mercor, etc.
- Data:
- Prompt–response pairs for specific tasks (summarization, coding, reasoning, customer support, etc.).
- Domain-specific instructions (legal, medical, finance, etc., with expert reviewers).
Preference and RLHF data
- Human raters compare multiple model outputs and choose preferred responses.
- Datasets like Anthropic HH-RLHF, UltraFeedback, SHP, plus proprietary internal data.
- Used for:
- Reward modeling.
- Direct Preference Optimization (DPO) and similar alignment techniques.
Red-teaming and safety data
- Human-generated adversarial prompts and failure cases.
- Used to train safety classifiers and improve robustness.
This human data is relatively small in volume but extremely high in impact.
6. Synthetic and model-generated data
Increasingly used to augment or specialize training.
-
Self-generated data
- Larger models generate:
- Rephrased instructions.
- Reasoning traces (chain-of-thought style).
- Q&A pairs from existing text.
- Larger models generate:
-
Distillation data
- Outputs from big models used to train smaller models.
-
Synthetic multimodal data
- Generated image–text pairs, dialogues, or code for specific tasks.
Benefits:
- Can target specific capabilities or rare scenarios.
- Helps when real data is scarce or sensitive.
Risks:
- Over-reliance can lead to “model collapse” or degraded diversity.
- Must be carefully filtered and balanced with real data.
7. User interaction data (with varying consent models)
Many AI companies also use product usage data to improve models.
-
Chat logs and conversations
- User interactions with chatbots (e.g., ChatGPT, Gemini, Copilot).
- Often used for:
- Fine-tuning.
- RLHF and safety improvements.
- Evaluating model behavior.
-
Telemetry and feedback
- Thumbs up/down, edits, reported issues, usage patterns.
Consent and privacy vary:
- Some companies default to using this data unless users opt out.
- Others require explicit opt-in, especially for enterprise or regulated contexts.
- Regulations like GDPR and the EU AI Act are pushing clearer consent and data-minimization practices.
8. Data curation pipelines
Collecting data is only the first step; the real work is in curation.
Typical pipeline stages:
-
Ingestion
- Pull from crawls, APIs, licenses, contractors, and synthetic generators.
-
Filtering
- Quality classifiers, language ID, spam detection, PII removal.
-
Deduplication
- Document-level and near-duplicate removal.
-
Balancing
- Adjust ratios of domains, languages, modalities, and tasks.
-
Documentation and provenance
- Track sources, licenses, and processing steps for each subset.
-
Evaluation
- Test subsets on benchmarks and internal evals before including them in training.
9. Legal and compliance considerations
Increasingly, data collection is shaped by regulation and policy:
-
Copyright and licensing
- Preference for permissively licensed or public-domain content.
- Growing number of explicit licensing deals with publishers and platforms.
-
Privacy and data protection
- GDPR, CCPA, and similar laws constrain use of personal data.
- EU AI Act and EDPB guidelines require:
- Legitimate interest assessments for scraping.
- Data minimization and accuracy checks.
- Mechanisms to honor opt-outs and robots.txt-like signals.
-
Opt-out mechanisms
- Some labs respect “no AI training” signals in robots.txt or via dedicated opt-out pages.
- Publishers and creators increasingly negotiate explicit terms.
Summary
In practice, a frontier AI company’s training data stack looks like:
- Open web crawls (Common Crawl + own crawlers) as the base.
- Curated open datasets for quality and diversity.
- Licensed publisher/platform feeds for high-quality, rights-cleared text and media.
- Code corpora from public repos and Q&A sites.
- Human-curated instruction and preference data from contractors and experts.
- Synthetic data to fill gaps and specialize capabilities.
- User interaction data (with varying consent) for fine-tuning and safety.
If you tell me what kind of AI system you’re interested in (e.g., general LLM, coding assistant, vertical-specific model, multimodal), I can outline a more specific data collection strategy and typical sources.
Brands