Prompt
What are the best sources of data for AI models?
Latest observation
The best sources of data for AI models depend on the modality (text, code, vision, speech, multimodal) and the goal (foundation pretraining, domain adaptation, instruction tuning, evaluation). In practice, teams mix open, licensed, and proprietary sources to balance scale, quality, and legal clarity.
1. Text and language data (for LLMs and NLP)
Large-scale web corpora
-
Common Crawl
- The default base layer for most LLMs.
- Petabytes of raw web text; requires heavy filtering and deduplication.
-
FineWeb / FineWeb-Edu, RedPajama-V2, Dolma, RefinedWeb
- Quality-filtered, deduplicated web subsets designed for LLM pretraining.
- Often used as the “high-quality” portion of the mix.
Curated open text
-
Wikipedia (all languages)
- Encyclopedic, relatively clean, multilingual.
-
Books and long-form text
- Project Gutenberg, Open Library, other public-domain or licensed book corpora.
-
Scientific and academic text
- arXiv, PubMed Central, Semantic Scholar Open Research Corpus (S2ORC).
-
Q&A and forums
- Stack Exchange, Stack Overflow, Quora, Reddit (via APIs or licensed datasets).
Code
- The Stack v2, CodeParrot, CodeSearchNet
- Large, multi-language code corpora from public repos and Q&A.
- Essential for code-capable models.
Instruction and preference data
-
Hugging Face Datasets (e.g., OpenHermes, Dolly, OASST, LIMA, SmolTalk)
- Instruction-tuning and chat-style data.
-
Anthropic HH-RLHF, UltraFeedback, SHP
- Human preference data for alignment (RLHF/DPO).
2. Vision data (for image and video models)
Image classification and detection
-
ImageNet
- 14M+ labeled images; foundational for vision models.
-
COCO, Open Images, Objects365, Visual Genome
- Detection, segmentation, and multi-label classification.
Multimodal image–text
-
LAION-5B, COYO-700M
- Billions of image–text pairs from the web.
- Used for CLIP-like models and text-to-image.
-
Conceptual Captions, MS-COCO Captions, Flickr30k
- High-quality image–caption pairs.
Video
-
WebVid, InternVid, Kinetics, Something-Something, HowTo100M
- Video–text or action-labeled datasets for video understanding and generation.
-
ShareGPT4V, LLaVA-Instruct, ShareGemini, and similar instruction-tuned multimodal sets
- Used in recent vision-language models (e.g., VITA-1.5, LLaVA variants).
3. Speech and audio data
-
LibriSpeech
- ~1,000 hours of English read speech; standard ASR benchmark.
-
Mozilla Common Voice
- Multilingual, crowdsourced speech with transcripts.
-
AudioSet
- 2M+ labeled sound clips across 632 classes; standard for audio classification.
-
VoxCeleb, TED-LIUM, MuAViC
- Speaker recognition, ASR on talks, and audio-visual speech translation.
-
CapSpeech and similar captioned audio datasets
- Audio–caption pairs for TTS and style modeling.
4. Multimodal and instruction-tuned data
-
FineVision (24M multimodal samples)
- Purpose-built for vision-language model (VLM) training.
-
LLaVA-style datasets (LLaVA-150K, LLaVA-Mixture, LVIS-Instruct, etc.)
- Image QA, reasoning, and instruction-following for VLMs.
-
ShareGPT4V, ALLaVA-Caption, ShareGemini
- Image and video captioning, visual QA, and video-based reasoning.
-
ScienceQA, ChatQA, and other reasoning-focused multimodal sets
- For math, science, and multi-step reasoning over images.
5. Domain-specific data
Depending on the vertical, teams add specialized corpora:
- Finance: SEC EDGAR filings, earnings call transcripts, news archives, market data.
- Healthcare: MIMIC, PubMed, clinical notes (with strict governance), medical guidelines.
- Legal: Court opinions, statutes, regulations (e.g., CaseLaw Access, EUR-Lex).
- Customer support: Historical tickets, chat logs, knowledge base articles.
These are often a mix of public, licensed, and internal data.
6. Licensed and commercial data
To improve quality and legal clarity, many teams license data:
- News and media: Reuters, AP, FT, major publishers.
- Social and Q&A platforms: Reddit, Stack Overflow, Quora, sometimes X/Twitter.
- Image/video/audio: Shutterstock, Getty, AP Images, stock media libraries.
- Specialized providers:
- Scale AI, Labelbox, Encord, Luel, Taskmonk for custom, rights-cleared multimodal datasets and human annotations.
Licensed data is especially important for:
- High-quality text and media.
- Domains with strong IP or privacy constraints.
- Commercial products where legal risk must be minimized.
7. Synthetic and model-generated data
Increasingly used to:
- Generate instruction–response pairs for specific tasks.
- Create reasoning traces and chain-of-thought style examples.
- Augment rare scenarios or underrepresented domains.
- Distill larger models into smaller ones.
Common sources:
- Outputs from larger LLMs/VLMs.
- Synthetic image–text pairs, dialogues, or code.
Must be balanced with real data to avoid quality degradation.
8. Repositories and marketplaces to discover datasets
- Hugging Face Datasets – Largest hub for NLP, vision, speech, and multimodal data.
- Kaggle Datasets – Broad, community-driven collection.
- Papers With Code – Datasets tied to SOTA papers and benchmarks.
- AWS Open Data, Google Public Datasets, Azure Open Datasets – Large-scale, cloud-hosted datasets.
- GitHub “awesome-datasets” lists – Curated collections for LLMs, multimodal, domain-specific data.
- Commercial marketplaces: AWS Data Exchange, Snowflake Marketplace, Nasdaq Data Link, Datarade, Neudata.
How to choose
-
For general-purpose LLMs:
→ Mix of Common Crawl + FineWeb/RedPajama/Dolma + Wikipedia + books + code + instruction/preference data. -
For vision-language models:
→ LAION/COYO + COCO/Captions + FineVision + LLaVA/ShareGPT4V-style instruction data. -
For speech/ASR:
→ LibriSpeech + Common Voice + AudioSet + domain-specific audio. -
For domain-specific models (finance, health, legal, support):
→ Add domain corpora (filings, papers, laws, tickets) plus licensed data where needed.
If you tell me your target model type and domain, I can suggest a concrete shortlist of datasets and where to get them.
Brands