Prompt
What are the best datasets for natural language processing?
Latest observation
Aug 31, 2026
Perplexity
The “best” NLP datasets depend on your task, but there’s a well-established core set that powers most modern models and benchmarks. Below are the most widely used, high-quality datasets by category.
1. General-purpose language modeling & pretraining
These are used to train or fine-tune large language models.
Common Crawl
- What: Massive web crawl (petabytes of raw HTML/text).
- Use: Foundation for many LLMs (e.g., variants of LLaMA, Falcon, etc.).
- Pros: Enormous scale, diverse domains.
- Cons: Noisy, requires heavy cleaning and filtering.
C4 (Colossal Clean Crawled Corpus)
- What: Cleaned version of Common Crawl used for T5.
- Use: Pretraining encoder–decoder and decoder-only models.
- Pros: Well-documented cleaning pipeline; widely benchmarked.
The Pile
- What: 800+ GB curated corpus from 22 sources (books, code, web, academic papers, etc.).
- Use: Pretraining and analysis of LLMs.
- Pros: Diverse, well-documented sources; good for research reproducibility.
RedPajama
- What: Open reproduction of LLaMA-style pretraining data (Common Crawl + code + books + etc.).
- Use: Training open LLMs with transparent data.
DOLMA, FineWeb, RefinedWeb
- What: Newer, heavily filtered web corpora designed for high-quality LLM training.
- Use: State-of-the-art open LLM pretraining.
2. Text classification & sentiment
IMDB Reviews
- Task: Binary sentiment classification (positive/negative).
- Size: 50k movie reviews.
- Use: Classic benchmark for sentiment analysis.
Amazon Reviews / Amazon Polarity
- Task: Sentiment / star-rating prediction.
- Use: Large-scale sentiment and rating prediction.
AG News
- Task: News topic classification (World, Sports, Business, Sci/Tech).
- Use: Short-text classification benchmark.
Yelp Review Polarity / Full
- Task: Sentiment (binary) and rating prediction from reviews.
- Use: Large-scale sentiment and rating tasks.
SST (Stanford Sentiment Treebank)
- Task: Fine-grained sentiment (very negative → very positive) with phrase-level annotations.
- Use: Sentiment at sentence and phrase level; tree-structured models.
3. Question answering (QA)
SQuAD (1.1 & 2.0)
- Task: Extractive QA on Wikipedia paragraphs.
- SQuAD 2.0 adds unanswerable questions.
- Use: Core benchmark for extractive QA models.
Natural Questions (Google)
- Task: Real Google search queries with long/short answers from Wikipedia.
- Use: Open-domain QA, retrieval-augmented models.
HotpotQA
- Task: Multi-hop QA requiring reasoning over multiple documents.
- Use: Multi-hop reasoning, explainability.
TriviaQA, WebQuestions, MS MARCO QA
- Task: Open-domain QA from web/search logs.
- Use: Large-scale QA, retrieval + reading.
DROP
- Task: QA requiring discrete reasoning (counting, sorting, arithmetic).
- Use: Numerical and logical reasoning in QA.
4. Summarization
CNN/DailyMail
- Task: News article → multi-sentence summary.
- Use: Abstractive summarization benchmark.
XSum
- Task: Highly abstractive summarization of BBC articles.
- Use: Challenging abstractive summarization.
PubMed / arXiv summarization datasets
- Task: Long scientific articles → abstract or summary.
- Use: Long-document summarization, domain-specific models.
Multi-News, SamSum, DialogSum
- Task: Multi-document or dialogue summarization.
- Use: Meeting notes, conversations, multi-source summarization.
5. Named Entity Recognition (NER) & structured extraction
CoNLL-2003
- Task: NER (Person, Location, Organization, Misc) in English and German.
- Use: Classic NER benchmark.
OntoNotes 5.0
- Task: Rich NER with many entity types.
- Use: Fine-grained entity recognition.
WikiANN, MultiCoNLL
- Task: Multilingual NER.
- Use: Cross-lingual NER, low-resource languages.
ACE, Few-NERD, etc.
- Task: Event extraction, fine-grained NER.
- Use: Information extraction pipelines.
6. Natural Language Inference (NLI) & semantic similarity
SNLI
- Task: Entailment/contradiction/neutral between sentence pairs.
- Use: Foundational NLI benchmark.
MultiNLI
- Task: Multi-genre NLI (spoken, written, fiction, etc.).
- Use: Robust NLI across domains.
QNLI, RTE, MNLI-mismatched
- Task: Variants of NLI used in GLUE/SuperGLUE.
STS Benchmark (STS-B), SICK-R
- Task: Semantic textual similarity (score 0–5).
- Use: Sentence embeddings, similarity models.
7. Coreference, parsing, and syntax
CoNLL-2012 (OntoNotes coreference)
- Task: Coreference resolution.
- Use: Mention clustering, anaphora resolution.
Penn Treebank, Universal Dependencies (UD)
- Task: Syntactic parsing (POS, dependencies).
- Use: Parsers, grammar-aware models, linguistic analysis.
8. Dialogue and conversational NLP
Cornell Movie Dialogs
- Task: Movie script dialogues.
- Use: Conversational modeling, chitchat.
DailyDialog, MultiWOZ, ConvAI2
- Task: Task-oriented and open-domain dialogue.
- Use: Dialogue state tracking, response generation.
PersonaChat, TopicalChat
- Task: Open-domain dialogue with personas/topics.
- Use: Long-term consistent chatbots.
9. Multilingual and cross-lingual datasets
XNLI
- Task: Multilingual NLI (15+ languages).
- Use: Cross-lingual transfer, multilingual encoders.
MLQA, XQuAD, TyDi QA
- Task: Multilingual QA.
OPUS, WMT parallel corpora
- Task: Machine translation.
- Use: MT models, cross-lingual pretraining.
mC4, CC100
- Task: Multilingual web corpora.
- Use: Pretraining multilingual LLMs (mT5, mBERT-style).
10. Code + natural language
CodeSearchNet
- Task: Code search, code–NL pairs across multiple languages.
CodeXGLUE, HumanEval, MBPP
- Task: Code generation, completion, translation.
- Use: Code LLMs, program synthesis.
The Stack, BigCode
- Task: Large-scale code corpora.
- Use: Pretraining code models (StarCoder, CodeLlama-style).
11. Benchmarks that combine many tasks
GLUE & SuperGLUE
- Collections of NLI, sentiment, QA, coreference, etc.
- Use: General NLP capability benchmarking.
BIG-bench, HELM, MMLU (mostly QA/knowledge)
- Use: Evaluating large models across many tasks and domains.
How to choose
- Pretraining LLMs: Common Crawl, C4, The Pile, RedPajama, DOLMA/FineWeb.
- Sentiment/classification: IMDB, SST, Amazon, Yelp, AG News.
- QA: SQuAD, Natural Questions, HotpotQA, DROP, MS MARCO.
- Summarization: CNN/DailyMail, XSum, PubMed/arXiv datasets.
- NER/IE: CoNLL-2003, OntoNotes, WikiANN.
- NLI/semantics: SNLI, MultiNLI, STS-B.
- Dialogue: MultiWOZ, DailyDialog, PersonaChat.
- Multilingual: XNLI, mC4, OPUS, UD.
- Code + NL: CodeSearchNet, The Stack, CodeXGLUE.
If you tell me your specific task (e.g., “fine-tune a small model for customer-support QA in English and Spanish”), I can suggest a minimal, high-impact set of datasets and how to combine them.
Brands