Prompt
Where can I find high-quality datasets for machine learning?
Latest observation
High-quality machine learning datasets come from a mix of open repositories, research benchmarks, cloud catalogs, and specialized marketplaces. The best sources depend on your modality (text, vision, speech, tabular, multimodal) and whether you need free research data or commercial-grade, labeled datasets.
1. General-purpose ML repositories
These are the first places to check for standard, well-curated datasets.
Kaggle Datasets
- What: Community-driven platform with 100k+ datasets.
- Data: Tabular, text, images, audio, time series, geospatial, etc.
- Strengths:
- Large variety, often with notebooks and discussions.
- Good for prototyping, competitions, and learning.
- Caveats: Quality and licensing vary; always check provenance and terms.
UCI Machine Learning Repository
- What: Classic academic repository.
- Data: ~600+ datasets for classification, regression, clustering, etc.
- Strengths:
- Clean, well-documented, widely used in research and teaching.
- Good for benchmarking algorithms.
- Caveats: Smaller scale; less suited for modern deep learning at scale.
OpenML
- What: Collaborative platform for datasets, tasks, and experiments.
- Data: Thousands of datasets with standardized metadata.
- Strengths:
- Integrated with AutoML and experiment tracking.
- Good for reproducible research.
Papers With Code
- What: Links datasets to research papers and leaderboards.
- Data: Datasets used in SOTA papers across CV, NLP, RL, etc.
- Strengths:
- Easy to find benchmark datasets tied to specific tasks.
- Clear performance baselines.
2. NLP and LLM datasets
Hugging Face Datasets
- What: Largest hub for NLP and multimodal datasets.
- Data: Text, speech, vision, multimodal; pretraining, fine-tuning, evaluation.
- Strengths:
- Tight integration with
transformersanddatasetslibraries. - Rich metadata, versioning, and community contributions.
- Tight integration with
- Use cases: Pretraining, instruction tuning, RAG, evaluation.
Common Crawl / FineWeb / RedPajama / Dolma
- What: Large-scale web text corpora.
- Strengths:
- Foundation for LLM pretraining.
- FineWeb and RedPajama are quality-filtered and more permissively licensed.
Specialized NLP datasets
- SQuAD, Natural Questions, HotpotQA – QA.
- GLUE, SuperGLUE, MMLU, Big-Bench – evaluation benchmarks.
- WMT, OPUS – translation.
- CNN/Daily Mail, XSum – summarization.
3. Computer vision datasets
ImageNet
- What: 14M+ labeled images across 20k+ categories.
- Use: Pretraining and benchmarking for image classification and detection.
COCO (Common Objects in Context)
- What: Images with detection, segmentation, and caption annotations.
- Use: Object detection, instance segmentation, vision-language models.
Open Images, Objects365, Visual Genome
- What: Large-scale detection and classification datasets.
- Use: Training robust detection models.
Specialized CV datasets
- KITTI, nuScenes, Waymo Open Dataset – autonomous driving.
- FER2013, AffectNet – emotion recognition.
- Food-101, Stanford Cars, CUB-200 – fine-grained classification.
4. Speech and audio datasets
LibriSpeech
- What: ~1,000 hours of English read speech from audiobooks.
- Use: ASR pretraining and benchmarking.
Mozilla Common Voice
- What: Multilingual, crowdsourced speech with transcripts.
- Use: ASR, voice assistants, multilingual models.
AudioSet, VoxCeleb, TED-LIUM
- AudioSet: Broad audio event classification.
- VoxCeleb: Speaker recognition.
- TED-LIUM: ASR on TED talks.
5. Multimodal and vision-language datasets
LAION-5B, COYO-700M
- What: Billions of image–text pairs from the web.
- Use: Pretraining CLIP-like models, text-to-image models.
Conceptual Captions, MS-COCO Captions, Flickr30k
- What: Images with human-written captions.
- Use: Image captioning, vision-language models.
VQA v2, GQA, TextVQA
- What: Visual question answering datasets.
- Use: VQA models, reasoning over images.
6. Tabular and business datasets
UCI (Adult, Bank Marketing, Credit Card Fraud, etc.)
- Use: Classification, regression, churn, fraud detection.
Kaggle business/finance datasets
- Examples: E-commerce transactions, credit risk, retail sales, stock data.
- Use: Applied ML, prototyping, competitions.
Government and macro datasets
- FRED, World Bank, OECD, Eurostat, Census
- Use: Forecasting, econometrics, policy analysis.
7. Cloud and open data catalogs
These aggregate large, often high-quality datasets across domains.
AWS Open Data Registry, Google Public Datasets, Azure Open Datasets
- What: Curated large-scale datasets hosted on cloud platforms.
- Data: Satellite imagery, genomics, transport, finance, public records, etc.
- Strengths:
- High-quality, well-maintained, often petabyte-scale.
- Easy to query with cloud-native tools.
Google Dataset Search
- What: Search engine for datasets across repositories.
- Use: Discovering domain-specific datasets when you don’t know where they live.
“Awesome Public Datasets” collections
- Community-curated lists (e.g., on GitHub) covering many domains: finance, health, climate, social science, etc.
8. Commercial dataset marketplaces
If you need licensed, production-grade, or domain-specific data.
AWS Data Exchange, Snowflake Marketplace, Nasdaq Data Link
- What: Marketplaces with datasets from multiple vendors.
- Data: Financial, retail, geospatial, alternative data, etc.
- Strengths:
- Clear licensing, quality controls, easy integration with cloud warehouses.
Specialized providers
- Defined.ai, Scale AI, Labelbox Datasets, Oxylabs video datasets, Titan Network (video)
- Use: High-quality labeled data for vision, speech, multimodal, and custom needs.
How to choose
-
For learning and prototyping:
→ Start with Kaggle, UCI, OpenML, and Hugging Face. -
For research and benchmarks:
→ Use Papers With Code, GLUE/SuperGLUE/MMLU, ImageNet, COCO, LibriSpeech, etc. -
For large-scale pretraining (LLMs, vision, multimodal):
→ Combine Common Crawl/FineWeb/RedPajama, LAION, ImageNet, LibriSpeech, etc. -
For production or domain-specific applications:
→ Add cloud open data and commercial marketplaces (AWS Data Exchange, Snowflake, specialized vendors).
If you tell me your task (e.g., “fraud detection,” “image classification,” “ASR,” “LLM for customer support”), I can suggest a short list of specific datasets and where to get them.
Brands