Prompt
What are the best datasets for financial research?
Latest observation
The "best" dataset for financial research depends heavily on your specific area of study—whether you are doing academic asset pricing, macroeconomic forecasting, high-frequency quantitative trading, or financial NLP/LLM research.
The top financial datasets are categorized below by research domain, accessibility, and use case.
1. Academic & Institutional Standard (The "Gold Standard")
Most empirical finance PhD research and institutional asset management models rely on these datasets, usually accessed via the Wharton Research Data Services (WRDS) platform.
- CRSP (Center for Research in Security Prices): The standard for US historical equity data (prices, returns, volume, indices, and corporate actions) dating back to 1926.
- Compustat (S&P Global): The standard source for standardized corporate balance sheets, income statements, cash flows, and line-item financial fundamentals.
- OptionMetrics (IvyDB): Historical option prices, implied volatility surfaces, and Greeks for US and global equity options.
- TAQ (Trade and Quote): Millisecond-level trade and quote tick-data for all US exchange-listed equities, essential for market microstructure research.
2. Macroeconomic & Asset Pricing Factor Libraries (Free & Public)
Essential for macroeconomic forecasting, baseline asset pricing models, and policy research.
- FRED (Federal Reserve Economic Data): Hosted by the St. Louis Fed, offering over 800,000 macroeconomic time series (interest rates, CPI, GDP, employment, monetary policy) accessible via a free API.
- Kenneth French Data Library: Publicly available benchmarks for factor-based asset pricing (Fama-French 3-factor, 5-factor, Momentum, size, value, and industry portfolios).
- World Bank Open Data & IMF Data: Comprehensive international statistics on balance of payments, sovereign debt, exchange rates, and cross-country economic development.
3. Corporate Filings, Textual, and NLP Datasets
Crucial for corporate governance, forensic accounting, financial sentiment analysis, and fine-tuning Financial LLMs (e.g., FinBERT, FinGPT).
- SEC EDGAR Database: Free access to official regulatory filings (10-K, 10-Q, 8-K, Form 4 insider trading) via the SEC API or tools like
sec-edgar-downloader. - FNSPID (Financial News and Stock Price Integration Dataset): A massive multimodal dataset pairing over 15 million financial headlines with stock price movements, designed for transformer and LLM training.
- Financial PhraseBank & FiQA: Benchmark datasets on Hugging Face specifically annotated for financial sentiment analysis and financial question-answering.
4. Market Microstructure & High-Frequency Order Books
Designed for deep learning, reinforcement learning, and algorithmic trading order book dynamics.
- LOBSTER: Reconstructs limit order book (LOB) data up to 50 levels of depth using NASDAQ TotalView-ITCH data, widely used in academic high-frequency trading papers.
- FI2010: The classic benchmark dataset for predicting mid-price movements from Limit Order Book depth via machine learning.
5. Quant Trading & Developer APIs (Free/Freemium Market Data)
Ideal for backtesting, personal quant projects, and live trading systems.
- Nasdaq Data Link (formerly Quandl): A major repository for historical market data, alternative data, commodity futures, and economic datasets.
- Polygon.io & Databento: Modern APIs providing real-time and granular historical market data (stocks, options, forex, crypto) with developer-friendly pricing and Python/Rust SDKs.
- Yahoo Finance (
yfinancewrapper): Free, easily accessible daily price history for prototyping and educational models. - Alpha Vantage / Financial Modeling Prep (FMP): APIs providing technical indicators, stock fundamentals, and live ticker updates.
6. Risk, Credit Scoring, and Alternative Data
Used for machine learning risk modeling, credit scoring, and non-traditional signals.
- Lending Club Dataset (Kaggle): Thousands of individual loan applications, repayment histories, and default records, commonly used for default prediction and credit risk modeling.
- GDELT & RavenPack: Global event news monitoring and automated sentiment tagging for macro-event-driven strategies.
- PitchBook & Crunchbase: The primary platforms for analyzing startup valuations, venture capital funding, and private equity deals.
Summary Recommendation Matrix
| Research Goal | Primary Datasets to Use |
|---|---|
| Academic Asset Pricing | CRSP, Compustat, Kenneth French Library |
| Macro & Policy Analysis | FRED, World Bank, IMF, ECB Statistical Data |
| Quant Trading & Backtesting | Polygon.io, Databento, Nasdaq Data Link |
| Financial LLMs & Text Mining | SEC EDGAR, FNSPID, Financial PhraseBank |
| Credit Scoring & Fraud ML | Lending Club, SEC Financial Statement Data, Kaggle Fraud sets |
Brands
Citations
- https://unidata.pro/blog/best-financial-ml-datasets/
- https://odsc.medium.com/best-financial-datasets-for-ai-data-science-in-2025-b11df09a22aa
- https://representai.co.uk/building-tools/top-financial-datasets-for-machine-learning-quantitative-finance-and-economic-research-2026-guide/
- https://imerit.ai/resources/blog/20-best-finance-economic-datasets-for-machine-learning-all-pbm/
- https://jdsemrau.substack.com/p/financial-datasets-a-comparative
- github.io
- https://www.reddit.com/r/quant/comments/1ae3pq3/interesting_proprietary_financial_databases_to/
Observation history
See 4 earlier observations
Create a free Obsurfable account to unlock historical responses, model comparisons, and deeper brand intelligence.
How did Obsurfable measure this prompt?
Obsurfable records AI answers to buyer-style prompts in its research corpus (5 observations for this page). Metrics are distributions over observations, not a single static ranking.
Which AI systems does Obsurfable collect answers from?
OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, and Meta AI.