Prompt

I'm building a model and need historical dataset APIs

Technology / API Platforms1 observationLast seen Oct 8, 2026

Latest observation

Oct 8, 2026OpenAI APIWeb search: off

Sure — if you’re building a model and need historical dataset APIs, here are some of the best options, grouped by use case.

General market / finance data

  • Alpha Vantage — stocks, FX, crypto, technical indicators, some fundamentals.
  • Polygon.io — strong equities, options, forex, crypto historical market data.
  • IEX Cloud — US equities-focused historical prices and fundamentals.
  • Tiingo — end-of-day and intraday equities, crypto, and news.
  • Nasdaq Data Link (Quandl) — broad collection of economic, market, and alternative datasets.

Macroeconomic / public data

  • FRED API — Federal Reserve economic data, rates, CPI, unemployment, GDP, etc.
  • World Bank API — global development indicators.
  • OECD API — macroeconomic and policy data.
  • IMF Data API — international macro data.
  • Eurostat API — European statistics.

News / sentiment / events

  • GDELT — global news and event database, very useful for NLP / event modeling.
  • NewsAPI — news articles, more limited historical depth depending on plan.
  • Event Registry — news/event aggregation.

Crypto historical data

  • CoinGecko API — prices, market cap, volume, some historical coverage.
  • CoinMarketCap API — crypto market data with historical endpoints.
  • Kaiko — institutional-grade crypto historical market data.
  • CryptoCompare — OHLCV, exchange data, social/sentiment data.

Weather / environmental data

  • NOAA API — climate and weather history.
  • Open-Meteo Archive API — weather history.
  • NASA Earthdata APIs — satellite and earth observation data.

Alternative / web-scale

  • Common Crawl — massive web archive.
  • OpenAlex API — scholarly metadata and citation graph.
  • Wikipedia dumps / MediaWiki API — historical page edits/content.

Sports

  • Sportsdata.io
  • Sportradar
  • The Odds API — more odds-focused, historical coverage varies.

Tips for choosing

When selecting an API for model training, check:

  • Historical depth: how far back it goes
  • Granularity: daily, hourly, tick-level
  • Coverage: assets, regions, or domains
  • Rate limits and bulk export options
  • Licensing: training/redistribution restrictions
  • Adjustments: splits/dividends for equities, survivorship bias handling
  • Missing data handling: gaps, revisions, timezone consistency

If you want the best fit

Tell me:

  1. What kind of model you’re building
  2. What domain you need historical data for
  3. Desired frequency (tick / minute / daily / monthly)
  4. Free vs paid preference

…and I can recommend the best APIs and even help you design the data pipeline.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.