Prompt

What are the best datasets for retail analytics?

Technology / Data Infrastructure6 observationsLast seen Sep 21, 2026

Latest observation

Sep 21, 2026PerplexityWeb search: on

For retail analytics, the best datasets fall into two groups: public/open benchmarks (great for learning, prototyping, and method comparison) and commercial panels/POS feeds (what serious retailers and CPG brands use for demand, pricing, and assortment decisions). Below is a focused shortlist of high‑quality options in 2026. unidata

Open / free datasets (strong starting points)

Transactional & basket data

  • UCI Online Retail II

    • ~1M transaction lines from a UK online gift wholesaler (2009–2011); ideal for RFM segmentation, cohort analysis, and basic demand modeling. unidata
  • Instacart Market Basket Analysis

    • 3M+ grocery orders, 200K+ users, with reorder flags; excellent for association rules, next‑basket prediction, and personalization. cllimber
  • Olist Brazilian E‑Commerce (Kaggle)

    • ~100K orders with products, sellers, payments, deliveries, and reviews; great for end‑to‑end relational modeling (orders → logistics → reviews). cllimber
  • Synthetic multi‑store POS dataset (Mendeley)

    • 970K+ line items across 30 stores over 24 months, with a ground‑truth event calendar (Black Friday, holidays, etc.); useful for forecasting, event‑effect estimation, and teaching without privacy constraints. data.mendeley

Demand forecasting benchmarks

  • M5 / Walmart Sales Forecasting (Kaggle)

    • 3,049 products, 42,840 series; canonical benchmark for hierarchical demand forecasting and promotion effects. deepsdata
  • UCI Hierarchical Sales Data

    • Small, purpose‑built dataset for testing reconciliation methods (bottom‑up, top‑down, MinT) rather than large‑scale training. unidata

Reviews & product text

  • Amazon Reviews (McAuley Lab)

    • Up to hundreds of millions of reviews; widely used for recommendation systems, sentiment analysis, and LLM‑based product understanding. cllimber
  • Womens E‑Commerce Clothing Reviews (Kaggle)

    • Compact, clean dataset for sentiment + rating prediction and review‑driven merchandising insights. cllimber

Web & session behavior

  • Retailrocket Ecommerce Dataset

    • Event‑level behavior (views, carts, purchases) plus catalog; good starter for session‑based recommendation and conversion modeling. cllimber
  • Google GA4 Demo Dataset (BigQuery)

    • Realistic GA4 schema with sessions, events, and revenue; ideal for ecommerce funnel analysis and attribution practice. cllimber

Macro & market context

  • US Census Bureau ecommerce sales (and FRED series), Eurostat retail trade
    • Official aggregates for market sizing, seasonality checks, and macro overlays. cllimber

Commercial / industry datasets (production grade)

  • dunnhumby

    • Loyalty, pricing, and promotion analytics tied to activation; strong for measuring promo ROI and price elasticity at household/SKU level. worldmetrics
  • NIQ (NielsenIQ)

    • Benchmark‑grade retail measurement and category planning; widely used by CPGs for share, velocity, and distribution metrics. worldmetrics
  • 84.51° (Kroger’s data arm), SPINS (natural/wellness), DataWeave

    • Retailer‑specific or category‑specialized panels; good for assortment, shelf, and natural‑channel strategy. worldmetrics
  • Retailer POS / marketplace feeds (e.g., Walmart Luminate, Amazon Brand Analytics, marketplace reports)

    • First‑party sales, search, and conversion data for your own SKUs; essential if you sell on those platforms. obsessai

How to choose

  • Learning / portfolio projects: UCI Online Retail II, Instacart, Olist, M5, and the synthetic POS dataset give you realistic schemas and clear tasks (segmentation, forecasting, recommendation). unidata
  • Demand forecasting research: Start with M5 and the UCI Hierarchical Sales Data; add Google Trends (BigQuery) for search‑demand context. unidata
  • Personalization / recsys: Instacart (grocery), Olist (marketplace), Retailrocket (events), and Amazon Reviews (text + ratings). cllimber
  • Real business decisions (pricing, promos, assortment): Combine your own POS/ecommerce exports with a panel like dunnhumby or NIQ for category context and benchmark metrics. worldmetrics

If you tell me your focus (e.g., forecasting, pricing, CLV, assortment) and whether you need open or commercial data, I can narrow this to a specific dataset stack and example queries.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (6 observations for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.