Prompt
How can I collect ecommerce datasets at scale?
Latest observation
Sep 1, 2026
Perplexity
To collect ecommerce datasets at scale, you need a pipeline that handles discovery, extraction, normalization, storage, and monitoring—while bypassing anti-bot systems and keeping data fresh. Below is a practical, production-oriented approach.
1. Define scope and data model
- Targets: Choose marketplaces/retailers (e.g., Amazon, Walmart, Shopify stores) and product sets (by category, brand, ASIN/SKU list).
- Fields: Product URL, SKU/ASIN/GTIN, title, brand, category path, variants, price (list/final), currency, discounts, availability, seller/offer info, ratings, review count, review text, timestamp, source.
- Refresh cadence: High-priority SKUs daily or hourly; long tail weekly.
2. Choose collection method
A. Use ecommerce scraping APIs (recommended for scale)
Providers like Bright Data, Oxylabs, ScrapingBee Ecommerce API, Apify, Olostep, Infatica offer specialized endpoints that:
- Handle proxies, CAPTCHAs, and anti-bot (Akamai, Cloudflare, PerimeterX).
- Render JavaScript.
- Return structured JSON or HTML for parsing.
Typical flow:
- Seed URLs: Start from category/search pages or known product URLs/ASINs.
- Two-phase scraping:
- Phase 1: Scrape category/search results to collect product URLs.
- Phase 2: Scrape each product URL for detailed data.
- Batching: Submit thousands of URLs in batches to the API; avoid per-request orchestration overhead.
B. Build custom scrapers (only if necessary)
If you must build in-house:
- Use headless browsers (Playwright, Puppeteer) behind a large residential/ISP proxy pool.
- Implement session reuse per proxy, randomized delays, and realistic headers.
- Add fallbacks (multiple proxy providers, different user-agents).
- Expect significant engineering effort to maintain as sites change defenses.
3. Extraction and normalization
- Parsing: Use XPath/CSS selectors, regex, or LLM-assisted extraction for tricky pages.
- Schema validation: Enforce required fields and types; reject malformed records early.
- Normalization:
- Standardize currencies, units, and price formats.
- Map categories to your taxonomy.
- Deduplicate by stable identifiers (ASIN/SKU/GTIN + marketplace + country).
- Change detection: Compare new values against last known; flag price drops, stock changes, new sellers, review spikes.
4. Storage and time-series design
- Store each scrape as a timestamped row in a relational or time-series table (or Parquet in a data lake).
- Key columns:
marketplace,country,product_id,variant_id,timestamp,price,currency,availability,seller_id,rating,review_count. - This enables:
- Historical price/availability trends.
- Diff-based alerts (e.g., competitor price < MAP).
- Aggregations by brand/category/seller.
5. Scheduling and monitoring
- Orchestration: Use cron, Airflow, Prefect, or cloud functions to run scrapes on schedule.
- Monitoring: Track:
- Volume (pages scraped, records stored).
- Freshness (max age per SKU).
- Failure rates and error types (blocked, parse error, timeout).
- Alerts: Notify when failure rate spikes, freshness SLA is breached, or key thresholds are crossed.
6. Compliance and risk management
- Review terms of service and robots.txt for each target; avoid authenticated or clearly restricted areas.
- Limit request rates to mimic human behavior; use geo-appropriate proxies and headers.
- Log provenance (source URL, timestamp, extraction method) for auditability.
Minimal example pattern (using an ecommerce scraping API)
# Pseudocode outline
API_KEY = "your_api_key"
BASE_URL = "https://api.provider.com/v1/ecommerce"
# 1. Seed: list of product URLs or ASINs
product_urls = load_seed_urls()
# 2. Batch scrape
batch_payload = {
"urls": product_urls,
"fields": ["title", "brand", "price", "currency", "availability", "seller", "rating", "reviews"],
"render_js": True
}
response = requests.post(
f"{BASE_URL}/scrape",
headers={"Authorization": f"Bearer {API_KEY}"},
json=batch_payload
)
records = response.json()["results"]
# 3. Normalize and validate
clean_records = []
for r in records:
if not is_valid_product(r):
continue
clean = normalize_product(r) # currency, price, categories, IDs
clean_records.append(clean)
# 4. Store as time-series
write_to_db(clean_records) # upsert by (marketplace, product_id, timestamp)
When to use managed vs in-house
- Use managed APIs if you need:
-
5 domains, >2k pages/day, geo-targeted content, or CAPTCHA/IP diversity.
- High reliability without maintaining scraper infrastructure.
-
- Consider in-house only for:
- ≤5 sites, ≤500 pages/day, stable targets, and you have dedicated engineering capacity.
If you share your target sites, scale (SKUs/pages per day), and stack (Python/Node, cloud), I can suggest a concrete toolset and a minimal pipeline blueprint.