Prompt
How do I extract product listings from ecommerce sites into CSV or a warehouse?
Latest observation
Sep 24, 2026
OpenAI APIWeb search: off
To extract product listings from ecommerce sites into CSV or a warehouse, you typically use one of these approaches:
1) Best option: Use the site’s API
If the ecommerce site offers a product API, that’s usually the cleanest and most reliable method.
Typical flow
- Authenticate with API keys/OAuth
- Pull product data in pages
- Normalize fields like:
- product_id
- title
- price
- currency
- brand
- category
- stock / availability
- image URLs
- product URL
- variants
- Write to:
- CSV, or
- a database / warehouse like PostgreSQL, BigQuery, Snowflake, Redshift
Pros
- Stable
- Less likely to break
- Usually compliant
2) Scrape HTML pages when no API exists
If there’s no API, you can scrape product pages directly.
Typical tools
- Python:
requests,BeautifulSoup,lxml - JavaScript:
Playwright,Puppeteer - For simple sites:
requests + BeautifulSoup - For dynamic sites:
PlaywrightorSelenium
Workflow
- Get product listing pages
- Extract product links
- Visit each product page
- Parse HTML for fields
- Save results to CSV or load into a warehouse
Common fields to extract
- product name
- SKU / ID
- price
- sale price
- availability
- rating
- review count
- description
- category
- images
- variant info
- breadcrumbs
- URL
3) Use structured data embedded in pages
Many ecommerce pages include schema.org JSON-LD or microdata.
Look for:
<script type="application/ld+json">Product,Offer,AggregateRating
This is often easier than scraping visible HTML because product info is already structured.
4) Use crawling frameworks for scale
If you need to scrape many sites or large catalogs, consider:
- Scrapy for Python crawling at scale
- Playwright for JavaScript-heavy sites
- Airflow / Dagster / Prefect for orchestration
- dbt for transforming loaded warehouse data
5) Load into CSV or warehouse
CSV
Use Python’s csv module or pandas:
import pandas as pd
df = pd.DataFrame(products) # products = list of dicts
df.to_csv("products.csv", index=False)
Warehouse
Common patterns:
- Save raw scraped data to cloud storage first
- Load into warehouse tables
- Transform into analytics tables
Example destinations
- BigQuery
- Snowflake
- Redshift
- Postgres
Typical pipeline
- Scrape products
- Store raw JSON/CSV in S3/GCS/Azure Blob
- Load into staging table
- Clean/transform into final schema
Example: simple product scraper in Python
Here’s a very basic example:
import requests
from bs4 import BeautifulSoup
import pandas as pd
url = "https://example.com/products"
html = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).text
soup = BeautifulSoup(html, "html.parser")
products = []
for item in soup.select(".product-card"):
title = item.select_one(".product-title").get_text(strip=True)
price = item.select_one(".product-price").get_text(strip=True)
link = item.select_one("a")["href"]
products.append({
"title": title,
"price": price,
"url": link
})
pd.DataFrame(products).to_csv("products.csv", index=False)
If the site is JavaScript-rendered
Use Playwright:
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com/products")
page.wait_for_load_state("networkidle")
cards = page.locator(".product-card")
products = []
for i in range(cards.count()):
card = cards.nth(i)
products.append({
"title": card.locator(".product-title").inner_text(),
"price": card.locator(".product-price").inner_text(),
})
browser.close()
Important considerations
- Check the site’s Terms of Service and robots.txt
- Respect rate limits
- Avoid login-protected/private data unless you have permission
- Use retries, throttling, and deduplication
- Expect HTML structures to change
Recommended architecture for production
If you want something robust:
- Crawler: Scrapy or Playwright
- Scheduler: Airflow/Prefect
- Raw storage: S3/GCS
- Warehouse: BigQuery/Snowflake/Postgres
- Transformations: dbt
- Monitoring: logs + alerts for parsing failures
If you want, I can give you:
- a ready-to-run Python scraper,
- a pipeline to BigQuery/Snowflake/Postgres, or
- a site-specific extraction approach if you share an example ecommerce page.