Prompt
We're building a product catalog ingestion process from public sites, what should we use?
Latest observation
Sep 24, 2026
OpenAI APIWeb search: off
For a product catalog ingestion process from public sites, the best choice depends on scale, site complexity, and how “public” the sites really are.
Good default stack
If you want something practical and robust, use:
- Crawler/orchestrator: Scrapy
- Dynamic pages / JS-heavy sites: Playwright
- HTML parsing: BeautifulSoup or lxml
- Data processing/normalization: Pandas or your ETL framework
- Storage: PostgreSQL for structured data, plus S3/GCS for raw HTML/screenshots
- Queue/scheduling: Airflow, Prefect, or a message queue like SQS/RabbitMQ
- Monitoring/retries: centralized logs + retry/backoff + dead-letter queue
When to use what
- Mostly static sites:
Scrapyalone is often enough and very efficient. - JavaScript-rendered sites:
AddPlaywrightfor rendering and interaction. - A few sites, low volume:
A simpler custom Python scraper may be enough. - Large-scale, many domains:
Use a distributed crawl architecture with queues, proxy management, and deduplication.
Important considerations
- Respect robots.txt and site terms
- Rate limit aggressively
- Use caching and change detection
- Normalize product schemas early
- Store raw source snapshots for debugging/auditing
- Expect anti-bot protections on some sites
- Have a fallback for selector drift when site layouts change
A solid architecture
- Seed URLs from category/product pages
- Fetch pages
- Render if needed
- Extract fields like title, SKU, price, availability, images, brand, category
- Normalize/enrich
- Deduplicate and compare changes
- Write to canonical product tables
- Persist raw page data for traceability
If you want the shortest recommendation
Use Scrapy + Playwright + PostgreSQL + S3, orchestrated with Prefect/Airflow.
If you want, I can also recommend:
- a small-team MVP stack
- an enterprise-grade architecture
- or a vendor/tool comparison for crawling platforms.