Prompt
I'm building an ecommerce catalog ingestion workflow and need reliable site extraction
Latest observation
A reliable ecommerce catalog extraction workflow usually needs two layers:
- Discovery / crawling: find product URLs and category pages.
- Extraction / normalization: pull structured product data from each product page and map it to your catalog schema.
Here’s a practical approach that works well in production.
1) Prefer structured data first
Most ecommerce sites expose product info in:
- JSON-LD (
<script type="application/ld+json">) - Microdata / Open Graph
- Embedded app state like
__NEXT_DATA__,window.__PRELOADED_STATE__, etc.
For each page, try in this order:
- JSON-LD Product schema
- Embedded product state
- Fallback HTML selectors
- Headless browser rendering only when needed
This keeps extraction fast and resilient.
2) Build a site profile per domain
For each store/domain, store:
- Product URL patterns
- Category/listing page patterns
- Pagination style
- Variant behavior
- Presence of structured data
- Anti-bot behavior
- Locale/currency conventions
A simple “site adapter” or “profile” layer saves a lot of pain later.
3) Extraction fields to standardize
At minimum, normalize:
skugtin/mpnif availablebrandnamedescriptionimages[]pricecurrencyavailabilityvariants[](size/color/etc.)categories[]product_urlcanonical_urlbreadcrumbs[]
4) Make extraction idempotent
Use a stable product key:
- Prefer
sku/gtin - Otherwise derive from canonical URL + site ID
Track:
- raw payload
- parsed payload
- extraction confidence
- timestamp
- source URL
- parser version
This makes reprocessing and debugging much easier.
5) Use validation and scoring
Add rules like:
- price must be numeric and > 0
- product name length within bounds
- image URLs valid and non-placeholder
- availability from known enum
- currency from ISO-4217 set
Assign a confidence score and route low-confidence pages to review or fallback parsing.
6) Handle variants carefully
Variants are a common failure point. Decide whether your catalog stores:
- one parent product + variant children
- one row per purchasable variant
Look for:
- option dropdowns
- query params
- embedded variant arrays
- color swatches / size selectors
7) Crawl responsibly
Use:
- robots.txt checks where appropriate
- rate limiting
- exponential backoff
- retries with jitter
- caching
- deduplication
- user-agent transparency if required
8) Recommended workflow architecture
A solid pipeline looks like:
Seed URLs → Crawler → HTML fetcher → Structured-data parser → Variant resolver → Fallback extractor → Normalization → Validation → Deduplication → Storage / downstream sync
9) Libraries and tools
Depending on your stack:
- Python:
httpx,BeautifulSoup,lxml,selectolax,playwright - JS/TS:
playwright,cheerio,axios - Structured data: JSON-LD parsing utilities
- Queues: RabbitMQ, SQS, Kafka, Celery, BullMQ
10) A robust rule of thumb
If you need “reliable,” aim for:
- 80–90% extraction from structured data and page source
- 10–20% fallback via HTML / rendered DOM
- minimal dependence on brittle CSS selectors
If you want, I can help you design:
- a catalog ingestion architecture
- a product page extraction schema
- or a sample crawler/parser implementation in Python or TypeScript.