Prompt
I'm building a pricing intelligence tool and need a scraping stack that won't break every week
Latest observation
Sep 24, 2026
OpenAI APIWeb search: off
For a pricing intelligence tool, the stack that “won’t break every week” is usually less about a single scraper and more about a resilient collection system with layered fallbacks.
Recommended stack
1) Crawling / orchestration
- Scrapy for fast, scalable crawling
- Playwright for JavaScript-heavy pages and anti-bot-tough sites
- Use scrapy-playwright if you want one pipeline with both static and dynamic handling
2) Extraction layer
- BeautifulSoup / lxml for static HTML parsing
- Prefer schema-based extractors that target structured data first:
- JSON-LD
- OpenGraph
- embedded product APIs
- network/XHR responses
- Only scrape rendered DOM when needed
3) Anti-breakage strategy
- Build multi-strategy extraction:
- official API if available
- hidden JSON / network calls
- HTML parsing
- rendered browser fallback
- Add selector fallbacks rather than one brittle CSS path
- Normalize everything into a product schema early
4) Proxy / network layer
- Rotating residential or ISP proxies for difficult targets
- Datacenter proxies for easier, high-volume sites
- Add:
- retry logic
- jittered backoff
- per-domain rate limits
- session persistence when needed
5) Change detection and monitoring
This is what prevents weekly breakage:
- Store HTML snapshots or extracted fields for sampled pages
- Add DOM diff alerts when page structure changes
- Set up field-level validation:
- price missing
- price format changed
- currency changed
- outlier detection
- Build a dashboard for:
- success rate by domain
- extraction coverage
- parse failure reasons
6) Queueing / scaling
- Redis + Celery or Kafka for job management
- PostgreSQL for canonical product/pricing data
- S3/GCS for raw HTML, screenshots, and response archives
7) Operational tooling
- Run on Docker
- Use Kubernetes only if scale truly requires it
- Add structured logs + tracing:
- request ID
- domain
- proxy used
- parser version
- extraction method
A practical “won’t break weekly” architecture
Scheduler
→ queues URLs
Fetcher
→ tries API / HTTP first
Browser fallback
→ Playwright only when needed
Parser
→ structured data first, DOM second
Validator
→ checks completeness and anomalies
Storage
→ raw + normalized + history
Monitor
→ alerts on failures and schema drift
What I would choose
If you want a reliable, maintainable setup:
- Scrapy for crawl management
- Playwright as a fallback, not default
- lxml for parsing
- PostgreSQL for structured output
- Redis for queues/caching
- S3 for raw captures
- Prometheus + Grafana for observability
- Sentry for exceptions
Most important design principle
Don’t build around selectors alone. Build around:
- data sources
- fallbacks
- monitoring
- versioned parsers
That’s what keeps a pricing intelligence tool stable over time.
If you want, I can sketch:
- a reference architecture diagram,
- a Python stack with package names, or
- a production-ready folder structure for this system.