Prompt
How do I collect public web data from multiple sites reliably?
Latest observation
To collect public web data from multiple sites reliably, think in terms of a pipeline: discover → fetch → parse → validate → store → monitor. The main challenge isn’t just getting the data once; it’s making the process robust against site changes, rate limits, and partial failures.
1) Start with the right source
Prefer, in this order:
- Official APIs
Most reliable, least brittle, usually easiest to maintain. - Feeds / structured exports
RSS/Atom, CSV exports, sitemaps, downloadable datasets. - HTML scraping
Use when no API exists. - Browser automation
Only when pages are heavily JavaScript-driven or require interaction.
If you can avoid scraping, do.
2) Respect site policies and legal constraints
Before collecting:
- Check robots.txt
- Review Terms of Service
- Don’t bypass authentication, paywalls, CAPTCHAs, or access controls
- Rate-limit requests to avoid abuse
- Identify your crawler with a clear User-Agent and contact info if appropriate
If the data is public, that doesn’t automatically mean unrestricted reuse or collection.
3) Design for reliability
Fetching
Use:
- Timeouts
- Retries with exponential backoff
- Connection pooling
- Caching
- Concurrency limits per domain
Key idea: do not hammer all sites equally. Each site should have its own fetch policy.
Example retry strategy
- Retry 3 times
- Backoff: 1s, 2s, 4s
- Retry only on transient failures:
- 429
- 500, 502, 503, 504
- network timeouts
Do not blindly retry parsing errors.
4) Separate crawling from parsing
Build two layers:
- Crawler/fetcher: gets raw HTML/JSON
- Parser/extractor: turns raw content into structured data
Why:
- You can re-parse old content when site markup changes
- You can debug extraction independently
- You can test parser changes on saved pages
Store raw responses when possible.
5) Normalize the data model
Different sites will name the same thing differently.
Create a canonical schema, for example:
source_sitesource_urlscraped_attitlepricecurrencyposted_atauthorbodyexternal_id
Normalize:
- dates to UTC
- currencies to ISO codes
- URLs to canonical form
- text to consistent encoding
- numbers to a standard numeric type
6) Make extraction resilient to markup changes
HTML changes constantly. Use strategies like:
- Prefer stable attributes (
data-*, itemprop, JSON-LD, schema.org) - Extract from structured embedded data when available
- Use multiple selectors with fallbacks
- Avoid brittle full-path CSS selectors when possible
Good idea:
- First try JSON-LD / embedded JSON
- Then semantic HTML
- Then fallback selectors
7) Handle pagination, deduplication, and incremental updates
Pagination
Support:
- page numbers
- cursor-based pagination
- “load more” endpoints
Deduplication
Deduplicate by:
- canonical URL
- source-specific ID
- content hash
- combination of fields
Incremental collection
Only fetch:
- new pages
- updated records
- changed timestamps
This saves bandwidth and reduces load on sites.
8) Use concurrency carefully
For multiple sites:
- Use per-domain rate limits
- Use a worker queue
- Keep a global concurrency cap
- Prioritize smaller or slower domains lower
Example policy:
- max 2 requests/sec per domain
- max 20 concurrent requests overall
- jitter in delays to avoid burst patterns
9) Monitor and alert
Reliability requires observability.
Track:
- request success rate
- HTTP status codes
- latency
- parsing success rate
- number of extracted records
- dedup rate
- schema validation failures
Alert on:
- sudden drop in records
- spike in 403/429/5xx
- parser exceptions
- unexpected field null rates
A site changing its HTML often looks like a “successful fetch” but “zero extracted items.”
10) Validate data quality
Before storing or using the data:
- enforce schema validation
- check required fields
- validate dates, URLs, numbers
- detect obviously bad values
- flag outliers
Examples:
- price is negative
- date is in the future by years
- title is empty but page is otherwise valid
11) Store raw and structured data
A robust pattern is:
- Raw layer: HTML/JSON as fetched
- Processed layer: structured records
- Curated layer: cleaned and deduplicated data
This helps with:
- debugging
- reprocessing after parser changes
- audits and reproducibility
12) Prefer idempotent jobs
Your crawler should be safe to rerun.
That means:
- re-fetching the same URL doesn’t duplicate data
- storage uses upserts or deterministic keys
- job state is checkpointed
- failures can resume from the last good point
13) Recommended architecture
A practical setup:
- Scheduler: decides what to crawl
- Queue: pending URLs/tasks
- Fetcher: downloads pages
- Parser: extracts structured data
- Validator: checks schema and quality
- Database / object storage: stores raw + parsed data
- Monitoring: logs and metrics
- Config per site: rate limits, selectors, pagination rules, auth if allowed
14) Tools you can use
Depending on your stack:
Python
requests/httpxfor HTTPBeautifulSouporlxmlfor parsingscrapyfor full crawling frameworkplaywrightfor JS-heavy pagespydanticfor validation
JavaScript/TypeScript
axios/fetchcheerioplaywright- queue libraries like BullMQ
Infrastructure
- PostgreSQL for structured storage
- S3/object storage for raw pages
- Redis for queues/rate limiting
- Prometheus/Grafana for monitoring
15) A simple reliability checklist
For each site, ask:
- Is there an API?
- What is the rate limit?
- What is the canonical identifier?
- How do I detect updates?
- What is the fallback if markup changes?
- How will I validate extracted data?
- How will I know if the site breaks tomorrow?
A minimal practical approach
If you’re starting from scratch:
- Make a list of target sites and their access methods
- Build one connector per site
- Store raw response bodies
- Parse into a shared schema
- Add retries/backoff and per-domain rate limits
- Add validation and deduplication
- Monitor record counts and error rates
- Revisit broken parsers quickly when pages change
If you want, I can also provide:
- a Python scraping architecture template
- a Scrapy project structure
- a per-domain rate limiting example
- or a sample schema for multi-site public data collection