Prompt
How do I create a reliable scraper for rapidly changing websites?
Latest observation
Creating a reliable scraper for a rapidly changing website is mostly about resilience, observability, and minimizing coupling to fragile page structure. Here’s a practical approach.
1) Prefer stable data sources over HTML scraping
Before scraping the rendered page, check for:
- Official APIs
- Hidden JSON endpoints used by the frontend
- GraphQL endpoints
- Network requests in browser dev tools
These are usually more stable than page markup and easier to parse.
2) Use selectors that are resilient
Avoid brittle selectors like long absolute paths or dynamic class names.
Better options:
- Semantic attributes like
data-* - Text anchors near the target element
- ARIA labels
- Stable IDs
- XPath/CSS selectors based on structure, not volatile styling
Example:
- Fragile:
.container > div:nth-child(3) > span.price - Better:
[data-testid="price"]orspan[aria-label="price"]
3) Build in multiple extraction strategies
Assume the site will change.
For each field, define fallback logic:
- Primary selector
- Secondary selector
- Extract from embedded JSON
- Extract from neighboring context
- Mark as missing and continue
This prevents one DOM change from breaking the whole scraper.
4) Separate fetching, parsing, and storage
Use a pipeline like:
- Fetcher: downloads HTML / JSON
- Parser: extracts fields
- Validator: checks completeness and format
- Writer: stores results
That makes debugging much easier when something breaks.
5) Add validation and anomaly detection
Detect when the site changes by checking:
- Missing fields
- Sudden drop in item count
- Unexpected HTML structure
- Changed content type
- Parsing errors
- Unusual response codes or redirects
Example alerts:
- “Parsed 0 products instead of usual 50”
- “Price field missing in 80% of records”
- “Response now contains login page”
6) Log raw snapshots for failed cases
When parsing fails, save:
- The raw HTML or JSON
- URL
- Timestamp
- Response headers
- Selector used
- Error details
This makes fixing the scraper much faster.
7) Use rate limiting and retries carefully
Rapid changes often come with anti-bot measures or unstable responses.
Use:
- Exponential backoff
- Retry on transient failures
- Randomized delay/jitter
- Concurrency limits
- Respect for robots.txt and site terms where applicable
Avoid aggressive retries that amplify blocking.
8) Handle dynamic content explicitly
If content is loaded by JavaScript:
- Use a browser automation tool only when necessary
- Wait for specific elements or network calls, not arbitrary sleep
- Prefer API calls discovered from the page over full browser rendering when possible
Tools:
- Playwright
- Selenium
- Puppeteer
9) Make selectors and field mappings configurable
Don’t hardcode everything in code. Put selectors in config files so you can update them quickly without redeploying the whole scraper.
Example config:
{
"product_name": ["[data-testid='name']", "h1.product-title"],
"price": ["[data-testid='price']", ".price"]
}
10) Version your parsers
When the website changes:
- Keep old parser versions
- Compare outputs between versions
- Roll back quickly if a fix causes regressions
11) Test against saved HTML samples
Maintain a fixture set of:
- Normal pages
- Edge cases
- Missing fields
- Updated versions after site changes
Run tests whenever you update parsing logic.
12) Use monitoring and alerts
Track:
- Success rate
- Field coverage
- Parse latency
- Block rate
- Response size
- Item count trends
This helps you detect failure before downstream data consumers do.
A practical architecture
A robust scraper often looks like this:
- Discover URLs
- Fetch page or API response
- Detect page type/version
- Parse with layered fallback extractors
- Validate fields
- Store result
- Log failures with raw input
- Alert on anomalies
If the site changes very frequently
Consider these additional tactics:
- Scrape from backend endpoints instead of frontend HTML
- Use machine-readable structured data if available (
JSON-LD, Open Graph, schema.org) - Monitor the site’s network requests periodically
- Keep parsers small and modular
- Use diffing to compare new page versions with old ones
Important legal/ethical note
Make sure your scraping complies with:
- The site’s terms of service
- Robots.txt where relevant
- Copyright/privacy rules
- Rate limits and access policies
If you want, I can also show you:
- a Python scraper template with fallback selectors, or
- a Playwright-based scraper that handles dynamic pages reliably.