Prompt
How do I crawl JavaScript-heavy sites at scale?
Latest observation
Crawling JavaScript-heavy sites at scale is mostly an engineering problem of rendering efficiently, extracting reliably, and controlling cost/latency.
1) Use the right crawler architecture
A good scalable setup is usually:
- Queue-based frontier: URLs go into a queue; workers pull tasks.
- Two-stage crawl:
- Fast fetch: try plain HTTP first.
- Render only when needed: use a browser engine for pages that require JS.
- Worker pool:
- HTTP workers for cheap pages
- Headless browser workers for JS pages
- Deduplication:
- Normalize URLs
- Track content hashes
- Avoid revisiting identical states
This keeps browser usage limited to pages that truly need it.
2) Detect when JS rendering is necessary
Don’t render everything. Common signals a page needs JS:
- Empty or sparse HTML in initial response
- Content loaded via XHR/fetch after load
- Important data missing from server HTML
- SPA frameworks like React/Vue/Angular
- Links or pagination built dynamically
Heuristics:
- Compare initial HTML to rendered DOM.
- Search for “app shell” patterns.
- Check network calls from a browser session and see if data comes from an API.
3) Prefer API extraction when possible
For many JS sites, the browser is just a UI layer over JSON APIs.
Best approach:
- Open the site once in DevTools/Playwright
- Inspect network requests
- Identify backend endpoints returning JSON
- Crawl those APIs directly
This is usually:
- Faster
- More reliable
- Cheaper
- Easier to scale
4) Use headless browsers efficiently
For pages that require rendering, use tools like:
- Playwright (recommended)
- Puppeteer
- Selenium (works, but generally less efficient for large-scale crawling)
Scaling tips:
- Reuse browser instances
- Open many pages in one browser process if stable
- Use browser contexts for isolation instead of new processes
- Block unnecessary resources:
- images
- fonts
- video
- analytics
- ads
- Set timeouts aggressively
- Wait for the right event, not arbitrary sleeps
Good wait strategy:
networkidlecan help, but some sites never go idle- Prefer waiting for a specific selector or API response
5) Render only until the data you need is available
If you need page title, product info, or article text:
- Wait for the main content selector
- Extract
- Stop the page early if possible
Avoid waiting for:
- all scripts
- all lazy-loaded assets
- animations
- endless polling
6) Handle infinite scroll and lazy loading
Many JS-heavy sites load content as you scroll.
Strategies:
- Scroll in increments and stop when:
- no new content appears
- API calls stop changing
- item count stabilizes
- Track page height and DOM item counts
- Use pagination APIs if available instead of scrolling
For scale, prefer direct API pagination over UI scrolling.
7) Build robust extraction
JavaScript sites change often. Make your scraper resilient:
- Use stable selectors, not brittle CSS paths
- Extract structured data from:
- JSON-LD
- embedded state objects
- hydration payloads
- API responses
- Keep parser logic separate from crawling logic
- Log extraction failures with HTML snapshots or response traces
8) Manage concurrency carefully
Headless browsers are expensive.
Practical tactics:
- Limit browser concurrency per machine
- Use autoscaling workers
- Split workloads:
- lightweight HTTP fetchers
- heavyweight browser renderers
- Use backpressure so browser queues don’t explode
Rule of thumb:
- HTTP crawling scales much more cheaply than full browser rendering.
- Use browsers as a last resort.
9) Be resilient to anti-bot measures
JS-heavy sites often also have anti-bot controls.
You may need:
- Session management
- Cookies / login handling
- Rate limiting
- Human-like pacing
- Retries with jitter
- Proxy rotation where appropriate
Also:
- Respect robots.txt and site terms
- Don’t overload sites
- Use crawl budgets per host
10) Store crawl state and observability
At scale, you need visibility into failure modes.
Track:
- Fetch success rate
- Render success rate
- Time to first byte
- Time to content ready
- Extraction success rate
- Per-host error rates
- Queue depth
Store:
- URL
- normalized URL
- fetch status
- final rendered URL
- response metadata
- content hash
- extracted entities
This helps you re-crawl efficiently and troubleshoot regressions.
11) A practical pattern that works well
A common scalable workflow:
- Enqueue seed URLs
- HTTP fetch page
- If data is present in HTML:
- extract and continue
- If not:
- check network/API clues
- if possible, switch to API crawling
- If still necessary:
- render with Playwright
- wait for target selector
- extract data
- Save results
- Discover more URLs and enqueue them
12) Recommended stack
A solid stack for JS-heavy crawling:
- Playwright for rendering
- Redis / Kafka / SQS for queues
- Postgres / Elasticsearch / object storage for results
- Docker for deployment
- Kubernetes / autoscaling workers for scale
- Proxy layer if needed
- Prometheus/Grafana for metrics
13) Cost-saving checklist
Before rendering a page, ask:
- Can I get the same data from an API?
- Is the content already in initial HTML?
- Can I extract from JSON-LD or embedded state?
- Can I block heavy resources?
- Can I stop rendering once the target element appears?
If yes to any, do that first.
14) Common mistake to avoid
The biggest mistake is treating every URL like it needs a full browser.
That leads to:
- high cost
- low throughput
- unstable crawls
- frequent timeouts
The winning strategy is: HTTP first, API second, browser last.
If you want, I can also give you:
- a Playwright-based crawling architecture,
- a sample worker design, or
- a code example for crawling JS-heavy pages at scale.