Prompt
I'm building a crawler that needs to render JavaScript pages at scale
Latest observation
If you’re crawling pages that need JavaScript rendering at scale, the main design choice is how much browser you really need per page. A full browser is expensive, so the usual goal is to render only when necessary, and keep everything else lightweight.
Practical architecture
1) Split URLs into tiers
Use a cheap first pass:
- fetch raw HTML with an HTTP client
- inspect for signs that JS rendering is needed:
- minimal HTML content
- content loaded from API/XHR calls
noscriptfallback only- app shell / empty root divs
- SPA frameworks (
__NEXT_DATA__,window.__INITIAL_STATE__, etc.)
Then route:
- Tier A: plain HTTP fetch
- Tier B: browser rendering only for pages that need it
This saves a lot of cost.
2) Use a browser pool, not one browser per request
At scale, launching browsers repeatedly is too slow.
Common setup:
- one browser process per worker node
- multiple isolated pages/tabs/contexts per browser
- reuse browser instances
- cap concurrency carefully
Good tools:
- Playwright: modern, reliable, great for automation
- Puppeteer: also solid, widely used
- Selenium: generally heavier; less ideal for high-throughput crawling
3) Block unnecessary resources
For most crawling tasks, you don’t need images, fonts, media, or ads.
Block:
- images
- video/audio
- fonts
- analytics
- trackers
- ads
This can cut load time and bandwidth dramatically.
4) Intercept network requests
Use request interception to:
- abort heavy resources
- log API calls
- capture JSON endpoints directly
- detect when the page is “done”
Often the best result is not the rendered DOM, but the underlying API response.
5) Set strong timeouts and heuristics
JS pages can hang forever if you wait on “network idle” blindly.
Use:
- navigation timeout
- DOM-ready timeout
- max total render time
- custom “page complete” signals
For example:
- wait for a selector
- wait for a specific API response
- wait for no network activity for N ms, with a hard upper bound
6) Use a distributed queue
At scale, a crawler should usually be:
- URL frontier / scheduler
- fetch/render workers
- deduplication layer
- result store
- retry/error handling
Message queues like:
- Redis streams / lists
- RabbitMQ
- Kafka
- SQS
help separate scheduling from rendering.
7) Use headless modes efficiently
Run in headless mode:
- no GUI
- disable unnecessary features
- run in containers if needed
But be careful:
- some sites fingerprint headless browsers
- you may need realistic user agents, viewport sizes, and browser fingerprints
8) Design for anti-bot resistance
At scale, sites may rate limit or block you.
Helpful techniques:
- rotate IPs responsibly
- respect robots and legal constraints
- randomize delays
- use realistic browser settings
- limit per-domain concurrency
- exponential backoff on errors / 429s
- maintain session cookies when needed
9) Persist the right artifact
Depending on your use case, save:
- final HTML
- extracted structured data
- screenshots
- network HAR
- rendered DOM snapshot
For debugging, screenshots + HAR files are extremely useful.
10) Measure cost per page
Track:
- median render time
- browser memory usage
- network bytes
- success rate
- per-domain block rate
- JS-needed ratio
Then optimize based on real data.
Recommended default stack
If you want a strong starting point:
- Python or Node.js
- Playwright
- Redis or RabbitMQ for queueing
- Docker for worker isolation
- PostgreSQL / S3 / object storage for results
- Prometheus + Grafana for metrics
A useful strategy
A lot of large crawlers use this pattern:
- Fetch HTML with plain HTTP
- Detect whether JS rendering is needed
- If yes, render with Playwright
- Block heavy assets
- Extract data from DOM and/or network responses
- Store results
- Retry with stricter or looser render rules if needed
This gets you better throughput than rendering everything.
Common mistakes
- rendering every page in a browser
- waiting too long for “network idle”
- not limiting concurrency per domain
- not blocking images/fonts
- not reusing browser instances
- relying only on DOM extraction when the API is easier
- ignoring memory leaks in long-lived workers
If you want, I can help you with one of these next:
- a reference architecture for a high-scale crawler
- a Playwright-based worker example
- a decision system for when to render JS
- a Kubernetes/container deployment approach