Prompt
What should I use for JS-rendered crawling at scale?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
If you mean crawling pages that require JavaScript to render, at scale, the usual choices are:
Best options
1. Playwright
Best default choice for JS-rendered crawling.
Why:
- Reliable modern browser automation
- Handles SPAs, hydration, lazy-loaded content, infinite scroll, etc.
- Good cross-browser support
- Better ergonomics and generally more robust than Selenium
- Can run headless and in parallel
Use it when:
- You need a real browser to execute JS
- You want a maintainable crawler
- You need to extract data from dynamic pages
2. Puppeteer
Great if you're mostly in the Chrome/Chromium ecosystem.
Why:
- Mature and widely used
- Good for simple or Chrome-only crawling
- Excellent ecosystem
Use it when:
- You only need Chromium
- You already have Puppeteer infrastructure
- You prefer its API or tooling
3. Browserless / managed browser infrastructure
If you want scale without managing browsers yourself, use a service like:
- Browserless
- Apify
- Zyte
- ScrapingBee
- Bright Data browser APIs
Use it when:
- You need to scale quickly
- You want to avoid maintaining fleets of headless browsers
- You need proxy management, retries, anti-bot handling, or browser pooling
What I’d recommend in practice
For most teams:
- Playwright + worker queue + autoscaled containers
- Add:
- request blocking for images/fonts/video if not needed
- concurrency limits per domain
- retry logic
- proxy rotation if necessary
- persistent browser contexts only when needed
For very large scale:
- Use Playwright in containers or serverless workers
- Backed by:
- a job queue (SQS, RabbitMQ, Redis queue, Kafka)
- browser pooling or one browser per worker
- observability (success rate, render time, captcha rate, block rate)
- Consider a managed browser provider if ops overhead gets too high
Important optimization: don’t render when you don’t have to
At scale, the biggest win is often:
- Try fetching the underlying API/XHR data
- If unavailable, use static HTML
- Only then use a real browser
This can reduce cost dramatically.
Avoid relying on
- Selenium for new projects unless you have legacy needs
- A single huge browser farm without queueing/backpressure
- Full rendering for every page if only a small subset actually needs it
Quick answer
- Default: Playwright
- If Chrome-only: Puppeteer
- If you want managed scale: Browserless/Apify/Zyte/Bright Data
- If possible: prefer API extraction over browser rendering
If you want, I can suggest a concrete architecture for 10k pages/day, 1M pages/day, or captcha-heavy sites.