Prompt
What should I use to crawl JavaScript-rendered pages?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
For JavaScript-rendered pages, you usually want a headless browser or a browser automation framework, not a basic HTTP crawler.
Good options
- Playwright — my top recommendation for most cases
- Modern, fast, reliable
- Works with Chromium, Firefox, WebKit
- Great for SPAs and dynamic content
- Puppeteer — good if you only need Chromium-based browsing
- Mature and popular
- Slightly narrower than Playwright
- Selenium — best if you need broad browser/grid support or already use it
- More traditional
- Often heavier and slower for crawling
- Scrapy + Playwright — great if you want a real crawler framework plus JS rendering
- Scrapy handles crawling logic, retries, pipelines, etc.
- Playwright handles page rendering
Rule of thumb
- Simple JS site / modern scraping → Playwright
- Need a full crawler framework → Scrapy + Playwright
- Chromium-only and existing Puppeteer code → Puppeteer
- Enterprise/browser compatibility needs → Selenium
Important note
If the data is loaded via API calls after the page renders, it’s often better to:
- inspect the network requests,
- call the underlying API directly,
- avoid browser rendering unless necessary.
If you want, I can recommend the best tool for your specific stack or show a minimal Playwright crawler example.