Prompt
I'm unhappy with Playwright for crawling at scale
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
That’s a fair complaint—Playwright is great for testing and selective browser automation, but it can be a rough fit for large-scale crawling.
Common pain points at scale:
- Heavy per-page overhead: full browser instances are expensive in CPU/RAM.
- Throughput limits: you’ll hit resource ceilings fast compared to HTTP-first crawlers.
- Operational complexity: browser crashes, timeouts, retries, and anti-bot handling pile up.
- Cost: cloud/browser infrastructure gets expensive quickly.
- Concurrency tradeoffs: more parallel browsers often means more instability, not more throughput.
What usually works better depends on your target pages:
If the site is mostly static or lightly dynamic
Use an HTTP client + parser stack:
httpx,aiohttp, orrequestslxml,BeautifulSoup,selectolax- Add a crawl framework like Scrapy for queueing, retries, deduping, throttling
This is usually the best option for scale.
If some JS is needed
Use a hybrid approach:
- Crawl mostly with HTTP
- Only send pages to a browser renderer when necessary
- Cache rendered output aggressively
- Consider route-based rules: “render only these URL patterns”
If you need browser automation at scale
You may want:
- browserless or managed headless browser infrastructure
- A crawler framework that supports browser pools
- Strong request interception to block images/fonts/video
- Tight concurrency control and queue-based job processing
Alternatives to consider
- Scrapy for core crawling
- Playwright only for rendering fallback
- Crawlee if you want a higher-level crawling abstraction
- Apify if you want managed infrastructure
- Puppeteer if you’re already in the Node ecosystem, though it won’t magically fix scale issues
Practical recommendation
A robust pattern is:
- Start with Scrapy/HTTP crawling
- Detect pages that need JS
- Escalate only those to Playwright
- Store rendered HTML and reuse it
- Keep browser concurrency low and HTTP concurrency high
If you want, I can help you design a high-scale crawling architecture that minimizes Playwright usage, or compare Scrapy vs Crawlee vs a custom HTTP crawler for your use case.