Prompt
What should I use instead of Playwright for large-scale crawling?
Latest observation
If your goal is large-scale crawling, Playwright is usually the wrong default because it’s heavy: every page can spin up a full browser context, which is expensive in CPU, RAM, and throughput.
Better options, depending on what you need
1) HTTP client + HTML parser
Best if the site content is mostly in the initial HTML or can be fetched via APIs.
- Node.js:
undici,got,axios+cheerio - Python:
httpx,aiohttp,requests+lxml,BeautifulSoup - Go:
colly,net/http+goquery
Why: fastest, cheapest, easiest to scale.
2) Scrapy
Best if you want a real crawling framework with concurrency, retries, deduping, pipelines, and scheduling.
- Great for:
- thousands to millions of pages
- queue-based crawling
- rate limiting and backoff
- parsing and data pipelines
Why: built for scale; much lighter than browser automation.
3) Apify Crawlee
If you liked Playwright’s convenience but want a crawler framework around it.
- Supports:
- HTTP crawling
- browser-based crawling when needed
- autoscaling
- request queues
- session management
Why: more crawler-oriented than raw Playwright.
4) Bright Data / ZenRows / ScrapingBee / browserless-type services
If you need to crawl sites with heavy anti-bot protection or JS rendering but don’t want to manage browsers yourself.
Why: offload browser management, proxies, fingerprinting, retries.
5) Headless browser only for “last mile”
A practical pattern is:
- Use HTTP crawling first
- Fall back to Playwright only for pages that truly require JS rendering
This hybrid approach is often the best at scale.
Rule of thumb
- Static or API-backed sites: use Scrapy / HTTP clients
- Large-scale structured crawling: use Scrapy
- Need JS occasionally: use hybrid HTTP + Playwright fallback
- Need JS everywhere and want managed infra: use Crawlee or a scraping platform
If you tell me:
- your language,
- whether the target pages are JS-heavy,
- your scale target, I can recommend the best replacement stack more specifically.