Prompt
What stack would you suggest for a crawler that can handle static HTML and JavaScript-rendered pages?
Latest observation
For a crawler that needs to handle both static HTML and JavaScript-rendered pages, I’d suggest a stack built around fast HTTP fetching + headless browser rendering + a queue-based orchestration layer.
Suggested stack
1) Core language/runtime
Python or Node.js
- Python is great if you want mature crawling, data processing, and scheduling tools.
- Node.js is very strong if a lot of your work depends on browser automation and JS-heavy sites.
If I had to pick one for general crawling: Python.
2) Static page fetching
Use a fast HTTP client:
- Python:
httpxoraiohttp - Node.js:
got,undici, or nativefetch
For parsing:
- BeautifulSoup4 or lxml in Python
- Cheerio in Node
This handles most pages efficiently without a browser.
3) JavaScript-rendered pages
Use a headless browser:
- Playwright is my top recommendation
- Alternatives: Puppeteer or Selenium
Playwright is especially good because:
- It’s reliable across Chromium/Firefox/WebKit
- Better modern API
- Good waiting/auto-sync behavior
- Handles SPA sites well
A common approach is:
- Try raw HTTP first
- If content is incomplete or JS-dependent, fall back to Playwright
4) Crawling orchestration
For scalability and resilience:
- Scrapy if you’re in Python and want a mature crawler framework
- Or build your own async crawler with:
asynciohttpxPlaywright- a queue like Redis
If you want a ready-made crawling framework in Python:
- Scrapy + Playwright integration is a very solid choice.
5) Queue / distributed processing
If crawling at medium or large scale:
- Redis for task queueing
- Celery or RQ for workers in Python
- RabbitMQ if you need stronger messaging guarantees
For simpler setups, Redis is often enough.
6) Storage
Depends on your data:
- PostgreSQL for structured crawl results
- MongoDB if records are semi-structured
- S3 / object storage for raw HTML, screenshots, page dumps
- Elasticsearch/OpenSearch if you need search/indexing
7) Anti-bot / reliability helpers
Useful additions:
- Proxy support
- Rate limiting / throttling
- Retry logic with backoff
- User-agent rotation
- Cookie/session handling
- CAPTCHA detection/reporting
Best practical stack recommendation
If you want the strongest general-purpose Python stack:
- Scrapy
- Playwright
- Redis
- PostgreSQL
- Docker
This gives you:
- Fast static crawling via Scrapy
- JS rendering via Playwright
- Easy scaling with Redis workers
- Durable data storage in PostgreSQL
If you prefer Node.js
A strong Node stack would be:
- Playwright
- BullMQ + Redis
- Cheerio
- PostgreSQL
- Docker
My recommendation in one line
Use Scrapy for the crawl engine, Playwright for JS rendering, Redis for queues, and PostgreSQL for storage.
If you want, I can also suggest:
- a minimal MVP stack,
- a production-scale distributed stack, or
- a reference architecture diagram for this crawler.