Prompt
I'm building a site monitoring pipeline, what crawling stack makes sense?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
It depends a lot on what “monitoring” means for you, but for most site-monitoring pipelines I’d think in three layers:
- Discovery / scheduling — what to crawl and when
- Fetch / render — how to retrieve pages reliably
- Extract / compare / alert — what changed and whether it matters
A practical stack that works well for most cases
1) Orchestration / queue
- Python + Celery / RQ / Dramatiq for simple-to-medium pipelines
- Kafka / RabbitMQ / Redis queues if you need scale or many workers
- Airflow if this is more of a scheduled data pipeline than a low-latency crawler
2) Crawling / fetching
Start with a split approach:
- HTTP-first fetcher for most pages:
- Python:
httpx,aiohttp, orrequests - Node:
undiciorgot
- Python:
- Headless browser only when needed:
- Playwright is usually the best default today
- Use it for JS-heavy pages, login flows, shadow DOM, and sites that require execution
This hybrid approach is usually much cheaper and more reliable than using a browser for everything.
3) Parsing / extraction
- BeautifulSoup or lxml for HTML parsing in Python
- parsel / cssselect if you like selector-based extraction
- For structured data, consider extracting:
- canonical URL
- title/meta description
- visible text snapshot
- key DOM blocks
- links
- schema.org / JSON-LD
4) Change detection
Depending on the type of monitoring:
- Text diff for content changes
- DOM tree diff for structural changes
- Screenshot diff for visual regressions
- Hashing normalized content to reduce noise
- Use rules to ignore:
- timestamps
- ads
- rotating banners
- personalization regions
5) Storage
- PostgreSQL for metadata, crawl jobs, snapshots index
- S3 / object storage for raw HTML, screenshots, and rendered artifacts
- Optional:
- Elasticsearch/OpenSearch if you need search over captured content
- Redis for dedupe / rate limiting / locks
Recommended stack by use case
A) Lightweight monitoring of many sites
Use:
- Python
- httpx + BeautifulSoup
- Playwright only for fallback
- Postgres + S3
- Celery + Redis
Good for:
- uptime/content monitoring
- detecting page changes
- moderate scale
B) JS-heavy or authenticated sites
Use:
- Playwright as primary renderer
- Python or Node orchestration
- Postgres + object storage
- Queue worker system
Good for:
- dashboards
- SPA sites
- logged-in flows
- sites with anti-bot protections
C) Large-scale crawling
Use:
- Scrapy as the core crawler
- Scrapy-Playwright for selective rendering
- Kafka/RabbitMQ
- Postgres + S3 + OpenSearch
- distributed workers on Kubernetes
Good for:
- broad crawling
- large link graphs
- high throughput
My default recommendation
If you’re starting a new monitoring pipeline today:
Python + httpx + BeautifulSoup/lxml + Playwright + Postgres + S3 + Redis/Celery
That gives you:
- cheap HTTP fetches for most pages
- browser rendering only when needed
- straightforward storage and retry handling
- an easy path to scale later
A few design choices that matter a lot
Use two fetch paths
- Fast path: plain HTTP
- Fallback path: Playwright rendering
This saves a ton of compute.
Normalize before diffing
Strip:
- scripts/styles
- timestamps
- whitespace noise
- tracking params
- session IDs
Keep raw artifacts
Store:
- raw HTML
- rendered HTML
- screenshot
- HTTP headers/status
- extracted normalized text
This makes debugging much easier.
Respect rate limits and robots where appropriate
Use:
- host-level concurrency caps
- exponential backoff
- user-agent identification
- politeness delays
If you want, I can help you choose a stack based on:
- how many sites/pages per day
- whether pages are JS-heavy
- whether you need screenshots
- whether you need login/session support
- your preferred language/runtime
If you share those constraints, I can suggest a concrete architecture.