Prompt
I'm building a browser automation service for scraping, what infrastructure do I need?
Latest observation
For a browser-automation scraping service, think in terms of orchestration, browser execution, resilience, and observability. The exact setup depends on scale and target sites, but a solid production architecture usually looks like this:
1) Core components
A. API / control plane
This is your public or internal service that:
- accepts scrape jobs
- validates input
- stores job metadata
- enqueues work
- tracks status/results
Typical stack:
- REST or gRPC API
- Postgres for job state
- Redis or a queue for dispatch
B. Job queue / scheduler
Use a queue so jobs are decoupled from workers and can retry cleanly.
Options:
- Redis + BullMQ
- RabbitMQ
- SQS
- Kafka if you need high throughput/event streaming
You’ll want:
- retries with backoff
- dead-letter queue
- concurrency limits
- prioritization if needed
C. Browser worker fleet
These are the machines/containers that actually run Chromium/Firefox/WebKit.
You have two main models:
Model 1: One browser per job
- Strong isolation
- Simple failure handling
- Higher resource usage
- Good for untrusted workloads or anti-bot-heavy sites
Model 2: Browser pool / session reuse
- More efficient
- Better throughput
- Harder to isolate state and failures
Typical worker stack:
- Docker containers running headless Chromium
- Playwright or Puppeteer
- Optional Xvfb only if you need non-headless rendering, though most headless setups don’t
D. Storage
You’ll likely need:
- Postgres: jobs, runs, user accounts, configs
- Object storage: screenshots, HTML snapshots, PDFs, HAR files, logs
- S3, GCS, Azure Blob, or MinIO
- Redis: caching, locks, rate limits, ephemeral session data
2) Infrastructure pieces
Compute
Depending on workload:
- Kubernetes if you want elastic scaling and isolation
- ECS/Fargate for simpler AWS-managed containers
- VMs if you want more control and predictable browser performance
For browser workloads, CPU and RAM matter a lot:
- Chromium can be memory-heavy
- Plan for per-browser memory budgets
- Separate “API” nodes from “worker” nodes
Networking
You’ll need:
- outbound internet access from workers
- egress IP management
- DNS reliability
- proxy support if you rotate IPs
If scraping at scale, consider:
- residential or datacenter proxy providers
- per-tenant proxy pools
- IP rotation and geo-targeting
- egress NAT gateways with good monitoring
Secrets management
Store:
- site credentials
- API keys
- proxy credentials
- session cookies or encryption keys
Use:
- AWS Secrets Manager / GCP Secret Manager / Vault / Kubernetes secrets with care
3) Browser-specific requirements
A. Container image
Build a stable browser image with:
- Chromium/Chrome
- Playwright/Puppeteer dependencies
- fonts, locales, certificates
- proper sandbox support if possible
Common issues:
- missing fonts causing layout differences
- sandbox restrictions in containers
/dev/shmtoo small for Chromium- kernel limits on file descriptors/processes
Useful configuration:
- increase
/dev/shmor mount a larger shared memory volume - tune ulimits
- run as non-root when possible
B. Session management
If you need login-based scraping:
- isolated browser contexts per task
- secure cookie storage
- session refresh logic
- MFA handling strategy if applicable
C. Anti-bot and fingerprinting
Many sites detect automation. If this is legitimate scraping and allowed:
- use realistic browser profiles
- manage headers consistently
- handle user agents, locale, timezone
- consider proxy quality and IP reputation
- add rate limiting and human-like pacing
Avoid trying to bypass site protections in ways that violate terms or law.
4) Reliability and scaling
A. Autoscaling
Scale workers based on:
- queue depth
- average job duration
- CPU/RAM usage
- browser crash rate
Common patterns:
- Kubernetes Horizontal Pod Autoscaler
- queue-length-based scaling
- separate pools for “light” and “heavy” jobs
B. Retries and idempotency
Scraping jobs fail a lot. Design for:
- retryable vs non-retryable errors
- page timeouts
- navigation failures
- selector timeouts
- stale sessions
Make jobs idempotent so retries don’t duplicate side effects.
C. Timeouts and watchdogs
Set:
- navigation timeout
- step timeout
- job timeout
- worker heartbeat timeout
Kill and restart workers that hang.
5) Observability
You’ll want:
- structured logs per job
- screenshots on failure
- HTML snapshots on failure
- traces for steps
- metrics:
- job success/failure rate
- queue latency
- page load times
- browser crash counts
- proxy error rates
- per-site success rate
Tooling:
- Prometheus + Grafana
- OpenTelemetry
- ELK/Opensearch
- Sentry for exceptions
6) Security and compliance
Because scraping can handle credentials and third-party sites:
- encrypt data at rest and in transit
- isolate tenants and jobs
- sanitize logs
- limit secret exposure in browser contexts
- audit access to scraped data
- respect robots.txt, site terms, and applicable laws
If you store customer-provided credentials or cookies, treat them like secrets.
7) Suggested reference architecture
A practical setup:
-
API service
- receives scrape request
- stores job in Postgres
- pushes to queue
-
Queue
- Redis/BullMQ or SQS
-
Worker fleet
- Docker containers on Kubernetes/ECS
- each worker executes one or a few browser jobs
- browser automation using Playwright
-
Storage
- Postgres for metadata
- S3 for artifacts
- Redis for locks/cache/rate limits
-
Observability
- logs, metrics, screenshots, traces
-
Proxy/egress layer
- optional proxy pool with rotation and monitoring
8) What you need at minimum to start
If you’re early-stage, a minimal production-ish stack could be:
- Node.js or Python API
- Postgres
- Redis queue
- Dockerized Playwright workers
- S3-compatible storage
- Prometheus/Grafana or a hosted logging/monitoring tool
Run workers on:
- a few VMs or a small Kubernetes cluster
This is usually enough to support an MVP and grow from there.
9) Key design choices to decide upfront
-
Playwright or Puppeteer?
- Playwright is often better for multi-browser support and stability
-
One browser per job or pooled browsers?
- One per job for simplicity and isolation
- Pooling for efficiency
-
Queue-backed or synchronous?
- Queue-backed is strongly recommended
-
Self-hosted or managed containers?
- Managed is easier operationally; self-hosted offers more control
-
Do you need proxies?
- Depends on target sites and geography
-
How will you store artifacts?
- Screenshots, HTML, PDFs, and traces should go to object storage
If you want, I can also give you:
- a reference architecture diagram
- a concrete AWS/GCP/Azure setup
- a low-cost MVP stack
- or a Kubernetes deployment pattern for Playwright workers