Prompt
I'm building a scraper for pages behind login walls; what infrastructure should I use?
Latest observation
For scraping pages behind login walls, the “right” infrastructure depends less on scraping itself and more on auth/session handling, browser automation, and operational safety. A practical stack usually looks like this:
Recommended setup
1) Use a browser automation layer
For login-walled pages, you’ll usually need a real browser:
- Playwright is often the best choice
- Great for modern sites
- Handles logins, cookies, MFA-ish flows better than basic HTTP scraping
- Easy to run headless in containers
- Selenium is still viable, especially if you already know it
- Puppeteer if you’re in Node and prefer Chrome-first tooling
Recommendation: start with Playwright unless you have a reason not to.
2) Run it in containers
Use Docker so your scraper is reproducible and deployable.
Typical container setup:
- Base image with Playwright browsers or Chromium installed
- Your scraper code
- Optional: a lightweight Linux distro if you need custom dependencies
This makes it easier to:
- deploy to cloud VMs
- scale workers horizontally
- keep browser dependencies stable
3) Use a job queue + worker model
For anything beyond a small personal scraper, use:
- Queue: Redis + BullMQ / RQ / Celery / Sidekiq
- Workers: one browser session per job or a small pool
This helps with:
- retries
- rate limiting
- backoff
- crash isolation
- scaling up/down cleanly
4) Persist sessions securely
Login-based scraping usually works best if you:
- log in once
- store session cookies / storage state
- reuse them until they expire
With Playwright, you can save:
storageState- cookies
- localStorage, depending on site behavior
Store these in:
- encrypted secrets storage
- a vault or secret manager
- or at minimum, protected blob storage with access control
Avoid hardcoding credentials into the codebase.
5) Add monitoring and observability
You’ll want visibility into:
- login failures
- captcha or MFA prompts
- expired sessions
- blocked requests
- DOM changes
- extraction errors
Useful tools:
- structured logs
- Sentry for exceptions
- Prometheus/Grafana for metrics
- screenshot/video capture on failures
Infra choices by scale
Small / prototype
- Local machine or one VM
- Playwright
- Docker
- simple cron or a basic queue
Medium
- Docker containers on a VM pool
- Redis queue
- 1+ worker services
- object storage for outputs and session artifacts
- monitoring + alerting
Larger scale
- Kubernetes or ECS
- autoscaling worker pods/tasks
- centralized secrets management
- browser session pooling
- queue-based orchestration
- strong observability and per-target rate limiting
Important operational considerations
Respect access rules
Only scrape systems you’re authorized to access. Login walls often imply:
- terms of service constraints
- user consent requirements
- rate limits
- audit/logging expectations
Handle anti-bot measures carefully
Some sites use:
- CSRF tokens
- device fingerprinting
- bot detection
- CAPTCHA
- MFA
If you encounter these, the safest route is usually:
- use official APIs if available
- request access/partnership
- avoid trying to bypass protections in ways that violate policy or law
Don’t over-optimize too early
A lot of scrapers fail because they use too much infrastructure too soon. Start with:
- one browser framework
- one worker
- session persistence
- logging Then scale only after the flow is stable.
A good default stack
If you want a concrete recommendation:
- Playwright
- Python or Node.js
- Docker
- Redis queue
- Postgres for metadata/output state
- S3-compatible storage for artifacts/screenshots/session backups
- Sentry + logs
- VMs or ECS/Kubernetes for deployment
If you want, I can propose:
- a minimal architecture for a single-site scraper,
- a production architecture for many sites, or
- a Playwright-based reference setup with containers and queues.