Prompt
Scrapy large crawl architecture
Latest observation
Here’s a practical Scrapy large-crawl architecture pattern that scales beyond a single spider/process and keeps the crawl maintainable, restartable, and observable.
1) Core goals for a large crawl
For large crawls, the architecture should support:
- High throughput without overloading target sites
- Deduplication across long-running jobs
- Restartability after crashes or deploys
- Distributed execution if one machine is not enough
- Clear separation between crawling, parsing, storage, and downstream processing
- Observability: metrics, logs, failed URL tracking, backpressure
2) Recommended high-level architecture
A common scalable design:
-
Job Scheduler / Controller
- Decides what to crawl, when, and with what priority.
- Stores crawl seeds and crawl state in a database or queue.
-
Distributed Request Queue
- Keeps pending URLs/requests.
- Options: Redis, Kafka, RabbitMQ, PostgreSQL queue table.
- Scrapy itself can integrate well with Redis via
scrapy-redis.
-
Spider Workers
- Multiple Scrapy workers run the same spider code.
- They pull requests from the queue and push discovered requests back.
-
Item Pipeline
- Validates, transforms, and sends extracted data to storage.
- Usually writes to:
- PostgreSQL / MySQL for structured data
- Elasticsearch/OpenSearch for search
- S3/GCS for raw payloads
- Kafka for downstream processing
-
State Store
- Tracks:
- seen URLs / fingerprints
- crawl progress
- per-domain stats
- retry/failure counts
- Can be Redis + persistent DB, or DB-only for durability.
- Tracks:
-
Monitoring + Alerting
- Prometheus/Grafana, ELK/OpenSearch, Sentry
- Track:
- requests/sec
- response codes
- item throughput
- queue depth
- retry rate
- ban/throttle signals
3) Single-machine scaling first
Before going distributed, maximize one machine safely:
Scrapy settings to tune
CONCURRENT_REQUESTSCONCURRENT_REQUESTS_PER_DOMAINDOWNLOAD_DELAYRANDOMIZE_DOWNLOAD_DELAYAUTOTHROTTLE_ENABLEDAUTOTHROTTLE_START_DELAYAUTOTHROTTLE_TARGET_CONCURRENCYRETRY_TIMESDOWNLOAD_TIMEOUTDNSCACHE_ENABLED
Good practices
- Use async-friendly parsing, avoid blocking I/O in callbacks
- Keep item pipelines lightweight
- Offload heavy work to downstream services
- Use
FEEDSor pipelines carefully; large synchronous writes can bottleneck - Compress and batch writes where possible
4) Distributed crawling architecture
If you need many workers, a common pattern is:
A. Redis-based queue
Use scrapy-redis style architecture:
- Redis stores:
- request queue
- duplicate filter keys
- optional scheduler state
- Multiple Scrapy workers consume the same queue
This is simple and effective for:
- URL frontier management
- distributed dedupe
- horizontal scaling
B. Kafka-based frontier
Better when:
- you need strong event streaming
- large-volume, durable, replayable queues
- integration with multiple consumers
Pattern:
- spiders consume from Kafka topics
- discovered URLs are published back to a frontier topic
- a separate service handles prioritization and dedupe
Kafka is more operationally complex but powerful for very large crawls.
5) Suggested component breakdown
Crawl orchestrator
Responsibilities:
- seeds jobs
- assigns domains/projects
- controls rate limits
- pauses/resumes crawls
- requeues failed jobs
Could be implemented with:
- Airflow
- Celery beat
- a custom FastAPI service
- Kubernetes CronJobs for scheduled runs
Frontier service
Responsibilities:
- URL normalization
- dedupe
- prioritization
- per-domain politeness
- depth control
- allowed domains / robots policy
This can be:
- Redis-backed
- DB-backed
- Kafka + small policy service
Spider service
Responsibilities:
- fetch HTML/API responses
- parse response
- extract items and next URLs
- emit structured data and discovered links
Keep spider logic stateless as much as possible.
Storage service
Responsibilities:
- item validation
- schema enforcement
- upsert logic
- raw data retention
Use idempotent writes:
- unique keys
- UPSERT / MERGE
- content hash for change detection
6) Crawl state and deduplication
Large crawls fail when state is not managed well.
URL deduplication
Use:
- normalized URL fingerprint
- canonicalization rules:
- lowercase host
- remove fragments
- normalize query params if needed
- strip tracking params (
utm_*, etc.)
Content deduplication
Useful when pages are similar or frequently updated:
- hash of normalized content
- compare last-seen hash
- only process changes
Crawl checkpoints
Store:
- last successful crawl time
- offsets / queue positions
- spider version
- job ID
- error status
This enables:
- resuming partial crawls
- incremental recrawls
- rollback after bad deployment
7) Politeness and anti-ban strategy
At scale, you need to avoid being blocked.
Use:
- per-domain concurrency limits
- adaptive throttling
- backoff on 429/403/5xx
- rotating user agents only if appropriate and ethical
- session management for authenticated sites
- respect robots.txt where required by policy
Don’t:
- hammer a domain with global concurrency
- retry aggressively on bans
- rely only on proxies to solve load issues
8) Item pipeline architecture
A good pattern is a multi-stage pipeline:
-
Validation
- required fields
- type checks
- schema validation
-
Normalization
- dates, currencies, text cleanup
- canonical IDs
-
Enrichment
- derived fields
- joins with reference data
- geocoding / categorization if needed
-
Persistence
- batch insert/upsert
- write raw record to object storage if necessary
-
Event emission
- send to Kafka / downstream queue for ML/search/analytics
Keep each stage idempotent.
9) Error handling patterns
For large crawls, error handling must be explicit:
- Transient errors: retry with exponential backoff
- Permanent errors: drop and log with reason
- Parsing errors: capture HTML snapshot + traceback
- Schema errors: route to dead-letter queue
- Bans/rate limits: reduce concurrency or pause domain
Maintain a dead-letter store for:
- failed URLs
- reason codes
- last exception
- retry count
- response body sample
10) Deployment options
Docker + Kubernetes
Best for serious scaling:
- run spider workers as pods
- autoscale on queue depth
- separate Redis/Kafka/Postgres services
- rolling updates
VM-based
Simpler to start:
- one or more workers on VMs
- systemd or supervisord
- cron/airflow for orchestration
CI/CD
Important for spider code changes:
- version spider code
- deploy with config changes separately
- keep backward compatibility in item schemas
11) Example reference architecture
Small-to-medium scale
- Scrapy workers on 1–3 VMs
- Redis for request queue + dedupe
- PostgreSQL for item storage
- S3 for raw archives
- Grafana + Prometheus for metrics
Large scale
- Kubernetes cluster
- Redis or Kafka frontier
- many spider replicas
- PostgreSQL/OpenSearch/S3
- Airflow or custom orchestrator
- centralized logs and metrics
- dead-letter queue and replay tooling
12) Practical design principles
- Keep spiders stateless
- Make pipelines idempotent
- Treat queue and state as first-class systems
- Normalize URLs aggressively
- Use backpressure and throttling
- Separate extraction from storage
- Build replayable crawls
- Observe everything
13) Minimal recommended stack
If you want a pragmatic default:
- Scrapy
- Redis for distributed queue + dedupe
- PostgreSQL for structured item storage and crawl state
- S3 for raw snapshots
- Prometheus + Grafana for metrics
- Kubernetes when worker count grows
- Optional: Kafka if you need event streaming and replay
If you want, I can also provide:
- a diagram of this architecture,
- a sample Scrapy project structure for distributed crawling, or
- a Redis/Kafka-based implementation blueprint.