Prompt

Scrapy large crawl architecture

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Here’s a practical Scrapy large-crawl architecture pattern that scales beyond a single spider/process and keeps the crawl maintainable, restartable, and observable.


1) Core goals for a large crawl

For large crawls, the architecture should support:

  • High throughput without overloading target sites
  • Deduplication across long-running jobs
  • Restartability after crashes or deploys
  • Distributed execution if one machine is not enough
  • Clear separation between crawling, parsing, storage, and downstream processing
  • Observability: metrics, logs, failed URL tracking, backpressure

2) Recommended high-level architecture

A common scalable design:

  1. Job Scheduler / Controller

    • Decides what to crawl, when, and with what priority.
    • Stores crawl seeds and crawl state in a database or queue.
  2. Distributed Request Queue

    • Keeps pending URLs/requests.
    • Options: Redis, Kafka, RabbitMQ, PostgreSQL queue table.
    • Scrapy itself can integrate well with Redis via scrapy-redis.
  3. Spider Workers

    • Multiple Scrapy workers run the same spider code.
    • They pull requests from the queue and push discovered requests back.
  4. Item Pipeline

    • Validates, transforms, and sends extracted data to storage.
    • Usually writes to:
      • PostgreSQL / MySQL for structured data
      • Elasticsearch/OpenSearch for search
      • S3/GCS for raw payloads
      • Kafka for downstream processing
  5. State Store

    • Tracks:
      • seen URLs / fingerprints
      • crawl progress
      • per-domain stats
      • retry/failure counts
    • Can be Redis + persistent DB, or DB-only for durability.
  6. Monitoring + Alerting

    • Prometheus/Grafana, ELK/OpenSearch, Sentry
    • Track:
      • requests/sec
      • response codes
      • item throughput
      • queue depth
      • retry rate
      • ban/throttle signals

3) Single-machine scaling first

Before going distributed, maximize one machine safely:

Scrapy settings to tune

  • CONCURRENT_REQUESTS
  • CONCURRENT_REQUESTS_PER_DOMAIN
  • DOWNLOAD_DELAY
  • RANDOMIZE_DOWNLOAD_DELAY
  • AUTOTHROTTLE_ENABLED
  • AUTOTHROTTLE_START_DELAY
  • AUTOTHROTTLE_TARGET_CONCURRENCY
  • RETRY_TIMES
  • DOWNLOAD_TIMEOUT
  • DNSCACHE_ENABLED

Good practices

  • Use async-friendly parsing, avoid blocking I/O in callbacks
  • Keep item pipelines lightweight
  • Offload heavy work to downstream services
  • Use FEEDS or pipelines carefully; large synchronous writes can bottleneck
  • Compress and batch writes where possible

4) Distributed crawling architecture

If you need many workers, a common pattern is:

A. Redis-based queue

Use scrapy-redis style architecture:

  • Redis stores:
    • request queue
    • duplicate filter keys
    • optional scheduler state
  • Multiple Scrapy workers consume the same queue

This is simple and effective for:

  • URL frontier management
  • distributed dedupe
  • horizontal scaling

B. Kafka-based frontier

Better when:

  • you need strong event streaming
  • large-volume, durable, replayable queues
  • integration with multiple consumers

Pattern:

  • spiders consume from Kafka topics
  • discovered URLs are published back to a frontier topic
  • a separate service handles prioritization and dedupe

Kafka is more operationally complex but powerful for very large crawls.


5) Suggested component breakdown

Crawl orchestrator

Responsibilities:

  • seeds jobs
  • assigns domains/projects
  • controls rate limits
  • pauses/resumes crawls
  • requeues failed jobs

Could be implemented with:

  • Airflow
  • Celery beat
  • a custom FastAPI service
  • Kubernetes CronJobs for scheduled runs

Frontier service

Responsibilities:

  • URL normalization
  • dedupe
  • prioritization
  • per-domain politeness
  • depth control
  • allowed domains / robots policy

This can be:

  • Redis-backed
  • DB-backed
  • Kafka + small policy service

Spider service

Responsibilities:

  • fetch HTML/API responses
  • parse response
  • extract items and next URLs
  • emit structured data and discovered links

Keep spider logic stateless as much as possible.

Storage service

Responsibilities:

  • item validation
  • schema enforcement
  • upsert logic
  • raw data retention

Use idempotent writes:

  • unique keys
  • UPSERT / MERGE
  • content hash for change detection

6) Crawl state and deduplication

Large crawls fail when state is not managed well.

URL deduplication

Use:

  • normalized URL fingerprint
  • canonicalization rules:
    • lowercase host
    • remove fragments
    • normalize query params if needed
    • strip tracking params (utm_*, etc.)

Content deduplication

Useful when pages are similar or frequently updated:

  • hash of normalized content
  • compare last-seen hash
  • only process changes

Crawl checkpoints

Store:

  • last successful crawl time
  • offsets / queue positions
  • spider version
  • job ID
  • error status

This enables:

  • resuming partial crawls
  • incremental recrawls
  • rollback after bad deployment

7) Politeness and anti-ban strategy

At scale, you need to avoid being blocked.

Use:

  • per-domain concurrency limits
  • adaptive throttling
  • backoff on 429/403/5xx
  • rotating user agents only if appropriate and ethical
  • session management for authenticated sites
  • respect robots.txt where required by policy

Don’t:

  • hammer a domain with global concurrency
  • retry aggressively on bans
  • rely only on proxies to solve load issues

8) Item pipeline architecture

A good pattern is a multi-stage pipeline:

  1. Validation

    • required fields
    • type checks
    • schema validation
  2. Normalization

    • dates, currencies, text cleanup
    • canonical IDs
  3. Enrichment

    • derived fields
    • joins with reference data
    • geocoding / categorization if needed
  4. Persistence

    • batch insert/upsert
    • write raw record to object storage if necessary
  5. Event emission

    • send to Kafka / downstream queue for ML/search/analytics

Keep each stage idempotent.


9) Error handling patterns

For large crawls, error handling must be explicit:

  • Transient errors: retry with exponential backoff
  • Permanent errors: drop and log with reason
  • Parsing errors: capture HTML snapshot + traceback
  • Schema errors: route to dead-letter queue
  • Bans/rate limits: reduce concurrency or pause domain

Maintain a dead-letter store for:

  • failed URLs
  • reason codes
  • last exception
  • retry count
  • response body sample

10) Deployment options

Docker + Kubernetes

Best for serious scaling:

  • run spider workers as pods
  • autoscale on queue depth
  • separate Redis/Kafka/Postgres services
  • rolling updates

VM-based

Simpler to start:

  • one or more workers on VMs
  • systemd or supervisord
  • cron/airflow for orchestration

CI/CD

Important for spider code changes:

  • version spider code
  • deploy with config changes separately
  • keep backward compatibility in item schemas

11) Example reference architecture

Small-to-medium scale

  • Scrapy workers on 1–3 VMs
  • Redis for request queue + dedupe
  • PostgreSQL for item storage
  • S3 for raw archives
  • Grafana + Prometheus for metrics

Large scale

  • Kubernetes cluster
  • Redis or Kafka frontier
  • many spider replicas
  • PostgreSQL/OpenSearch/S3
  • Airflow or custom orchestrator
  • centralized logs and metrics
  • dead-letter queue and replay tooling

12) Practical design principles

  • Keep spiders stateless
  • Make pipelines idempotent
  • Treat queue and state as first-class systems
  • Normalize URLs aggressively
  • Use backpressure and throttling
  • Separate extraction from storage
  • Build replayable crawls
  • Observe everything

13) Minimal recommended stack

If you want a pragmatic default:

  • Scrapy
  • Redis for distributed queue + dedupe
  • PostgreSQL for structured item storage and crawl state
  • S3 for raw snapshots
  • Prometheus + Grafana for metrics
  • Kubernetes when worker count grows
  • Optional: Kafka if you need event streaming and replay

If you want, I can also provide:

  1. a diagram of this architecture,
  2. a sample Scrapy project structure for distributed crawling, or
  3. a Redis/Kafka-based implementation blueprint.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.