Prompt
How do I queue and retry large crawl jobs?
Latest observation
To queue and retry large crawl jobs, the usual pattern is to make crawling asynchronous, chunked, and idempotent.
Recommended approach
1. Put crawl requests onto a queue
Instead of doing the crawl in the request/response path:
- API receives crawl job request
- Store a
crawl_jobrecord in a database - Enqueue a message like:
job_id- target URL / seed list
- crawl parameters
- attempt count
Workers then pull jobs from the queue and process them in the background.
Good queue options:
- SQS
- RabbitMQ
- Redis Queue / Celery
- Kafka for very high throughput
- Cloud-native job runners like Cloud Tasks, Sidekiq, BullMQ, etc.
2. Break large crawls into smaller units
A “large crawl” should not be one giant job. Split it into:
- seed jobs
- page fetch jobs
- discovery jobs for new URLs
- content processing jobs
This lets you retry only the failed piece rather than the entire crawl.
Example:
- Job A: crawl domain root
- Job B: crawl discovered URLs batch 1
- Job C: crawl discovered URLs batch 2
3. Make every crawl step idempotent
Retries will happen, so each worker should be safe to run more than once.
Use:
- Unique job IDs
- Deduplication keys per URL + crawl run
- “Upsert” writes instead of inserts
- A status table like:
pendingin_progresscompletedfailedretrying
If a page was already fetched successfully, retrying shouldn’t duplicate it.
4. Track attempts and failure reasons
Store metadata for each job:
attempt_countlast_errornext_retry_atstatusstarted_atfinished_at
This makes it easy to:
- retry only transient failures
- stop after max attempts
- inspect problematic URLs
5. Use exponential backoff for retries
For transient failures like:
- network timeouts
- 429 rate limits
- temporary 5xx responses
Retry with backoff:
- 1 min
- 5 min
- 15 min
- 1 hour
Add jitter to avoid retry storms.
Example rule:
- retry up to 5 times
- only retry on timeout / 429 / 5xx
- do not retry on 4xx like 404 or 403 unless your business logic says otherwise
6. Persist crawl state externally
Don’t keep critical crawl state only in memory.
Persist:
- visited URLs
- discovered URLs
- frontier / pending URL set
- per-URL status
- checkpoints for each batch
That way, if a worker crashes, another worker can resume.
7. Use checkpoints for very large crawls
For multi-hour or multi-million-page crawls, checkpoint progress regularly.
Example:
- every 1,000 URLs processed
- save frontier state
- save last successful batch
- resume from checkpoint after failure
8. Separate “job orchestration” from “fetching”
A common design:
- orchestrator: creates batches, assigns work, handles retries
- fetch workers: download pages
- parser workers: extract links/content
- storage workers: write results
This improves resilience and scaling.
Simple architecture example
- User submits crawl for
example.com - API creates
crawl_job = 123 - Enqueue seed URL job
- Worker fetches page, extracts 200 links
- Worker writes results and enqueues 200 child URL jobs
- Each child job is retried independently if it fails
- Orchestrator marks job complete when frontier is empty
Practical retry policy
- Retry only on transient errors
- Use max retry count
- Exponential backoff + jitter
- Dead-letter queue after max retries
- Manual requeue for dead-lettered jobs
Important gotchas
- Duplicate processing: always assume at-least-once delivery
- Queue explosion: batch or rate-limit URL discovery
- Politeness: respect robots.txt, crawl-delay, domain rate limits
- Memory usage: don’t store the whole frontier in RAM for large crawls
- Timeouts: set worker timeouts lower than queue visibility timeouts if using SQS-like systems
If you want a concrete implementation
I can show you:
- a Redis/Celery design,
- an AWS SQS + worker design,
- or a SQL-backed job table pattern.
If you tell me your stack, I can sketch the exact queue/retry setup.