Prompt
How do I crawl multiple domains on a schedule?
Latest observation
To crawl multiple domains on a schedule, you usually want a small system with four parts:
- A scheduler to decide when crawls run
- A crawl queue to hold work for each domain
- A crawler worker to fetch pages
- Storage/reporting for results, errors, and crawl state
Common setup
1) Define crawl jobs per domain
Keep a config like:
[
{"domain": "example.com", "schedule": "daily"},
{"domain": "example.org", "schedule": "hourly"},
{"domain": "example.net", "schedule": "weekly"}
]
Each domain gets its own:
- start URL(s)
- crawl depth limits
- allowed paths / disallowed paths
- rate limit
- schedule
2) Use a scheduler
Pick one:
- Cron for simple setups
- Celery Beat / Celery workers for Python
- Airflow if you need pipelines and observability
- Cloud scheduler (AWS EventBridge, GCP Cloud Scheduler, Azure) for managed scheduling
The scheduler triggers jobs like:
crawl example.comcrawl example.org
3) Queue crawl tasks
For multiple domains, don’t crawl everything in one process. Put URLs or domain jobs onto a queue:
- Redis Queue
- RabbitMQ
- SQS
- Kafka
This lets you:
- run crawls in parallel
- isolate failures
- retry safely
- control per-domain concurrency
4) Crawl politely
For each domain:
- respect
robots.txt - use delays/rate limits
- set a clear user-agent
- avoid too many concurrent requests to the same host
- stop at allowed depth/URL count
5) Store crawl state
Save:
- last crawl time
- discovered URLs
- page status codes
- hashes/content for change detection
- failures/retry counts
This helps you do incremental crawls and avoid re-fetching everything.
Example architecture
Cron / Scheduler
→ enqueues jobs for each domain
Queue
→ workers consume crawl jobs
Crawler workers
→ fetch URLs, extract links, enqueue more URLs if allowed
Database / Storage
→ results and crawl history
Python example with Celery
If you want a practical pattern:
- Celery Beat schedules domain jobs
- Celery workers perform crawling
Pseudo-structure:
# tasks.py
from celery import Celery
app = Celery("crawler", broker="redis://localhost:6379/0")
@app.task
def crawl_domain(domain):
# fetch seed URLs, crawl pages, store results
pass
Scheduled by Celery Beat:
beat_schedule = {
"crawl-example-com": {
"task": "tasks.crawl_domain",
"schedule": crontab(hour=2, minute=0),
"args": ("example.com",),
},
}
For many domains, you can generate these schedules dynamically from a database.
If you need this at scale
Consider:
- per-domain concurrency limits
- distributed deduplication of URLs
- content fingerprinting for change detection
- headless browser only when needed (avoid for every page)
- separate queues by domain or priority
Simpler option
If you only have a handful of domains:
- store domain list in a database or config file
- run a daily/hourly cron job
- the job loops through domains and invokes the crawler once per domain
That’s often enough to start.
Important caveats
- Make sure you’re allowed to crawl the sites
- Check
robots.txt - Follow site terms and legal restrictions
- Avoid aggressive crawl rates
If you want, I can give you:
- a Python/Celery implementation,
- a Node.js implementation, or
- a Docker + cron + Redis architecture.