Prompt
I'm building a pipeline to revisit pages daily, what infrastructure do I need?
Latest observation
To revisit pages daily, you generally need a small crawling/scheduling system with these core pieces:
1) URL inventory
A place to store the pages you want to revisit:
- URL
- last crawled time
- crawl frequency / priority
- status / error history
- content fingerprint or etag/last-modified if available
This can be a database table, key-value store, or queue-backed metadata store.
2) Scheduler
Something that decides what to crawl next and when:
- runs daily
- selects URLs due for refresh
- avoids overloading the same domain
- can re-queue failed pages
- supports priority and rate limits
Common options:
- cron + worker jobs
- Airflow / Dagster / Prefect
- a custom scheduler service
- queue delay jobs
3) Fetching workers
A set of workers that:
- download the page
- handle redirects, timeouts, retries
- respect robots.txt / crawl rules if needed
- set a sensible user agent
- optionally use conditional requests (
If-None-Match,If-Modified-Since) to reduce bandwidth
These are usually horizontally scalable containers or serverless jobs.
4) Queue / job system
A message queue to distribute work:
- SQS, RabbitMQ, Kafka, Redis queues, Pub/Sub, etc.
- scheduler enqueues URLs
- workers consume URLs
This helps with retries, scaling, and backpressure.
5) Storage
You’ll want somewhere to store:
- raw HTML or fetched content
- extracted text/metadata
- crawl logs and statuses
- change history / diffs if you track updates
Typical choices:
- object storage for raw pages
- relational DB for metadata
- search index or document store for content
6) Change detection
If the goal is “revisit” because pages may change, you need logic to detect changes:
- compare hashes of content
- compare extracted fields
- use HTTP headers like ETag / Last-Modified
- store snapshots for diffing
7) Monitoring and alerting
Important for daily crawling:
- success/failure rates
- queue depth
- per-domain error rates
- latency
- dropped jobs
- alerting on systemic failures
8) Compliance and politeness controls
Especially if crawling external sites:
- robots.txt handling
- domain rate limits
- backoff on errors
- user-agent identification
- legal/ToS review
Simple reference architecture
Scheduler → Queue → Crawler workers → Storage
↘ Logs/metrics/alerts
Minimal stack for a small pipeline
If this is modest scale, you can start with:
- Postgres for URL state
- Cron for daily scheduling
- Celery/RQ/Sidekiq workers
- Redis as queue
- S3/GCS for raw HTML snapshots
- Prometheus/Grafana or cloud monitoring for observability
If you expect scale
Consider:
- distributed queue
- autoscaling workers
- per-domain politeness manager
- deduplication service
- content fingerprinting pipeline
- separate fetch/extract/index stages
If you want, I can sketch:
- a minimal architecture for 1k–100k pages/day, or
- a production-grade architecture for millions of pages/day.