Prompt
How do I crawl and parse data from millions of pages?
Latest observation
Crawling and parsing millions of pages is mainly a systems engineering problem, not just a scraping problem. The winning pattern is:
- Discover URLs
- Schedule and fetch at scale
- Parse/extract in a streaming pipeline
- Deduplicate and normalize
- Store efficiently
- Monitor, retry, and respect site limits
Here’s a practical blueprint.
1) Decide what kind of pages you’re crawling
Different sources need different strategies:
- Static HTML: easiest; use HTTP fetch + HTML parser.
- JS-rendered pages: may need headless browsers, but these are expensive at scale.
- APIs behind pages: best if available; use the API instead of the page.
- Already-archived web: consider Common Crawl or other datasets before crawling from scratch.
If you can get the data from an API or bulk dump, do that first. It will be much cheaper and more reliable.
2) Build a distributed crawl pipeline
For millions of pages, avoid a single monolithic script.
Core components
- URL frontier / queue: holds URLs to crawl
- Fetcher workers: download pages
- Parser workers: extract links/data
- Deduplication store: prevent revisiting the same URL/content
- Persistence layer: store raw HTML and parsed records
- Scheduler: enforces politeness, retries, prioritization
Common architectures
- Message queue-based: Kafka / RabbitMQ / SQS + worker pool
- Batch-based: generate URL batches, process with Spark/Flink/Beam
- Crawler framework: Scrapy + distributed queue, or custom service
For very large scale, a queue-based or stream-processing architecture is usually best.
3) Separate fetching from parsing
Do not parse deeply inside your fetcher if volume is high.
Better flow
- Fetcher downloads HTML
- Store raw response immediately
- Emit a job/event for parsing
- Parser extracts fields and links
- Links go back to the frontier
This separation gives you:
- retries without refetching everything
- easier scaling
- ability to re-parse old HTML when extraction logic changes
4) Use deduplication aggressively
Millions of pages means lots of duplicates and near-duplicates.
URL dedup
Normalize URLs:
- lowercase host
- remove fragments
- sort/remove tracking parameters
- canonicalize trailing slashes
- normalize default ports
Content dedup
Use hashes:
- exact hash of raw content
- near-duplicate detection via SimHash / MinHash if needed
This saves bandwidth, storage, and parsing time.
5) Be polite and resilient
At scale, bad crawl behavior gets you blocked quickly.
Respect:
robots.txt- rate limits per domain
- crawl delays
- user-agent identification
Robust fetching:
- timeouts
- retries with exponential backoff
- redirect handling
- connection pooling
- compressed responses
- per-domain throttling
A good crawler is “many domains, but slow per domain.”
6) Parse with tools suited to scale
For HTML:
- lxml: fast and robust
- selectolax: very fast parser
- BeautifulSoup: easy, but slower
- Go/Rust parsers: good if you need maximum throughput
For extracting structured data:
- XPath / CSS selectors
- regex only for small, well-defined patterns
- DOM cleanup before extraction if pages are messy
For millions of pages:
Prefer deterministic selectors and simple extraction rules over complex heuristics when possible.
7) Store raw and extracted data separately
Raw layer
Store fetched pages in:
- object storage (S3/GCS/Azure Blob)
- compressed format: gzip/zstd
- partitioned by date/domain/source
Structured layer
Store parsed output in:
- Parquet for analytics
- JSONL for simplicity
- database if you need transactional access
A common design:
- raw HTML in object storage
- metadata in a relational DB
- parsed records in Parquet on a data lake
8) Design for incremental crawling
Don’t recrawl everything every time.
Track:
- last fetched time
- HTTP status
- ETag / Last-Modified
- content hash
- change frequency per domain/page
Then do:
- frequent recrawls for fast-changing pages
- slow recrawls for stable pages
- conditional requests where possible (
If-Modified-Since,If-None-Match)
This cuts costs dramatically.
9) Handle malformed and heterogeneous pages
At scale, expect:
- broken HTML
- mixed encodings
- mobile/desktop variants
- anti-bot pages
- soft 404s
- captchas
Mitigations:
- encoding detection
- robust HTML parser recovery
- content-type checks
- page classification before extraction
- fallback rules for known templates
10) Measure everything
You need observability.
Track:
- fetch success rate
- latency by domain
- HTTP status codes
- parse success rate
- extraction coverage
- duplicate rate
- queue depth
- retry rate
- ban/captcha rate
Without metrics, large crawling projects drift and fail quietly.
11) A practical stack
Here’s a reasonable stack for millions of pages:
Small-to-medium scale
- Python + Scrapy
- Redis or RabbitMQ for queueing
- PostgreSQL for metadata
- S3 for raw HTML
- Parquet for output
Larger scale
- Kafka or SQS for scheduling
- Multiple fetcher services in containers
- lxml/selectolax parsers
- S3 + Parquet
- Spark/Flink/Beam for downstream processing
- Prometheus/Grafana for monitoring
12) Sample pipeline design
A simple production architecture:
- Seed URL list loaded into queue
- Fetcher pulls URL
- Applies domain politeness and robots check
- Downloads page
- Saves raw response to storage
- Emits parse job
- Parser extracts:
- target fields
- outgoing links
- canonical URL
- metadata
- New links are normalized and deduped
- Unseen links are pushed back to queue
- Metrics collected at every step
13) If you’re just starting, don’t overbuild
Start with:
- a small seed set
- one worker
- URL normalization
- raw storage
- basic parser
- deduplication
- retry logic
Then scale horizontally.
Most “millions of pages” projects fail because they start with browser automation everywhere or try to parse everything in one script.
14) Important legal/ethical note
Crawl only where you have permission or a legitimate basis to do so. Respect site policies, robots rules where applicable, rate limits, and privacy/security constraints. Avoid collecting sensitive personal data unless you have a lawful reason and a compliance process.
Recommended implementation approach
If you want a concrete path:
- Use Scrapy for crawling if pages are mostly static
- Use S3 + JSONL/Parquet for storage
- Add Redis/Kafka for distributed scheduling
- Use lxml/selectolax for fast parsing
- Add domain-level throttling and URL canonicalization
- Build monitoring + dedup + retries from day one
If you want, I can give you:
- a reference architecture diagram,
- a Scrapy-based implementation, or
- a Python pseudocode crawler for millions of pages.