Prompt

How do I crawl and parse data from millions of pages?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Crawling and parsing millions of pages is mainly a systems engineering problem, not just a scraping problem. The winning pattern is:

  1. Discover URLs
  2. Schedule and fetch at scale
  3. Parse/extract in a streaming pipeline
  4. Deduplicate and normalize
  5. Store efficiently
  6. Monitor, retry, and respect site limits

Here’s a practical blueprint.


1) Decide what kind of pages you’re crawling

Different sources need different strategies:

  • Static HTML: easiest; use HTTP fetch + HTML parser.
  • JS-rendered pages: may need headless browsers, but these are expensive at scale.
  • APIs behind pages: best if available; use the API instead of the page.
  • Already-archived web: consider Common Crawl or other datasets before crawling from scratch.

If you can get the data from an API or bulk dump, do that first. It will be much cheaper and more reliable.


2) Build a distributed crawl pipeline

For millions of pages, avoid a single monolithic script.

Core components

  • URL frontier / queue: holds URLs to crawl
  • Fetcher workers: download pages
  • Parser workers: extract links/data
  • Deduplication store: prevent revisiting the same URL/content
  • Persistence layer: store raw HTML and parsed records
  • Scheduler: enforces politeness, retries, prioritization

Common architectures

  • Message queue-based: Kafka / RabbitMQ / SQS + worker pool
  • Batch-based: generate URL batches, process with Spark/Flink/Beam
  • Crawler framework: Scrapy + distributed queue, or custom service

For very large scale, a queue-based or stream-processing architecture is usually best.


3) Separate fetching from parsing

Do not parse deeply inside your fetcher if volume is high.

Better flow

  • Fetcher downloads HTML
  • Store raw response immediately
  • Emit a job/event for parsing
  • Parser extracts fields and links
  • Links go back to the frontier

This separation gives you:

  • retries without refetching everything
  • easier scaling
  • ability to re-parse old HTML when extraction logic changes

4) Use deduplication aggressively

Millions of pages means lots of duplicates and near-duplicates.

URL dedup

Normalize URLs:

  • lowercase host
  • remove fragments
  • sort/remove tracking parameters
  • canonicalize trailing slashes
  • normalize default ports

Content dedup

Use hashes:

  • exact hash of raw content
  • near-duplicate detection via SimHash / MinHash if needed

This saves bandwidth, storage, and parsing time.


5) Be polite and resilient

At scale, bad crawl behavior gets you blocked quickly.

Respect:

  • robots.txt
  • rate limits per domain
  • crawl delays
  • user-agent identification

Robust fetching:

  • timeouts
  • retries with exponential backoff
  • redirect handling
  • connection pooling
  • compressed responses
  • per-domain throttling

A good crawler is “many domains, but slow per domain.”


6) Parse with tools suited to scale

For HTML:

  • lxml: fast and robust
  • selectolax: very fast parser
  • BeautifulSoup: easy, but slower
  • Go/Rust parsers: good if you need maximum throughput

For extracting structured data:

  • XPath / CSS selectors
  • regex only for small, well-defined patterns
  • DOM cleanup before extraction if pages are messy

For millions of pages:

Prefer deterministic selectors and simple extraction rules over complex heuristics when possible.


7) Store raw and extracted data separately

Raw layer

Store fetched pages in:

  • object storage (S3/GCS/Azure Blob)
  • compressed format: gzip/zstd
  • partitioned by date/domain/source

Structured layer

Store parsed output in:

  • Parquet for analytics
  • JSONL for simplicity
  • database if you need transactional access

A common design:

  • raw HTML in object storage
  • metadata in a relational DB
  • parsed records in Parquet on a data lake

8) Design for incremental crawling

Don’t recrawl everything every time.

Track:

  • last fetched time
  • HTTP status
  • ETag / Last-Modified
  • content hash
  • change frequency per domain/page

Then do:

  • frequent recrawls for fast-changing pages
  • slow recrawls for stable pages
  • conditional requests where possible (If-Modified-Since, If-None-Match)

This cuts costs dramatically.


9) Handle malformed and heterogeneous pages

At scale, expect:

  • broken HTML
  • mixed encodings
  • mobile/desktop variants
  • anti-bot pages
  • soft 404s
  • captchas

Mitigations:

  • encoding detection
  • robust HTML parser recovery
  • content-type checks
  • page classification before extraction
  • fallback rules for known templates

10) Measure everything

You need observability.

Track:

  • fetch success rate
  • latency by domain
  • HTTP status codes
  • parse success rate
  • extraction coverage
  • duplicate rate
  • queue depth
  • retry rate
  • ban/captcha rate

Without metrics, large crawling projects drift and fail quietly.


11) A practical stack

Here’s a reasonable stack for millions of pages:

Small-to-medium scale

  • Python + Scrapy
  • Redis or RabbitMQ for queueing
  • PostgreSQL for metadata
  • S3 for raw HTML
  • Parquet for output

Larger scale

  • Kafka or SQS for scheduling
  • Multiple fetcher services in containers
  • lxml/selectolax parsers
  • S3 + Parquet
  • Spark/Flink/Beam for downstream processing
  • Prometheus/Grafana for monitoring

12) Sample pipeline design

A simple production architecture:

  1. Seed URL list loaded into queue
  2. Fetcher pulls URL
  3. Applies domain politeness and robots check
  4. Downloads page
  5. Saves raw response to storage
  6. Emits parse job
  7. Parser extracts:
    • target fields
    • outgoing links
    • canonical URL
    • metadata
  8. New links are normalized and deduped
  9. Unseen links are pushed back to queue
  10. Metrics collected at every step

13) If you’re just starting, don’t overbuild

Start with:

  • a small seed set
  • one worker
  • URL normalization
  • raw storage
  • basic parser
  • deduplication
  • retry logic

Then scale horizontally.

Most “millions of pages” projects fail because they start with browser automation everywhere or try to parse everything in one script.


14) Important legal/ethical note

Crawl only where you have permission or a legitimate basis to do so. Respect site policies, robots rules where applicable, rate limits, and privacy/security constraints. Avoid collecting sensitive personal data unless you have a lawful reason and a compliance process.


Recommended implementation approach

If you want a concrete path:

  • Use Scrapy for crawling if pages are mostly static
  • Use S3 + JSONL/Parquet for storage
  • Add Redis/Kafka for distributed scheduling
  • Use lxml/selectolax for fast parsing
  • Add domain-level throttling and URL canonicalization
  • Build monitoring + dedup + retries from day one

If you want, I can give you:

  1. a reference architecture diagram,
  2. a Scrapy-based implementation, or
  3. a Python pseudocode crawler for millions of pages.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.