Prompt

How do I crawl millions of pages without my own infra?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

If you want to crawl millions of pages without running your own infrastructure, the practical answer is: don’t try to “scrape the web” directly from one machine. Use a combination of managed crawling tools, public datasets, and serverless/batch compute.

Best options

1) Use a managed crawling platform

These handle queueing, retries, scaling, proxies, and sometimes rendering.

Examples:

  • Apify
  • Bright Data Web Scraper / Data Collector
  • Zyte
  • ScrapingBee
  • Oxylabs
  • Common Crawl if you just need already-crawled web data

Good for:

  • large-scale scraping jobs
  • JS-heavy sites
  • proxy management
  • avoiding ops work

Tradeoff:

  • can get expensive at very large scale
  • still need to respect site policies and legal constraints

2) Use serverless / batch compute

If you need custom crawling logic but not your own servers, run the crawler in:

  • AWS Lambda
  • AWS ECS/Fargate
  • Google Cloud Run
  • Google Batch
  • Azure Container Apps
  • Fly.io Machines
  • Render Background Workers
  • Temporal Cloud for orchestration

Typical pattern:

  • a queue of URLs
  • workers pull URLs in parallel
  • store results in object storage / database
  • throttle per domain

Good for:

  • custom logic
  • pay-per-use
  • no server maintenance

Tradeoff:

  • you still design the crawler architecture
  • need careful rate limiting and deduplication

3) Use existing web datasets

If your goal is data extraction rather than live crawling, this is often the cheapest route.

Sources:

  • Common Crawl
  • Internet Archive
  • Wikidata / Wikipedia dumps
  • domain-specific datasets
  • commercial data providers

Good for:

  • huge scale
  • historical snapshots
  • avoiding bot defenses entirely

Tradeoff:

  • data may be stale or incomplete
  • not all sites are represented

A scalable architecture without owning infra

A common approach:

  1. Seed URLs
  2. Put them in a managed queue
  3. Use serverless workers to fetch pages
  4. Store HTML/metadata in S3/GCS/Azure Blob
  5. Extract links and enqueue new URLs
  6. Deduplicate with a URL hash set
  7. Enforce per-domain politeness
  8. Export cleaned data to your warehouse or search index

Managed components you can use:

  • Queue: SQS / PubSub / Azure Queue
  • Storage: S3 / GCS / Blob
  • Compute: Lambda / Cloud Run / Fargate
  • DB for state: DynamoDB / Firestore / Postgres
  • Orchestration: Step Functions / Workflows / Temporal Cloud

Key engineering requirements at million-page scale

Must-haves

  • URL deduplication
  • Retry logic
  • Per-domain rate limiting
  • Backoff on 429/403/5xx
  • Canonicalization of URLs
  • Content fingerprinting
  • Checkpointing
  • Monitoring and cost controls

Nice-to-haves

  • robots.txt awareness
  • JS rendering only when needed
  • language/content filters
  • automatic sitemap discovery
  • proxy rotation if allowed and necessary

What I’d recommend based on your goal

If you want the fastest path

Use Apify or Zyte.

If you want custom control but no servers

Use Cloud Run + a managed queue + object storage.

If you just need data at scale

Use Common Crawl or a commercial dataset.


Important note

Crawling millions of pages can easily run into:

  • site terms of service violations
  • copyright issues
  • anti-bot protections
  • privacy issues

So it’s best to crawl:

  • your own sites
  • sites that allow it
  • public datasets
  • APIs where available

If you want, I can give you:

  1. a cheap architecture on AWS/GCP,
  2. a tool-by-tool comparison, or
  3. a reference crawler design for million-page scale.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.