Prompt

What should I use to crawl a million pages a month?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

For 1 million pages/month, the best choice depends on how heavy the pages are, how fresh the data must be, and whether you need rendering/login/JS. Roughly, that’s only about 23 pages/minute on average, so the scale itself isn’t huge—but reliability matters more than raw speed.

Good options

1. Simple static pages: Python + Scrapy

Best if pages are mostly HTML and don’t require JavaScript rendering.

  • Very efficient and lightweight
  • Built-in throttling, retries, deduping
  • Easy to run on a single VM or a small cluster
  • Good ecosystem for parsing/storing data

Use this if: you want control, low cost, and pages are crawl-friendly.


2. JavaScript-heavy sites: Playwright (or Puppeteer) + queue

Best if pages need browser rendering.

  • Handles modern sites well
  • More resource-intensive than Scrapy
  • Usually pair with a task queue like Celery, RabbitMQ, SQS, or Kafka

Use this if: you need to render pages, interact with buttons, or handle anti-bot-ish frontends.


3. Managed crawling platform: Apify, Zyte, Bright Data, Oxylabs

Best if you want less ops work.

  • Proxy management included
  • Scaling and retries are easier
  • Often useful for hard targets
  • Higher cost, less control

Use this if: you care more about reliability and speed to production than running infrastructure yourself.


My practical recommendation

If you’re starting from scratch:

  • Static pages: Scrapy + PostgreSQL/S3 + Redis queue
  • Dynamic pages: Playwright + queue + rotating proxies
  • At scale / hard sites: managed service like Zyte or Bright Data

Infrastructure estimate

For 1M pages/month, you can often do this with:

  • 1–3 modest servers for static crawling
  • More if using browsers (Playwright can be 5–20x heavier)
  • Add:
    • Redis for scheduling/dupe filtering
    • PostgreSQL for metadata
    • S3/GCS for raw HTML if needed

Important caveats

Make sure you consider:

  • robots.txt / site terms
  • rate limiting and politeness
  • retries and backoff
  • deduplication
  • monitoring for blocks/errors

If you want, I can recommend a setup based on:

  1. static vs JS pages,
  2. how many domains, and
  3. whether you need full HTML or just extracted fields.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.