Prompt
What should I use for large-scale URL discovery and crawling?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
For large-scale URL discovery and crawling, you usually want a crawler framework or distributed crawler setup, not just a simple HTTP library.
Good options by use case
1. General-purpose large-scale crawling
- Scrapy (Python)
- Great for building crawlers quickly
- Mature ecosystem
- Can be scaled with Scrapy-Redis, Scrapy Cluster, or your own queue system
- Apache Nutch
- Designed for large-scale web crawling
- Built on Hadoop ecosystem
- Good if you already use big-data infrastructure
- StormCrawler
- Java-based, distributed, good for enterprise-scale crawling
2. URL discovery specifically
If you mainly need to find URLs rather than fully crawl content:
- Sitemaps
- Best starting point when available
- Common Crawl
- Huge public web crawl dataset
- Search engine APIs
- For discovery via indexed results
- Focused crawlers built on:
- Scrapy
- Playwright/Puppeteer if the site is JS-heavy
3. JavaScript-heavy sites
- Playwright
- Puppeteer
- Sometimes paired with a crawler queue for scale
4. Very high scale / custom infrastructure
If you’re crawling millions or billions of URLs:
- Use a distributed queue (Kafka, RabbitMQ, Redis, SQS)
- Use workers with:
- HTTP client
- robots.txt handling
- rate limiting / politeness
- deduplication
- canonicalization
- Store discovered URLs in a dedup/index layer like:
- Redis Bloom filter
- Elasticsearch
- PostgreSQL with unique constraints
- Bigtable / Cassandra / DynamoDB at larger scale
Practical recommendation
- If you want the easiest strong default: Scrapy
- If you need distributed crawling at enterprise scale: Apache Nutch or StormCrawler
- If sites are JS-heavy: Playwright + queue-based crawler
- If you need URL discovery at internet scale: use a distributed crawler architecture, not a single tool
What I’d recommend in most cases
A solid stack is:
- Scrapy for crawling logic
- Redis/Kafka/SQS for the URL frontier
- PostgreSQL/Elasticsearch/Redis Bloom for dedup and storage
- Playwright only for pages that truly need rendering
If you want, I can suggest a stack based on:
- your language preference,
- target scale,
- whether the sites are JS-heavy,
- and whether you need just discovery or full content extraction.