Prompt
JavaScript crawling at scale
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
If you mean crawling websites at scale with JavaScript, the main challenge is that many sites render content dynamically, so you need to balance coverage, speed, reliability, and cost.
Key approaches
1) Prefer HTTP fetching first
- Use plain
fetch,axios, orgotto request pages directly. - Parse HTML with
cheerio,parse5, orjsdom. - This is much faster and cheaper than a browser.
Best when:
- Content is server-rendered
- You only need links, metadata, or static content
2) Use a headless browser only when needed
- Use Playwright or Puppeteer for JS-heavy pages.
- Browser crawling is expensive, so reserve it for:
- React/Vue/SPA pages
- Content loaded after navigation
- Pages requiring interactions
3) Build a queue-based crawler
At scale, avoid recursive crawling. Use:
- A URL frontier / queue
- A dedupe store for visited URLs
- Per-domain rate limits
- Retry logic and backoff
Common pattern:
- Producer adds URLs to queue
- Workers fetch pages
- Extract new URLs
- Normalize + dedupe
- Enqueue discovered URLs
4) Normalize URLs carefully
To avoid duplicates:
- Remove fragments:
#section - Handle trailing slashes consistently
- Resolve relative URLs
- Canonicalize query params if appropriate
- Respect
rel="canonical"when useful
5) Be polite and compliant
- Check
robots.txt - Honor crawl delays where applicable
- Identify your crawler with a clear User-Agent
- Throttle requests per host
- Avoid abusive parallelism
Scaling considerations
Concurrency
- Start with low concurrency per domain
- Scale globally, not per-site, unless you control the target
- In Node.js, use worker pools and async I/O, not blocking loops
Storage
- Store:
- URL
- fetch status
- fetch time
- content hash
- extracted links
- metadata
- Use a database or queue system:
- Redis
- PostgreSQL
- Kafka / RabbitMQ / SQS for larger systems
Fault tolerance
- Expect:
- timeouts
- 429s
- 503s
- transient DNS failures
- Use retries with exponential backoff
- Persist crawl state so jobs can resume
Rendering strategy
A practical pattern:
- Try plain HTTP fetch
- Detect if content is insufficient
- Escalate to browser rendering only for those pages
Useful libraries in JavaScript
- Playwright — best modern browser automation
- Puppeteer — popular Chromium automation
- Cheerio — fast HTML parsing
- got / axios / native
fetch— HTTP requests - BullMQ — Redis-backed job queue
- p-limit — concurrency control
- robots-parser — robots.txt handling
- Bottleneck — rate limiting
Example architecture
- Scheduler: seeds and prioritizes URLs
- Fetcher workers: HTTP-first, browser fallback
- Parser: extracts links/content
- Normalizer + deduper: prevents revisits
- Storage: saves content and crawl state
- Rate limiter: per domain / per IP
- Monitoring: metrics, logs, failure alerts
Example: minimal HTTP crawler in Node.js
import * as cheerio from 'cheerio';
const seen = new Set();
const queue = ['https://example.com'];
while (queue.length) {
const url = queue.shift();
if (seen.has(url)) continue;
seen.add(url);
try {
const res = await fetch(url, {
headers: { 'User-Agent': 'MyCrawler/1.0' }
});
if (!res.ok) continue;
const html = await res.text();
const $ = cheerio.load(html);
$('a[href]').each((_, el) => {
const href = $(el).attr('href');
try {
const next = new URL(href, url).toString();
if (!seen.has(next)) queue.push(next);
} catch {}
});
console.log('Crawled:', url);
} catch (err) {
console.error('Failed:', url, err.message);
}
}
This works for small jobs, but for scale you’d want:
- persistent queue
- dedupe in Redis/DB
- concurrency limits
- retries and rate limiting
If you want, I can also provide:
- a production-grade Node.js crawler architecture
- a Playwright-based crawler with queue + dedupe
- a multi-worker design using Redis/BullMQ
- a robots.txt-aware crawler template