Prompt

JavaScript crawling at scale

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

If you mean crawling websites at scale with JavaScript, the main challenge is that many sites render content dynamically, so you need to balance coverage, speed, reliability, and cost.

Key approaches

1) Prefer HTTP fetching first

  • Use plain fetch, axios, or got to request pages directly.
  • Parse HTML with cheerio, parse5, or jsdom.
  • This is much faster and cheaper than a browser.

Best when:

  • Content is server-rendered
  • You only need links, metadata, or static content

2) Use a headless browser only when needed

  • Use Playwright or Puppeteer for JS-heavy pages.
  • Browser crawling is expensive, so reserve it for:
    • React/Vue/SPA pages
    • Content loaded after navigation
    • Pages requiring interactions

3) Build a queue-based crawler

At scale, avoid recursive crawling. Use:

  • A URL frontier / queue
  • A dedupe store for visited URLs
  • Per-domain rate limits
  • Retry logic and backoff

Common pattern:

  • Producer adds URLs to queue
  • Workers fetch pages
  • Extract new URLs
  • Normalize + dedupe
  • Enqueue discovered URLs

4) Normalize URLs carefully

To avoid duplicates:

  • Remove fragments: #section
  • Handle trailing slashes consistently
  • Resolve relative URLs
  • Canonicalize query params if appropriate
  • Respect rel="canonical" when useful

5) Be polite and compliant

  • Check robots.txt
  • Honor crawl delays where applicable
  • Identify your crawler with a clear User-Agent
  • Throttle requests per host
  • Avoid abusive parallelism

Scaling considerations

Concurrency

  • Start with low concurrency per domain
  • Scale globally, not per-site, unless you control the target
  • In Node.js, use worker pools and async I/O, not blocking loops

Storage

  • Store:
    • URL
    • fetch status
    • fetch time
    • content hash
    • extracted links
    • metadata
  • Use a database or queue system:
    • Redis
    • PostgreSQL
    • Kafka / RabbitMQ / SQS for larger systems

Fault tolerance

  • Expect:
    • timeouts
    • 429s
    • 503s
    • transient DNS failures
  • Use retries with exponential backoff
  • Persist crawl state so jobs can resume

Rendering strategy

A practical pattern:

  1. Try plain HTTP fetch
  2. Detect if content is insufficient
  3. Escalate to browser rendering only for those pages

Useful libraries in JavaScript

  • Playwright — best modern browser automation
  • Puppeteer — popular Chromium automation
  • Cheerio — fast HTML parsing
  • got / axios / native fetch — HTTP requests
  • BullMQ — Redis-backed job queue
  • p-limit — concurrency control
  • robots-parser — robots.txt handling
  • Bottleneck — rate limiting

Example architecture

  • Scheduler: seeds and prioritizes URLs
  • Fetcher workers: HTTP-first, browser fallback
  • Parser: extracts links/content
  • Normalizer + deduper: prevents revisits
  • Storage: saves content and crawl state
  • Rate limiter: per domain / per IP
  • Monitoring: metrics, logs, failure alerts

Example: minimal HTTP crawler in Node.js

import * as cheerio from 'cheerio';

const seen = new Set();
const queue = ['https://example.com'];

while (queue.length) {
  const url = queue.shift();
  if (seen.has(url)) continue;
  seen.add(url);

  try {
    const res = await fetch(url, {
      headers: { 'User-Agent': 'MyCrawler/1.0' }
    });
    if (!res.ok) continue;

    const html = await res.text();
    const $ = cheerio.load(html);

    $('a[href]').each((_, el) => {
      const href = $(el).attr('href');
      try {
        const next = new URL(href, url).toString();
        if (!seen.has(next)) queue.push(next);
      } catch {}
    });

    console.log('Crawled:', url);
  } catch (err) {
    console.error('Failed:', url, err.message);
  }
}

This works for small jobs, but for scale you’d want:

  • persistent queue
  • dedupe in Redis/DB
  • concurrency limits
  • retries and rate limiting

If you want, I can also provide:

  1. a production-grade Node.js crawler architecture
  2. a Playwright-based crawler with queue + dedupe
  3. a multi-worker design using Redis/BullMQ
  4. a robots.txt-aware crawler template

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.