Prompt

How to crawl JavaScript sites with retries and queues

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To crawl JavaScript-heavy sites reliably, you usually want a two-layer system:

  1. A queue to manage crawl jobs/pages and control concurrency
  2. Retries to handle transient failures like timeouts, 429s, flaky rendering, or navigation errors

A common setup is:

  • Crawler: Playwright or Puppeteer for JS rendering
  • Queue: BullMQ, RabbitMQ, SQS, Redis-based queue, or a simple in-memory queue for small jobs
  • Retry logic: exponential backoff + max attempts
  • Deduplication: avoid crawling the same URL repeatedly
  • Persistence: store visited URLs and results in DB

Recommended architecture

1) Use a queue for URLs

Each job represents:

  • URL
  • retry count
  • priority
  • depth
  • parent URL or source

Example job payload:

{
  "url": "https://example.com/products",
  "depth": 1,
  "attempt": 0
}

2) Use a headless browser

For JS sites, fetch() alone often won’t work because content is rendered client-side.

Use:

  • Playwright (recommended)
  • Puppeteer
  • Selenium only if necessary

3) Add retries with backoff

Retry on:

  • timeout
  • navigation failure
  • browser disconnect
  • HTTP 429/5xx
  • empty/partial content when expected content not loaded

Use exponential backoff:

  • attempt 1: immediate
  • attempt 2: wait 1s
  • attempt 3: wait 2s
  • attempt 4: wait 4s

4) Limit concurrency

Too many browser tabs/pages will:

  • slow everything down
  • trigger bot defenses
  • cause memory issues

Typical starting point:

  • 2–5 browser instances
  • 5–10 pages total depending on host

5) Extract new links and enqueue them

After rendering a page:

  • collect links
  • normalize them
  • filter by domain/rules
  • enqueue unseen links

Example with Node.js + Playwright + a simple queue

Basic idea

  • Pull next URL from queue
  • Open page
  • Wait for content to load
  • Scrape data
  • Enqueue discovered links
  • Retry if failure

Pseudocode

while (queue not empty) {
  job = dequeue()
  try {
    result = crawlPage(job.url)
    save(result)
    enqueueLinks(result.links)
  } catch (err) {
    if (job.attempt < maxRetries) {
      requeue(job, backoff(job.attempt))
    } else {
      markFailed(job)
    }
  }
}

Playwright example with retries

import { chromium } from 'playwright';

async function sleep(ms) {
  return new Promise(resolve => setTimeout(resolve, ms));
}

async function crawlPage(browser, url, attempt = 0, maxRetries = 3) {
  const page = await browser.newPage();

  try {
    await page.goto(url, { waitUntil: 'networkidle', timeout: 30000 });

    // Optional: wait for a key selector that indicates the page is ready
    // await page.waitForSelector('.product-card', { timeout: 10000 });

    const title = await page.title();
    const links = await page.$$eval('a[href]', as =>
      as.map(a => a.href)
    );

    await page.close();
    return { url, title, links };
  } catch (err) {
    await page.close().catch(() => {});

    if (attempt < maxRetries) {
      const delay = Math.pow(2, attempt) * 1000;
      await sleep(delay);
      return crawlPage(browser, url, attempt + 1, maxRetries);
    }

    throw err;
  }
}

async function main() {
  const browser = await chromium.launch({ headless: true });

  const seedUrls = ['https://example.com'];
  const visited = new Set();
  const queue = [...seedUrls];

  while (queue.length > 0) {
    const url = queue.shift();
    if (visited.has(url)) continue;
    visited.add(url);

    try {
      const result = await crawlPage(browser, url);

      console.log('Crawled:', result.url, result.title);

      // enqueue discovered links
      for (const link of result.links) {
        if (link.startsWith('https://example.com') && !visited.has(link)) {
          queue.push(link);
        }
      }
    } catch (err) {
      console.error('Failed:', url, err.message);
    }
  }

  await browser.close();
}

main();

Better: use a real queue system

If you want scale, use something like BullMQ + Redis.

Why BullMQ?

  • persistent jobs
  • built-in retries
  • delayed jobs
  • concurrency control
  • easy backoff policies

Example retry settings

{
  attempts: 5,
  backoff: {
    type: 'exponential',
    delay: 1000
  }
}

BullMQ-style flow

  1. Add seed URLs as jobs
  2. Worker processes job with Playwright
  3. On success, enqueue discovered links
  4. On failure, BullMQ retries automatically

Worker example

import { Worker, Queue } from 'bullmq';
import { chromium } from 'playwright';

const queue = new Queue('crawl', { connection: { host: '127.0.0.1', port: 6379 } });

const worker = new Worker('crawl', async job => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage();

  try {
    await page.goto(job.data.url, { waitUntil: 'networkidle', timeout: 30000 });

    const title = await page.title();
    const links = await page.$$eval('a[href]', as => as.map(a => a.href));

    for (const link of links) {
      await queue.add('crawl', { url: link }, {
        attempts: 5,
        backoff: { type: 'exponential', delay: 1000 },
        removeOnComplete: true,
        removeOnFail: false
      });
    }

    await browser.close();
    return { title };
  } catch (err) {
    await browser.close();
    throw err;
  }
}, {
  connection: { host: '127.0.0.1', port: 6379 },
  concurrency: 3
});

Key retry rules

Retry only transient errors

Good retry candidates:

  • timeout
  • DNS hiccup
  • 429 rate limit
  • 5xx errors
  • browser disconnect
  • navigation interrupted

Do not retry blindly on:

  • 404
  • authentication errors
  • blocked access due to policy
  • invalid URLs

Best practices for JS crawling

1) Use request interception

Block heavy resources to speed up crawling:

  • images
  • video
  • fonts
  • ads
  • analytics

Example:

await page.route('**/*', route => {
  const type = route.request().resourceType();
  if (['image', 'font', 'media'].includes(type)) {
    return route.abort();
  }
  route.continue();
});

2) Wait for the right signal

Don’t always rely on networkidle. Sometimes SPAs keep polling forever.

Instead:

  • wait for a specific element
  • wait for content length
  • wait for a known API response

3) Deduplicate URLs

Normalize:

  • remove hash fragments
  • sort query params if needed
  • strip tracking params like utm_*

4) Handle infinite scroll

For pages that load content on scroll:

  • scroll incrementally
  • detect new content
  • stop when no more changes

5) Persist visited state

Use Redis or a DB so restarts don’t recrawl everything.


Practical recommendation

If you’re starting from scratch:

  • Playwright for rendering
  • BullMQ + Redis for queueing and retries
  • Postgres or Redis for visited URL tracking
  • Exponential backoff retries
  • Concurrency limit of 2–5 initially

If you want, I can give you:

  1. a production-ready Node.js crawler with BullMQ + Playwright, or
  2. a Python version using Playwright + Redis/Celery.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.