Prompt
How to crawl JavaScript sites with retries and queues
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
To crawl JavaScript-heavy sites reliably, you usually want a two-layer system:
- A queue to manage crawl jobs/pages and control concurrency
- Retries to handle transient failures like timeouts, 429s, flaky rendering, or navigation errors
A common setup is:
- Crawler: Playwright or Puppeteer for JS rendering
- Queue: BullMQ, RabbitMQ, SQS, Redis-based queue, or a simple in-memory queue for small jobs
- Retry logic: exponential backoff + max attempts
- Deduplication: avoid crawling the same URL repeatedly
- Persistence: store visited URLs and results in DB
Recommended architecture
1) Use a queue for URLs
Each job represents:
- URL
- retry count
- priority
- depth
- parent URL or source
Example job payload:
{
"url": "https://example.com/products",
"depth": 1,
"attempt": 0
}
2) Use a headless browser
For JS sites, fetch() alone often won’t work because content is rendered client-side.
Use:
- Playwright (recommended)
- Puppeteer
- Selenium only if necessary
3) Add retries with backoff
Retry on:
- timeout
- navigation failure
- browser disconnect
- HTTP 429/5xx
- empty/partial content when expected content not loaded
Use exponential backoff:
- attempt 1: immediate
- attempt 2: wait 1s
- attempt 3: wait 2s
- attempt 4: wait 4s
4) Limit concurrency
Too many browser tabs/pages will:
- slow everything down
- trigger bot defenses
- cause memory issues
Typical starting point:
- 2–5 browser instances
- 5–10 pages total depending on host
5) Extract new links and enqueue them
After rendering a page:
- collect links
- normalize them
- filter by domain/rules
- enqueue unseen links
Example with Node.js + Playwright + a simple queue
Basic idea
- Pull next URL from queue
- Open page
- Wait for content to load
- Scrape data
- Enqueue discovered links
- Retry if failure
Pseudocode
while (queue not empty) {
job = dequeue()
try {
result = crawlPage(job.url)
save(result)
enqueueLinks(result.links)
} catch (err) {
if (job.attempt < maxRetries) {
requeue(job, backoff(job.attempt))
} else {
markFailed(job)
}
}
}
Playwright example with retries
import { chromium } from 'playwright';
async function sleep(ms) {
return new Promise(resolve => setTimeout(resolve, ms));
}
async function crawlPage(browser, url, attempt = 0, maxRetries = 3) {
const page = await browser.newPage();
try {
await page.goto(url, { waitUntil: 'networkidle', timeout: 30000 });
// Optional: wait for a key selector that indicates the page is ready
// await page.waitForSelector('.product-card', { timeout: 10000 });
const title = await page.title();
const links = await page.$$eval('a[href]', as =>
as.map(a => a.href)
);
await page.close();
return { url, title, links };
} catch (err) {
await page.close().catch(() => {});
if (attempt < maxRetries) {
const delay = Math.pow(2, attempt) * 1000;
await sleep(delay);
return crawlPage(browser, url, attempt + 1, maxRetries);
}
throw err;
}
}
async function main() {
const browser = await chromium.launch({ headless: true });
const seedUrls = ['https://example.com'];
const visited = new Set();
const queue = [...seedUrls];
while (queue.length > 0) {
const url = queue.shift();
if (visited.has(url)) continue;
visited.add(url);
try {
const result = await crawlPage(browser, url);
console.log('Crawled:', result.url, result.title);
// enqueue discovered links
for (const link of result.links) {
if (link.startsWith('https://example.com') && !visited.has(link)) {
queue.push(link);
}
}
} catch (err) {
console.error('Failed:', url, err.message);
}
}
await browser.close();
}
main();
Better: use a real queue system
If you want scale, use something like BullMQ + Redis.
Why BullMQ?
- persistent jobs
- built-in retries
- delayed jobs
- concurrency control
- easy backoff policies
Example retry settings
{
attempts: 5,
backoff: {
type: 'exponential',
delay: 1000
}
}
BullMQ-style flow
- Add seed URLs as jobs
- Worker processes job with Playwright
- On success, enqueue discovered links
- On failure, BullMQ retries automatically
Worker example
import { Worker, Queue } from 'bullmq';
import { chromium } from 'playwright';
const queue = new Queue('crawl', { connection: { host: '127.0.0.1', port: 6379 } });
const worker = new Worker('crawl', async job => {
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
try {
await page.goto(job.data.url, { waitUntil: 'networkidle', timeout: 30000 });
const title = await page.title();
const links = await page.$$eval('a[href]', as => as.map(a => a.href));
for (const link of links) {
await queue.add('crawl', { url: link }, {
attempts: 5,
backoff: { type: 'exponential', delay: 1000 },
removeOnComplete: true,
removeOnFail: false
});
}
await browser.close();
return { title };
} catch (err) {
await browser.close();
throw err;
}
}, {
connection: { host: '127.0.0.1', port: 6379 },
concurrency: 3
});
Key retry rules
Retry only transient errors
Good retry candidates:
- timeout
- DNS hiccup
- 429 rate limit
- 5xx errors
- browser disconnect
- navigation interrupted
Do not retry blindly on:
- 404
- authentication errors
- blocked access due to policy
- invalid URLs
Best practices for JS crawling
1) Use request interception
Block heavy resources to speed up crawling:
- images
- video
- fonts
- ads
- analytics
Example:
await page.route('**/*', route => {
const type = route.request().resourceType();
if (['image', 'font', 'media'].includes(type)) {
return route.abort();
}
route.continue();
});
2) Wait for the right signal
Don’t always rely on networkidle. Sometimes SPAs keep polling forever.
Instead:
- wait for a specific element
- wait for content length
- wait for a known API response
3) Deduplicate URLs
Normalize:
- remove hash fragments
- sort query params if needed
- strip tracking params like
utm_*
4) Handle infinite scroll
For pages that load content on scroll:
- scroll incrementally
- detect new content
- stop when no more changes
5) Persist visited state
Use Redis or a DB so restarts don’t recrawl everything.
Practical recommendation
If you’re starting from scratch:
- Playwright for rendering
- BullMQ + Redis for queueing and retries
- Postgres or Redis for visited URL tracking
- Exponential backoff retries
- Concurrency limit of 2–5 initially
If you want, I can give you:
- a production-ready Node.js crawler with BullMQ + Playwright, or
- a Python version using Playwright + Redis/Celery.