Prompt

Playwright crawler for thousands of pages

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

If you need a Playwright crawler for thousands of pages, the main challenge is usually not “how do I click and navigate?” but:

  1. Throughput — how many pages per minute you can process
  2. Stability — avoiding crashes, timeouts, memory leaks
  3. Politeness — rate limiting, retries, backoff
  4. Scalability — queueing and concurrency control

Here’s a solid pattern to use.


Recommended architecture

Use:

  • One browser instance
  • A small pool of pages/contexts
  • A job queue
  • Concurrency limit
  • Retries + timeout handling
  • Persistence for visited URLs / results

For thousands of pages, avoid launching a new browser per page. That’s usually too slow and memory-heavy.


Good default setup

  • Browser: Chromium
  • Use headless
  • Concurrency: start with 3–10 pages at a time
  • Reuse contexts if possible, but isolate if needed
  • Add randomized delays if the target site is sensitive

Example crawler in Node.js

const { chromium } = require('playwright');

const START_URLS = [
  'https://example.com/page1',
  'https://example.com/page2',
  // ...
];

const CONCURRENCY = 5;
const NAV_TIMEOUT = 30000;

async function crawlPage(page, url) {
  try {
    await page.goto(url, { waitUntil: 'domcontentloaded', timeout: NAV_TIMEOUT });

    // Optional: wait for some selector if needed
    // await page.waitForSelector('main', { timeout: 10000 });

    const title = await page.title();
    const content = await page.textContent('body');

    console.log({ url, title, contentLength: content?.length || 0 });

    return { url, title, content };
  } catch (err) {
    console.error(`Failed: ${url}`, err.message);
    return { url, error: err.message };
  }
}

async function run() {
  const browser = await chromium.launch({ headless: true });
  const results = [];

  let index = 0;
  async function worker(workerId) {
    const context = await browser.newContext();
    const page = await context.newPage();

    while (index < START_URLS.length) {
      const current = index++;
      const url = START_URLS[current];

      const result = await crawlPage(page, url);
      results.push(result);

      // small delay to avoid hammering the server
      await page.waitForTimeout(200);
    }

    await context.close();
  }

  const workers = Array.from({ length: CONCURRENCY }, (_, i) => worker(i));
  await Promise.all(workers);

  await browser.close();
  return results;
}

run().catch(err => {
  console.error(err);
  process.exit(1);
});

Important improvements for large-scale crawling

1. Use a queue instead of shared index

The shared index++ pattern is simple, but for production use a proper queue or task system.

Good options:

  • Bottleneck
  • p-queue
  • BullMQ if you need persistence/distributed workers
  • A database-backed queue for fault tolerance

2. Retry failed pages

Network issues happen. Add 2–3 retries with exponential backoff.

Pseudo-logic:

  • Try page
  • If timeout / transient error:
    • wait 1s
    • retry
    • wait 2s
    • retry
  • If still failing, mark failed

3. Avoid memory leaks

For large crawls:

  • Reuse pages, but periodically recreate them
  • Close popups / extra tabs
  • Avoid storing huge DOM/text in memory unless needed
  • Keep results in a file/db, not just arrays

A common pattern is:

  • process 100–500 pages
  • refresh page/context/browser if memory grows

4. Use domcontentloaded instead of networkidle unless needed

networkidle can hang on sites with analytics or long-polling.

Prefer:

await page.goto(url, { waitUntil: 'domcontentloaded' });

Then wait for a specific element if required.


5. Block unnecessary resources

If you only need text/HTML, you can speed things up by blocking images, fonts, and media.

await page.route('**/*', route => {
  const resourceType = route.request().resourceType();
  if (['image', 'font', 'media'].includes(resourceType)) {
    route.abort();
  } else {
    route.continue();
  }
});

This can significantly improve speed.


6. Persist checkpoints

For thousands of pages, don’t rely on in-memory state only.

Save:

  • completed URLs
  • failed URLs
  • extracted data
  • timestamps

If the process dies halfway, you can resume.


If you need very high scale

If “thousands” means:

  • 10k–100k pages
  • multiple domains
  • continuous crawling

Then consider:

  • multiple worker processes
  • distributed queue
  • Docker/Kubernetes
  • separate fetch vs parse pipeline

Playwright is great for JS-heavy pages, but if many pages are static, a plain HTTP fetcher is much faster. Use Playwright only where browser rendering is necessary.


Practical tips

  • Set --disable-dev-shm-usage in Docker
  • Increase system file descriptor limits if running many pages
  • Use page.setExtraHTTPHeaders() for consistent crawling
  • Set a realistic user agent if needed
  • Handle captcha / anti-bot carefully and legally

Best pattern summary

For thousands of pages:

  • 1 browser
  • N concurrent contexts/pages
  • queue-based scheduler
  • retry logic
  • resource blocking
  • checkpoints
  • periodic recycling

If you want, I can give you one of these next:

  1. a production-ready Playwright crawler template
  2. a version with Bottleneck/p-queue
  3. a distributed crawler with BullMQ
  4. a crawler that extracts structured data from pages

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.