Prompt
Playwright crawler for thousands of pages
Latest observation
If you need a Playwright crawler for thousands of pages, the main challenge is usually not “how do I click and navigate?” but:
- Throughput — how many pages per minute you can process
- Stability — avoiding crashes, timeouts, memory leaks
- Politeness — rate limiting, retries, backoff
- Scalability — queueing and concurrency control
Here’s a solid pattern to use.
Recommended architecture
Use:
- One browser instance
- A small pool of pages/contexts
- A job queue
- Concurrency limit
- Retries + timeout handling
- Persistence for visited URLs / results
For thousands of pages, avoid launching a new browser per page. That’s usually too slow and memory-heavy.
Good default setup
- Browser: Chromium
- Use headless
- Concurrency: start with 3–10 pages at a time
- Reuse contexts if possible, but isolate if needed
- Add randomized delays if the target site is sensitive
Example crawler in Node.js
const { chromium } = require('playwright');
const START_URLS = [
'https://example.com/page1',
'https://example.com/page2',
// ...
];
const CONCURRENCY = 5;
const NAV_TIMEOUT = 30000;
async function crawlPage(page, url) {
try {
await page.goto(url, { waitUntil: 'domcontentloaded', timeout: NAV_TIMEOUT });
// Optional: wait for some selector if needed
// await page.waitForSelector('main', { timeout: 10000 });
const title = await page.title();
const content = await page.textContent('body');
console.log({ url, title, contentLength: content?.length || 0 });
return { url, title, content };
} catch (err) {
console.error(`Failed: ${url}`, err.message);
return { url, error: err.message };
}
}
async function run() {
const browser = await chromium.launch({ headless: true });
const results = [];
let index = 0;
async function worker(workerId) {
const context = await browser.newContext();
const page = await context.newPage();
while (index < START_URLS.length) {
const current = index++;
const url = START_URLS[current];
const result = await crawlPage(page, url);
results.push(result);
// small delay to avoid hammering the server
await page.waitForTimeout(200);
}
await context.close();
}
const workers = Array.from({ length: CONCURRENCY }, (_, i) => worker(i));
await Promise.all(workers);
await browser.close();
return results;
}
run().catch(err => {
console.error(err);
process.exit(1);
});
Important improvements for large-scale crawling
1. Use a queue instead of shared index
The shared index++ pattern is simple, but for production use a proper queue or task system.
Good options:
- Bottleneck
- p-queue
- BullMQ if you need persistence/distributed workers
- A database-backed queue for fault tolerance
2. Retry failed pages
Network issues happen. Add 2–3 retries with exponential backoff.
Pseudo-logic:
- Try page
- If timeout / transient error:
- wait 1s
- retry
- wait 2s
- retry
- If still failing, mark failed
3. Avoid memory leaks
For large crawls:
- Reuse pages, but periodically recreate them
- Close popups / extra tabs
- Avoid storing huge DOM/text in memory unless needed
- Keep results in a file/db, not just arrays
A common pattern is:
- process 100–500 pages
- refresh page/context/browser if memory grows
4. Use domcontentloaded instead of networkidle unless needed
networkidle can hang on sites with analytics or long-polling.
Prefer:
await page.goto(url, { waitUntil: 'domcontentloaded' });
Then wait for a specific element if required.
5. Block unnecessary resources
If you only need text/HTML, you can speed things up by blocking images, fonts, and media.
await page.route('**/*', route => {
const resourceType = route.request().resourceType();
if (['image', 'font', 'media'].includes(resourceType)) {
route.abort();
} else {
route.continue();
}
});
This can significantly improve speed.
6. Persist checkpoints
For thousands of pages, don’t rely on in-memory state only.
Save:
- completed URLs
- failed URLs
- extracted data
- timestamps
If the process dies halfway, you can resume.
If you need very high scale
If “thousands” means:
- 10k–100k pages
- multiple domains
- continuous crawling
Then consider:
- multiple worker processes
- distributed queue
- Docker/Kubernetes
- separate fetch vs parse pipeline
Playwright is great for JS-heavy pages, but if many pages are static, a plain HTTP fetcher is much faster. Use Playwright only where browser rendering is necessary.
Practical tips
- Set
--disable-dev-shm-usagein Docker - Increase system file descriptor limits if running many pages
- Use
page.setExtraHTTPHeaders()for consistent crawling - Set a realistic user agent if needed
- Handle captcha / anti-bot carefully and legally
Best pattern summary
For thousands of pages:
- 1 browser
- N concurrent contexts/pages
- queue-based scheduler
- retry logic
- resource blocking
- checkpoints
- periodic recycling
If you want, I can give you one of these next:
- a production-ready Playwright crawler template
- a version with Bottleneck/p-queue
- a distributed crawler with BullMQ
- a crawler that extracts structured data from pages