Prompt
Bright Data web crawler setup
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
Here’s a practical Bright Data web crawler setup you can use as a starting point.
1) Decide what type of crawler you need
Bright Data is usually used in one of these ways:
- Web Unlocker / Scraping Browser for JavaScript-heavy sites
- Residential / Datacenter proxies for custom crawlers
- Bright Data Collector for no-code/low-code extraction
- Browser automation with Playwright / Selenium + Bright Data proxies
If you’re building your own crawler, the most common approach is:
- Playwright or Puppeteer
- routed through Bright Data proxies
- with retry, throttling, and deduplication
2) Create your Bright Data setup
In the Bright Data dashboard:
- Create an account
- Choose a product:
- Proxy network or Web Unlocker
- Create a zone
- Note:
- username
- password
- proxy host
- proxy port
Example proxy endpoint often looks like:
brd.superproxy.io:22225
3) Basic proxy configuration
Example: Node.js with Playwright
Install:
npm install playwright
Example crawler:
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch({
headless: true,
proxy: {
server: 'http://brd.superproxy.io:22225',
username: 'YOUR_BRIGHT_DATA_USERNAME',
password: 'YOUR_BRIGHT_DATA_PASSWORD'
}
});
const page = await browser.newPage();
await page.goto('https://example.com', { waitUntil: 'networkidle' });
console.log(await page.title());
await browser.close();
})();
4) Add crawl logic
A crawler should usually include:
- URL queue
- visited set / dedupe
- rate limiting
- retry with backoff
- HTML extraction
- error handling
Example simplified flow:
const urls = [
'https://example.com',
'https://example.com/about'
];
for (const url of urls) {
try {
await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
const content = await page.textContent('body');
console.log(url, content.slice(0, 200));
} catch (err) {
console.error('Failed:', url, err.message);
}
}
5) If using Bright Data Web Unlocker
If the site is difficult to scrape, Web Unlocker may be easier than managing your own browser automation.
Typical pattern:
- Send requests through Bright Data’s unlocker endpoint
- Let Bright Data handle rotation, CAPTCHA, and blocking
You’ll usually configure:
- target URL
- zone
- headers / geo / render options depending on your plan
6) Best practices
- Respect site terms and robots.txt where applicable
- Use delays between requests
- Rotate session identifiers when needed
- Keep concurrency low at first
- Log status codes and failures
- Store crawled data incrementally
7) Example architecture for a real crawler
A solid production setup:
- Queue: Redis / RabbitMQ / simple DB table
- Workers: Playwright or HTTP client
- Proxy layer: Bright Data zone
- Storage: PostgreSQL / S3 / Elasticsearch
- Monitoring: logs + retry counts + success rate
8) Troubleshooting tips
If requests fail:
- Verify proxy credentials
- Check zone is active
- Make sure the host/port are correct
- Try different session persistence settings
- Use browser automation instead of raw HTTP for JS-heavy sites
- Inspect response codes and proxy errors
If you want, I can give you one of these next:
- A complete Bright Data crawler in Node.js
- A Python crawler with Playwright
- A Bright Data Web Unlocker example
- A production-ready crawler architecture