Prompt

Bright Data web crawler setup

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Here’s a practical Bright Data web crawler setup you can use as a starting point.

1) Decide what type of crawler you need

Bright Data is usually used in one of these ways:

  • Web Unlocker / Scraping Browser for JavaScript-heavy sites
  • Residential / Datacenter proxies for custom crawlers
  • Bright Data Collector for no-code/low-code extraction
  • Browser automation with Playwright / Selenium + Bright Data proxies

If you’re building your own crawler, the most common approach is:

  • Playwright or Puppeteer
  • routed through Bright Data proxies
  • with retry, throttling, and deduplication

2) Create your Bright Data setup

In the Bright Data dashboard:

  1. Create an account
  2. Choose a product:
    • Proxy network or Web Unlocker
  3. Create a zone
  4. Note:
    • username
    • password
    • proxy host
    • proxy port

Example proxy endpoint often looks like:

brd.superproxy.io:22225

3) Basic proxy configuration

Example: Node.js with Playwright

Install:

npm install playwright

Example crawler:

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({
    headless: true,
    proxy: {
      server: 'http://brd.superproxy.io:22225',
      username: 'YOUR_BRIGHT_DATA_USERNAME',
      password: 'YOUR_BRIGHT_DATA_PASSWORD'
    }
  });

  const page = await browser.newPage();
  await page.goto('https://example.com', { waitUntil: 'networkidle' });

  console.log(await page.title());

  await browser.close();
})();

4) Add crawl logic

A crawler should usually include:

  • URL queue
  • visited set / dedupe
  • rate limiting
  • retry with backoff
  • HTML extraction
  • error handling

Example simplified flow:

const urls = [
  'https://example.com',
  'https://example.com/about'
];

for (const url of urls) {
  try {
    await page.goto(url, { waitUntil: 'networkidle', timeout: 60000 });
    const content = await page.textContent('body');
    console.log(url, content.slice(0, 200));
  } catch (err) {
    console.error('Failed:', url, err.message);
  }
}

5) If using Bright Data Web Unlocker

If the site is difficult to scrape, Web Unlocker may be easier than managing your own browser automation.

Typical pattern:

  • Send requests through Bright Data’s unlocker endpoint
  • Let Bright Data handle rotation, CAPTCHA, and blocking

You’ll usually configure:

  • target URL
  • zone
  • headers / geo / render options depending on your plan

6) Best practices

  • Respect site terms and robots.txt where applicable
  • Use delays between requests
  • Rotate session identifiers when needed
  • Keep concurrency low at first
  • Log status codes and failures
  • Store crawled data incrementally

7) Example architecture for a real crawler

A solid production setup:

  • Queue: Redis / RabbitMQ / simple DB table
  • Workers: Playwright or HTTP client
  • Proxy layer: Bright Data zone
  • Storage: PostgreSQL / S3 / Elasticsearch
  • Monitoring: logs + retry counts + success rate

8) Troubleshooting tips

If requests fail:

  • Verify proxy credentials
  • Check zone is active
  • Make sure the host/port are correct
  • Try different session persistence settings
  • Use browser automation instead of raw HTTP for JS-heavy sites
  • Inspect response codes and proxy errors

If you want, I can give you one of these next:

  1. A complete Bright Data crawler in Node.js
  2. A Python crawler with Playwright
  3. A Bright Data Web Unlocker example
  4. A production-ready crawler architecture

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.