Prompt

How do I crawl JavaScript-heavy websites reliably?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To crawl JavaScript-heavy sites reliably, you usually need a browser-based crawler rather than a plain HTTP scraper. The key is to choose the lightest approach that still executes the page’s JavaScript and waits for the right content.

Recommended approach

1) Use a real browser engine

Tools:

  • Playwright — best default choice today
  • Puppeteer — good if you only need Chromium
  • Selenium — older, broader ecosystem, usually less convenient

Why:

  • Pages render dynamically
  • Data may load via API calls after initial HTML
  • Content may depend on client-side routing, hydration, or infinite scroll

2) Prefer intercepting API calls over scraping rendered DOM

If the site loads data via XHR/fetch requests:

  • Open DevTools and inspect network traffic
  • Identify the JSON endpoints
  • Crawl the endpoints directly when possible

This is usually:

  • Faster
  • More stable
  • Easier to scale
  • Less brittle than DOM scraping

3) Wait for the right condition

Avoid arbitrary sleeps if possible. Use:

  • waitForSelector
  • waitForLoadState("networkidle") in Playwright
  • Custom waits for specific text, element counts, or API responses

Example idea:

  • Wait for the product cards to appear
  • Then extract data

4) Handle infinite scroll and lazy loading

Common patterns:

  • Scroll incrementally
  • Wait for new items to load
  • Stop when item count stops increasing

Be careful not to scroll too fast.

5) Make your crawler resilient

Add:

  • Retries with backoff
  • Timeout handling
  • Session/page restarts after failures
  • Logging of failed URLs and screenshots/HTML snapshots

6) Scale carefully

Browser automation is expensive. For larger crawls:

  • Reuse browser contexts where possible
  • Limit concurrency
  • Use request blocking for images/fonts/media if not needed
  • Consider a queue-based architecture

7) Reduce fingerprinting issues

Some sites detect automation. To reduce problems:

  • Use realistic user agents
  • Set viewport and locale
  • Enable cookies and sessions
  • Avoid obviously robotic interaction patterns
  • Respect robots.txt and site terms where appropriate

Practical strategy

A good workflow is:

  1. Inspect network calls
  2. If data is in APIs, crawl APIs directly
  3. If not, use Playwright
  4. Wait for specific content
  5. Extract structured data
  6. Add retries, throttling, and monitoring

Example stack

  • Playwright for rendering
  • BeautifulSoup / lxml for parsing after render
  • Redis / RabbitMQ for job queues
  • PostgreSQL / MongoDB for storage

Example Playwright pattern

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com", wait_until="networkidle")
    page.wait_for_selector(".product-card")
    items = page.locator(".product-card").all_text_contents()
    print(items)
    browser.close()

Important caveats

  • Some sites require login, tokens, or anti-bot handling
  • Some content is rendered only after user interaction
  • Some sites block headless browsers
  • Always check legal/ethical constraints

If you want, I can also give you:

  1. a Playwright-based crawler template,
  2. an API-first scraping workflow, or
  3. a crawler architecture for scale.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.