Prompt

How do I crawl JavaScript-heavy sites at scale?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Crawling JavaScript-heavy sites at scale is mostly an engineering problem of rendering efficiently, extracting reliably, and controlling cost/latency.

1) Use the right crawler architecture

A good scalable setup is usually:

  • Queue-based frontier: URLs go into a queue; workers pull tasks.
  • Two-stage crawl:
    1. Fast fetch: try plain HTTP first.
    2. Render only when needed: use a browser engine for pages that require JS.
  • Worker pool:
    • HTTP workers for cheap pages
    • Headless browser workers for JS pages
  • Deduplication:
    • Normalize URLs
    • Track content hashes
    • Avoid revisiting identical states

This keeps browser usage limited to pages that truly need it.

2) Detect when JS rendering is necessary

Don’t render everything. Common signals a page needs JS:

  • Empty or sparse HTML in initial response
  • Content loaded via XHR/fetch after load
  • Important data missing from server HTML
  • SPA frameworks like React/Vue/Angular
  • Links or pagination built dynamically

Heuristics:

  • Compare initial HTML to rendered DOM.
  • Search for “app shell” patterns.
  • Check network calls from a browser session and see if data comes from an API.

3) Prefer API extraction when possible

For many JS sites, the browser is just a UI layer over JSON APIs.

Best approach:

  • Open the site once in DevTools/Playwright
  • Inspect network requests
  • Identify backend endpoints returning JSON
  • Crawl those APIs directly

This is usually:

  • Faster
  • More reliable
  • Cheaper
  • Easier to scale

4) Use headless browsers efficiently

For pages that require rendering, use tools like:

  • Playwright (recommended)
  • Puppeteer
  • Selenium (works, but generally less efficient for large-scale crawling)

Scaling tips:

  • Reuse browser instances
  • Open many pages in one browser process if stable
  • Use browser contexts for isolation instead of new processes
  • Block unnecessary resources:
    • images
    • fonts
    • video
    • analytics
    • ads
  • Set timeouts aggressively
  • Wait for the right event, not arbitrary sleeps

Good wait strategy:

  • networkidle can help, but some sites never go idle
  • Prefer waiting for a specific selector or API response

5) Render only until the data you need is available

If you need page title, product info, or article text:

  • Wait for the main content selector
  • Extract
  • Stop the page early if possible

Avoid waiting for:

  • all scripts
  • all lazy-loaded assets
  • animations
  • endless polling

6) Handle infinite scroll and lazy loading

Many JS-heavy sites load content as you scroll.

Strategies:

  • Scroll in increments and stop when:
    • no new content appears
    • API calls stop changing
    • item count stabilizes
  • Track page height and DOM item counts
  • Use pagination APIs if available instead of scrolling

For scale, prefer direct API pagination over UI scrolling.

7) Build robust extraction

JavaScript sites change often. Make your scraper resilient:

  • Use stable selectors, not brittle CSS paths
  • Extract structured data from:
    • JSON-LD
    • embedded state objects
    • hydration payloads
    • API responses
  • Keep parser logic separate from crawling logic
  • Log extraction failures with HTML snapshots or response traces

8) Manage concurrency carefully

Headless browsers are expensive.

Practical tactics:

  • Limit browser concurrency per machine
  • Use autoscaling workers
  • Split workloads:
    • lightweight HTTP fetchers
    • heavyweight browser renderers
  • Use backpressure so browser queues don’t explode

Rule of thumb:

  • HTTP crawling scales much more cheaply than full browser rendering.
  • Use browsers as a last resort.

9) Be resilient to anti-bot measures

JS-heavy sites often also have anti-bot controls.

You may need:

  • Session management
  • Cookies / login handling
  • Rate limiting
  • Human-like pacing
  • Retries with jitter
  • Proxy rotation where appropriate

Also:

  • Respect robots.txt and site terms
  • Don’t overload sites
  • Use crawl budgets per host

10) Store crawl state and observability

At scale, you need visibility into failure modes.

Track:

  • Fetch success rate
  • Render success rate
  • Time to first byte
  • Time to content ready
  • Extraction success rate
  • Per-host error rates
  • Queue depth

Store:

  • URL
  • normalized URL
  • fetch status
  • final rendered URL
  • response metadata
  • content hash
  • extracted entities

This helps you re-crawl efficiently and troubleshoot regressions.

11) A practical pattern that works well

A common scalable workflow:

  1. Enqueue seed URLs
  2. HTTP fetch page
  3. If data is present in HTML:
    • extract and continue
  4. If not:
    • check network/API clues
    • if possible, switch to API crawling
  5. If still necessary:
    • render with Playwright
    • wait for target selector
    • extract data
  6. Save results
  7. Discover more URLs and enqueue them

12) Recommended stack

A solid stack for JS-heavy crawling:

  • Playwright for rendering
  • Redis / Kafka / SQS for queues
  • Postgres / Elasticsearch / object storage for results
  • Docker for deployment
  • Kubernetes / autoscaling workers for scale
  • Proxy layer if needed
  • Prometheus/Grafana for metrics

13) Cost-saving checklist

Before rendering a page, ask:

  • Can I get the same data from an API?
  • Is the content already in initial HTML?
  • Can I extract from JSON-LD or embedded state?
  • Can I block heavy resources?
  • Can I stop rendering once the target element appears?

If yes to any, do that first.

14) Common mistake to avoid

The biggest mistake is treating every URL like it needs a full browser.

That leads to:

  • high cost
  • low throughput
  • unstable crawls
  • frequent timeouts

The winning strategy is: HTTP first, API second, browser last.

If you want, I can also give you:

  1. a Playwright-based crawling architecture,
  2. a sample worker design, or
  3. a code example for crawling JS-heavy pages at scale.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.