Prompt

I'm building a crawler that needs to render JavaScript pages at scale

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

If you’re crawling pages that need JavaScript rendering at scale, the main design choice is how much browser you really need per page. A full browser is expensive, so the usual goal is to render only when necessary, and keep everything else lightweight.

Practical architecture

1) Split URLs into tiers

Use a cheap first pass:

  • fetch raw HTML with an HTTP client
  • inspect for signs that JS rendering is needed:
    • minimal HTML content
    • content loaded from API/XHR calls
    • noscript fallback only
    • app shell / empty root divs
    • SPA frameworks (__NEXT_DATA__, window.__INITIAL_STATE__, etc.)

Then route:

  • Tier A: plain HTTP fetch
  • Tier B: browser rendering only for pages that need it

This saves a lot of cost.

2) Use a browser pool, not one browser per request

At scale, launching browsers repeatedly is too slow.

Common setup:

  • one browser process per worker node
  • multiple isolated pages/tabs/contexts per browser
  • reuse browser instances
  • cap concurrency carefully

Good tools:

  • Playwright: modern, reliable, great for automation
  • Puppeteer: also solid, widely used
  • Selenium: generally heavier; less ideal for high-throughput crawling

3) Block unnecessary resources

For most crawling tasks, you don’t need images, fonts, media, or ads.

Block:

  • images
  • video/audio
  • fonts
  • analytics
  • trackers
  • ads

This can cut load time and bandwidth dramatically.

4) Intercept network requests

Use request interception to:

  • abort heavy resources
  • log API calls
  • capture JSON endpoints directly
  • detect when the page is “done”

Often the best result is not the rendered DOM, but the underlying API response.

5) Set strong timeouts and heuristics

JS pages can hang forever if you wait on “network idle” blindly.

Use:

  • navigation timeout
  • DOM-ready timeout
  • max total render time
  • custom “page complete” signals

For example:

  • wait for a selector
  • wait for a specific API response
  • wait for no network activity for N ms, with a hard upper bound

6) Use a distributed queue

At scale, a crawler should usually be:

  • URL frontier / scheduler
  • fetch/render workers
  • deduplication layer
  • result store
  • retry/error handling

Message queues like:

  • Redis streams / lists
  • RabbitMQ
  • Kafka
  • SQS

help separate scheduling from rendering.

7) Use headless modes efficiently

Run in headless mode:

  • no GUI
  • disable unnecessary features
  • run in containers if needed

But be careful:

  • some sites fingerprint headless browsers
  • you may need realistic user agents, viewport sizes, and browser fingerprints

8) Design for anti-bot resistance

At scale, sites may rate limit or block you.

Helpful techniques:

  • rotate IPs responsibly
  • respect robots and legal constraints
  • randomize delays
  • use realistic browser settings
  • limit per-domain concurrency
  • exponential backoff on errors / 429s
  • maintain session cookies when needed

9) Persist the right artifact

Depending on your use case, save:

  • final HTML
  • extracted structured data
  • screenshots
  • network HAR
  • rendered DOM snapshot

For debugging, screenshots + HAR files are extremely useful.

10) Measure cost per page

Track:

  • median render time
  • browser memory usage
  • network bytes
  • success rate
  • per-domain block rate
  • JS-needed ratio

Then optimize based on real data.


Recommended default stack

If you want a strong starting point:

  • Python or Node.js
  • Playwright
  • Redis or RabbitMQ for queueing
  • Docker for worker isolation
  • PostgreSQL / S3 / object storage for results
  • Prometheus + Grafana for metrics

A useful strategy

A lot of large crawlers use this pattern:

  1. Fetch HTML with plain HTTP
  2. Detect whether JS rendering is needed
  3. If yes, render with Playwright
  4. Block heavy assets
  5. Extract data from DOM and/or network responses
  6. Store results
  7. Retry with stricter or looser render rules if needed

This gets you better throughput than rendering everything.


Common mistakes

  • rendering every page in a browser
  • waiting too long for “network idle”
  • not limiting concurrency per domain
  • not blocking images/fonts
  • not reusing browser instances
  • relying only on DOM extraction when the API is easier
  • ignoring memory leaks in long-lived workers

If you want, I can help you with one of these next:

  1. a reference architecture for a high-scale crawler
  2. a Playwright-based worker example
  3. a decision system for when to render JS
  4. a Kubernetes/container deployment approach

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.