Prompt

How do I build a scraper that handles JavaScript-rendered pages?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

To scrape JavaScript-rendered pages, you usually need a browser automation tool instead of a basic HTTP request library. The browser loads the page, runs JavaScript, and lets you extract the final rendered DOM.

Common approaches

1) Use a headless browser

Good options:

  • Playwright (recommended)
  • Selenium
  • Puppeteer (Node.js)

These tools can:

  • load pages like a real browser
  • wait for JS content to appear
  • click buttons / scroll / fill forms
  • extract rendered HTML or structured data

Recommended stack: Playwright

Install

Python

pip install playwright
playwright install

Basic example

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com", wait_until="networkidle")
    
    # Extract text after JS renders
    content = page.text_content("body")
    print(content)
    
    browser.close()

Wait for specific elements

If content loads later, wait for a selector:

page.goto("https://example.com")
page.wait_for_selector(".product-card")
items = page.locator(".product-card").all_text_contents()
print(items)

Best practices

1) Prefer APIs if available

Many JS-rendered sites fetch data from backend APIs. If you can find the API request, scraping it is:

  • faster
  • more reliable
  • less resource-intensive

Use browser devtools → Network tab to inspect requests.


2) Wait intelligently

Avoid fixed sleeps unless necessary. Use:

  • wait_for_selector
  • wait_for_load_state("networkidle")
  • explicit waits for known content

3) Handle pagination and scrolling

Some sites lazy-load content on scroll. Example:

for _ in range(5):
    page.mouse.wheel(0, 2000)
    page.wait_for_timeout(1000)

4) Rotate headers / user agents carefully

Sometimes sites block default automation signatures. You may need:

  • a realistic user-agent
  • proper language headers
  • cookies/session handling

5) Respect rate limits and robots.txt

Be mindful of:

  • site terms of service
  • legal/ethical restrictions
  • request volume
  • anti-bot measures

When to use which tool

  • Playwright: best overall for modern JS-heavy sites
  • Selenium: broad compatibility, older ecosystem
  • Scrapy + Playwright: great for larger scraping projects
  • Requests/HTTPX: only if the data is available in raw HTML or APIs

A practical workflow

  1. Open the page in browser
  2. Inspect Network requests
  3. Check whether data comes from an API
  4. If yes, scrape the API directly
  5. If not, use Playwright/Selenium to render the page
  6. Wait for the target DOM nodes
  7. Extract the content
  8. Store results and add retries/error handling

Example: scraping rendered data with Playwright

from playwright.sync_api import sync_playwright

url = "https://example.com/products"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto(url, wait_until="domcontentloaded")
    page.wait_for_selector(".product")

    products = []
    for card in page.locator(".product").all():
        products.append({
            "name": card.locator(".name").text_content(),
            "price": card.locator(".price").text_content()
        })

    print(products)
    browser.close()

If you want, I can show you:

  1. a Python Playwright scraper template,
  2. a Node.js Puppeteer version, or
  3. how to extract API calls from a JS-rendered site.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.