Prompt

How do I pull structured data from pages with infinite scroll?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Pulling structured data from infinite-scroll pages usually means you need to do two things:

  1. Figure out how the page loads more content
  2. Extract the data once it’s loaded

There are a few common approaches.


1) Use the underlying API if possible

This is usually the best option.

Many infinite-scroll pages don’t load HTML directly. They fetch JSON from an API as you scroll.

How to find it

  • Open DevTools → Network
  • Scroll the page
  • Look for requests like:
    • fetch
    • xhr
    • graphql
  • Inspect the response:
    • If it’s JSON, you can often scrape the API directly.

Why this is best

  • Faster
  • More reliable than browser automation
  • Easier to parse structured data
  • Less likely to break if the UI changes

Example pattern

You may see something like:

  • GET /api/items?page=1
  • GET /api/items?cursor=abc123
  • POST /graphql

Then you can replicate those requests with Python requests, httpx, or similar.


2) Use browser automation to scroll and collect rendered content

If there’s no usable API, use a browser automation tool like:

  • Playwright
  • Selenium
  • Puppeteer

General workflow

  • Open the page
  • Scroll to the bottom repeatedly
  • Wait for new items to load
  • Extract item elements from the DOM

Things to watch for

  • Lazy loading delays
  • “Load more” buttons
  • Virtualized lists (old items disappear from DOM)
  • Anti-bot measures

Example with Playwright (Python)

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com")

    last_height = 0
    while True:
        page.mouse.wheel(0, 2000)
        page.wait_for_timeout(1500)

        height = page.evaluate("document.body.scrollHeight")
        if height == last_height:
            break
        last_height = height

    items = page.locator(".item")
    data = []
    for i in range(items.count()):
        data.append(items.nth(i).inner_text())

    print(data)
    browser.close()

3) If the page uses “Load more” instead of infinite scroll

Sometimes scrolling triggers the same behavior as clicking a button.

In that case:

  • Click the button repeatedly
  • Wait for new content
  • Continue until the button disappears or no more content loads

4) Handle cursor-based pagination

A lot of infinite-scroll APIs use a cursor instead of page numbers.

Example:

  • First request returns:
    {
      "items": [...],
      "next_cursor": "xyz"
    }
    
  • Next request uses that cursor:
    /api/items?cursor=xyz
    

You keep requesting until next_cursor is null or missing.

This is often the cleanest structured-data approach.


5) Deal with virtualized lists

Some sites only render visible rows for performance.

That means:

  • The DOM may only contain 20–50 visible items
  • Scrolling causes old rows to be removed

If that happens:

  • Don’t rely only on final DOM state
  • Extract data incrementally as you scroll
  • Or use the API directly if possible

6) Wait for network activity, not just time

Instead of fixed sleep delays, it’s better to wait for:

  • a specific selector to appear
  • network requests to finish
  • item count to increase

This makes scraping more reliable.


7) Normalize the data afterward

Once extracted, convert the data into a structured format like:

  • JSON
  • CSV
  • a database table

For example:

  • product name
  • price
  • rating
  • URL
  • description

If you’re scraping text from DOM elements, it helps to build a schema first.


Recommended strategy

In order of preference:

  1. Find and call the API directly
  2. Use Playwright/Selenium if the API is hidden or protected
  3. Scroll and extract incrementally
  4. Store results in a structured schema

Practical checklist

  • Open DevTools Network tab
  • Scroll and inspect requests
  • Identify JSON API or GraphQL
  • Recreate request in code
  • Paginate with cursor/page params
  • Stop when no new items arrive
  • Save results as structured output

If you want, I can show you:

  • a Python Playwright scraper for infinite scroll
  • a requests-based API scraper
  • or how to inspect network calls to find the hidden endpoint

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.