Prompt
How do I pull structured data from pages with infinite scroll?
Latest observation
Pulling structured data from infinite-scroll pages usually means you need to do two things:
- Figure out how the page loads more content
- Extract the data once it’s loaded
There are a few common approaches.
1) Use the underlying API if possible
This is usually the best option.
Many infinite-scroll pages don’t load HTML directly. They fetch JSON from an API as you scroll.
How to find it
- Open DevTools → Network
- Scroll the page
- Look for requests like:
fetchxhrgraphql
- Inspect the response:
- If it’s JSON, you can often scrape the API directly.
Why this is best
- Faster
- More reliable than browser automation
- Easier to parse structured data
- Less likely to break if the UI changes
Example pattern
You may see something like:
GET /api/items?page=1GET /api/items?cursor=abc123POST /graphql
Then you can replicate those requests with Python requests, httpx, or similar.
2) Use browser automation to scroll and collect rendered content
If there’s no usable API, use a browser automation tool like:
- Playwright
- Selenium
- Puppeteer
General workflow
- Open the page
- Scroll to the bottom repeatedly
- Wait for new items to load
- Extract item elements from the DOM
Things to watch for
- Lazy loading delays
- “Load more” buttons
- Virtualized lists (old items disappear from DOM)
- Anti-bot measures
Example with Playwright (Python)
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com")
last_height = 0
while True:
page.mouse.wheel(0, 2000)
page.wait_for_timeout(1500)
height = page.evaluate("document.body.scrollHeight")
if height == last_height:
break
last_height = height
items = page.locator(".item")
data = []
for i in range(items.count()):
data.append(items.nth(i).inner_text())
print(data)
browser.close()
3) If the page uses “Load more” instead of infinite scroll
Sometimes scrolling triggers the same behavior as clicking a button.
In that case:
- Click the button repeatedly
- Wait for new content
- Continue until the button disappears or no more content loads
4) Handle cursor-based pagination
A lot of infinite-scroll APIs use a cursor instead of page numbers.
Example:
- First request returns:
{ "items": [...], "next_cursor": "xyz" } - Next request uses that cursor:
/api/items?cursor=xyz
You keep requesting until next_cursor is null or missing.
This is often the cleanest structured-data approach.
5) Deal with virtualized lists
Some sites only render visible rows for performance.
That means:
- The DOM may only contain 20–50 visible items
- Scrolling causes old rows to be removed
If that happens:
- Don’t rely only on final DOM state
- Extract data incrementally as you scroll
- Or use the API directly if possible
6) Wait for network activity, not just time
Instead of fixed sleep delays, it’s better to wait for:
- a specific selector to appear
- network requests to finish
- item count to increase
This makes scraping more reliable.
7) Normalize the data afterward
Once extracted, convert the data into a structured format like:
- JSON
- CSV
- a database table
For example:
- product name
- price
- rating
- URL
- description
If you’re scraping text from DOM elements, it helps to build a schema first.
Recommended strategy
In order of preference:
- Find and call the API directly
- Use Playwright/Selenium if the API is hidden or protected
- Scroll and extract incrementally
- Store results in a structured schema
Practical checklist
- Open DevTools Network tab
- Scroll and inspect requests
- Identify JSON API or GraphQL
- Recreate request in code
- Paginate with cursor/page params
- Stop when no new items arrive
- Save results as structured output
If you want, I can show you:
- a Python Playwright scraper for infinite scroll
- a requests-based API scraper
- or how to inspect network calls to find the hidden endpoint