Prompt

I'm building a scraper for real estate listings with pagination and dynamic content

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

A good way to approach a real estate scraper with pagination and dynamic content is to split it into 4 parts:

  1. Discover the listing URLs
  2. Handle pagination
  3. Render dynamic content
  4. Extract and store data

1) Choose the right tool

If the site loads content with JavaScript, use a browser automation tool instead of plain requests:

  • Playwright: best modern option
  • Selenium: widely used, a bit heavier
  • requests + BeautifulSoup: only if the page HTML already contains the listings

For dynamic real estate sites, I’d usually recommend Playwright.


2) Pagination strategies

Common pagination patterns:

A. Page number links

Example:

  • /listings?page=1
  • /listings?page=2

You can loop until no results are found.

B. “Load more” button

Click the button repeatedly until it disappears or no new listings appear.

C. Infinite scroll

Scroll down, wait for more content, repeat until the count stops increasing.


3) Extract listing data

Typical fields:

  • title
  • price
  • address/location
  • beds/baths/sqft
  • listing URL
  • image URL
  • agent/agency
  • description

Try to extract from the listing cards first, then visit the detail page only if needed.


4) Example with Playwright

Here’s a simple Python example for a paginated site:

from playwright.sync_api import sync_playwright

def scrape_listings(base_url):
    listings = []

    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()

        current_page = 1
        while True:
            url = f"{base_url}?page={current_page}"
            page.goto(url, wait_until="networkidle")

            cards = page.locator(".listing-card")
            count = cards.count()

            if count == 0:
                break

            for i in range(count):
                card = cards.nth(i)
                title = card.locator(".title").inner_text().strip()
                price = card.locator(".price").inner_text().strip()
                link = card.locator("a").get_attribute("href")

                listings.append({
                    "title": title,
                    "price": price,
                    "link": link
                })

            current_page += 1

        browser.close()

    return listings

5) Example for infinite scroll

from playwright.sync_api import sync_playwright
import time

def scrape_infinite_scroll(url):
    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto(url, wait_until="networkidle")

        last_height = 0
        while True:
            page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
            time.sleep(2)

            new_height = page.evaluate("document.body.scrollHeight")
            if new_height == last_height:
                break
            last_height = new_height

        # then extract listings
        cards = page.locator(".listing-card")
        results = []
        for i in range(cards.count()):
            card = cards.nth(i)
            results.append({
                "title": card.locator(".title").inner_text().strip(),
                "price": card.locator(".price").inner_text().strip()
            })

        browser.close()
        return results

6) Tips for dynamic sites

  • Use wait_for_selector() to wait for listings to appear
  • Prefer explicit waits over fixed sleeps
  • Inspect network requests in DevTools — sometimes data comes from a JSON API you can call directly
  • Add rate limiting and retries
  • Use rotating headers or proxies only if appropriate and allowed

7) Important caution

Before scraping, check:

  • the site’s Terms of Service
  • robots.txt
  • whether the site provides an API or RSS feed

Also avoid aggressive scraping that could overload their servers.


If you want, I can help you with one of these next:

  1. a Playwright scraper template
  2. a pagination scraper for a specific site
  3. a load-more / infinite-scroll scraper
  4. a parser for listing detail pages

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as Circuit. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.