Prompt
I'm building a scraper for real estate listings with pagination and dynamic content
Latest observation
A good way to approach a real estate scraper with pagination and dynamic content is to split it into 4 parts:
- Discover the listing URLs
- Handle pagination
- Render dynamic content
- Extract and store data
1) Choose the right tool
If the site loads content with JavaScript, use a browser automation tool instead of plain requests:
- Playwright: best modern option
- Selenium: widely used, a bit heavier
- requests + BeautifulSoup: only if the page HTML already contains the listings
For dynamic real estate sites, I’d usually recommend Playwright.
2) Pagination strategies
Common pagination patterns:
A. Page number links
Example:
/listings?page=1/listings?page=2
You can loop until no results are found.
B. “Load more” button
Click the button repeatedly until it disappears or no new listings appear.
C. Infinite scroll
Scroll down, wait for more content, repeat until the count stops increasing.
3) Extract listing data
Typical fields:
- title
- price
- address/location
- beds/baths/sqft
- listing URL
- image URL
- agent/agency
- description
Try to extract from the listing cards first, then visit the detail page only if needed.
4) Example with Playwright
Here’s a simple Python example for a paginated site:
from playwright.sync_api import sync_playwright
def scrape_listings(base_url):
listings = []
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
current_page = 1
while True:
url = f"{base_url}?page={current_page}"
page.goto(url, wait_until="networkidle")
cards = page.locator(".listing-card")
count = cards.count()
if count == 0:
break
for i in range(count):
card = cards.nth(i)
title = card.locator(".title").inner_text().strip()
price = card.locator(".price").inner_text().strip()
link = card.locator("a").get_attribute("href")
listings.append({
"title": title,
"price": price,
"link": link
})
current_page += 1
browser.close()
return listings
5) Example for infinite scroll
from playwright.sync_api import sync_playwright
import time
def scrape_infinite_scroll(url):
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto(url, wait_until="networkidle")
last_height = 0
while True:
page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
time.sleep(2)
new_height = page.evaluate("document.body.scrollHeight")
if new_height == last_height:
break
last_height = new_height
# then extract listings
cards = page.locator(".listing-card")
results = []
for i in range(cards.count()):
card = cards.nth(i)
results.append({
"title": card.locator(".title").inner_text().strip(),
"price": card.locator(".price").inner_text().strip()
})
browser.close()
return results
6) Tips for dynamic sites
- Use
wait_for_selector()to wait for listings to appear - Prefer explicit waits over fixed sleeps
- Inspect network requests in DevTools — sometimes data comes from a JSON API you can call directly
- Add rate limiting and retries
- Use rotating headers or proxies only if appropriate and allowed
7) Important caution
Before scraping, check:
- the site’s Terms of Service
robots.txt- whether the site provides an API or RSS feed
Also avoid aggressive scraping that could overload their servers.
If you want, I can help you with one of these next:
- a Playwright scraper template
- a pagination scraper for a specific site
- a load-more / infinite-scroll scraper
- a parser for listing detail pages