Prompt
How do I extract real estate listings and addresses from websites?
Latest observation
Extracting real estate listings and addresses from websites is usually done with a mix of web scraping, HTML parsing, and sometimes browser automation if the site is dynamic.
1) First, check the rules
Before scraping, make sure to review:
- the site’s Terms of Service
- robots.txt
- any API they provide
Some sites prohibit automated extraction, especially for real estate data.
2) Common approaches
A. If the data is in the HTML
Use a scraper to request pages and parse the content.
Typical tools:
- Python:
requests,BeautifulSoup,lxml - JavaScript/Node:
axios,cheerio
Example in Python:
import requests
from bs4 import BeautifulSoup
url = "https://example.com/listings"
html = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).text
soup = BeautifulSoup(html, "html.parser")
for card in soup.select(".listing-card"):
title = card.select_one(".title").get_text(strip=True)
address = card.select_one(".address").get_text(strip=True)
print(title, address)
B. If the site loads data with JavaScript
Many real estate sites render listings dynamically. In that case:
- inspect network requests in browser dev tools
- look for underlying JSON/API endpoints
- use browser automation if needed
Tools:
- Python:
playwright,selenium - Node:
puppeteer,playwright
C. If there’s a JSON API behind the page
This is often the cleanest method.
In browser dev tools:
- Open the page
- Go to Network
- Filter by Fetch/XHR
- Reload page
- Look for responses containing listing data
Then call that endpoint directly with code.
3) Extracting addresses specifically
Addresses may appear:
- in visible text
- in embedded JSON-LD
- in data attributes
- in API responses
Example: JSON-LD extraction
import requests, json
from bs4 import BeautifulSoup
html = requests.get("https://example.com/listing/123").text
soup = BeautifulSoup(html, "html.parser")
for script in soup.find_all("script", type="application/ld+json"):
data = json.loads(script.string)
if isinstance(data, dict) and data.get("@type") in ["House", "Apartment", "Residence"]:
print(data.get("address"))
Example: finding address in page text
address = soup.select_one(".listing-address").get_text(" ", strip=True)
4) Handling pagination
Real estate sites usually have multiple pages or “load more” buttons. You’ll need to:
- loop through page numbers
- follow next-page links
- or use the API parameters
Example:
for page in range(1, 6):
url = f"https://example.com/listings?page={page}"
# fetch and parse
5) If the page uses anti-bot protections
Some sites use:
- rate limiting
- CAPTCHAs
- Cloudflare
- required sessions/cookies
You should avoid bypassing protections if it violates site rules. Safer options:
- use an official API
- request permission
- scrape at low volume
- cache results
6) Best practice workflow
- Inspect the page source
- Check for JSON/API endpoints
- Use selectors for listing cards and addresses
- Handle pagination
- Store results in CSV/JSON/database
- Add throttling and error handling
7) Example output format
You might save fields like:
- listing_id
- address
- city
- state
- zip
- price
- bedrooms
- bathrooms
- url
8) If you want, I can help you with:
- a Python scraper for a specific site
- extracting data from HTML vs JavaScript-rendered pages
- building a CSV export
- parsing addresses and geocoding
- writing a scraper that respects rate limits
If you share a sample URL or page HTML, I can show you the exact selectors or code structure.