Prompt
How do I scrape real estate listings and normalize the fields?
Latest observation
Scraping real estate listings and normalizing the fields is usually a 2-part job:
- Collect the raw listing data
- Map/clean it into a consistent schema
A good approach depends on whether you’re scraping from HTML pages, JSON APIs, or both.
1) Define a canonical schema first
Before scraping, decide what your normalized record should look like. For example:
{
"source": "zillow",
"listing_id": "12345",
"url": "https://...",
"title": "2 Bed Condo in Downtown",
"property_type": "condo",
"status": "for_sale",
"price": 525000,
"currency": "USD",
"beds": 2,
"baths": 2,
"sqft": 1100,
"lot_sqft": null,
"year_built": 2008,
"address": {
"line1": "123 Main St",
"city": "Austin",
"state": "TX",
"postal_code": "78701",
"country": "US"
},
"latitude": 30.2672,
"longitude": -97.7431,
"description": "..."
}
Typical normalized fields:
source,source_listing_id,urltitle,descriptionstatus(for_sale, for_rent, sold, pending)property_type(house, condo, townhouse, land, multi_family)price,currencybeds,baths,sqft,lot_sqftaddresscomponentslat,lonyear_builthoa_fee,taxes,parking, etc. if available
2) Scrape raw data
Prefer APIs/JSON when available
Many listing sites render data in embedded JSON or XHR calls. That’s often easier and more reliable than parsing HTML.
Common places to look:
<script type="application/ld+json">- embedded page state like
__NEXT_DATA__,window.__INITIAL_STATE__ - network requests returning JSON
- structured data in HTML attributes
Use HTML scraping only when needed
If you scrape HTML:
- use
requests+BeautifulSoupfor static pages - use
PlaywrightorSeleniumfor JS-rendered sites
Basic example:
import requests
from bs4 import BeautifulSoup
url = "https://example.com/listing/123"
html = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).text
soup = BeautifulSoup(html, "html.parser")
title = soup.select_one("h1").get_text(strip=True)
price = soup.select_one(".price").get_text(strip=True)
3) Normalize field names and values
This is the key part: different sites use different labels and formats.
Example mapping
Raw source fields might be:
beds,bedrooms,bdbaths,bathrooms,basqft,livingArea,homeSizeprice,list_price,$propertyType,type
Normalize them by creating a mapping table:
FIELD_MAP = {
"bedrooms": "beds",
"beds": "beds",
"bd": "beds",
"bathrooms": "baths",
"ba": "baths",
"livingArea": "sqft",
"homeSize": "sqft",
"list_price": "price"
}
4) Clean and convert values
Numeric parsing
Real estate sites often format numbers as strings:
"$525,000"→525000"1,100 sqft"→1100"2.5 baths"→2.5
Example helper:
import re
def parse_number(text):
if text is None:
return None
m = re.search(r"[\d,.]+", str(text))
if not m:
return None
return float(m.group(0).replace(",", ""))
Then cast to int where appropriate:
price = int(parse_number("$525,000")) if parse_number("$525,000") else None
sqft = int(parse_number("1,100 sqft")) if parse_number("1,100 sqft") else None
beds = parse_number("2 beds")
baths = parse_number("2.5 baths")
Normalize property type
Sites may say:
Single FamilySFHHouseDetached
Map them to canonical values:
PROPERTY_TYPE_MAP = {
"single family": "house",
"sfh": "house",
"house": "house",
"condo": "condo",
"townhouse": "townhouse",
"townhome": "townhouse",
"multi family": "multi_family",
"land": "land"
}
Normalize status
Map to standard statuses:
active,for sale,available→for_salepending,under contract→pendingsold,closed→soldfor rent,rental→for_rent
5) Standardize addresses
Address normalization often needs separate parsing:
- street line
- unit/suite
- city
- state
- postal code
- country
Example tools:
usaddressfor US addresseslibpostalfor international addresses
You can also keep the original formatted address and store parsed components separately.
6) Handle missing or inconsistent data
Expect:
- missing bedrooms/baths
- multiple price values (list price, estimated value, reduced price)
- area in different units (
sqft,m²) - mixed types (
"2"vs2vs2.0) - ranges (
"2-3 beds")
Rules to define:
- Use
nullfor unknowns - Prefer listing price over estimate if both exist
- Convert units consistently
- Preserve original raw values for traceability
A useful pattern is to store both:
raw_*fields- normalized fields
Example:
{
"raw_price": "$525,000",
"price": 525000
}
7) Build a transformation pipeline
A practical pipeline looks like this:
- Fetch page/API
- Extract raw listing object
- Map raw keys to canonical keys
- Parse/clean values
- Validate schema
- Save to database/CSV
Example in Python:
def normalize_listing(raw):
return {
"source": raw.get("source"),
"listing_id": raw.get("id") or raw.get("listingId"),
"url": raw.get("url"),
"title": raw.get("title"),
"property_type": normalize_property_type(raw.get("propertyType")),
"status": normalize_status(raw.get("status")),
"price": parse_money(raw.get("price")),
"beds": parse_number(raw.get("beds")),
"baths": parse_number(raw.get("baths")),
"sqft": parse_number(raw.get("sqft")),
"year_built": parse_number(raw.get("yearBuilt")),
"address": normalize_address(raw.get("address")),
"latitude": raw.get("lat"),
"longitude": raw.get("lon"),
"description": raw.get("description"),
}
8) Validate output
Use a schema validator to catch bad records:
pydanticmarshmallow- JSON Schema
Example with Pydantic:
from pydantic import BaseModel
from typing import Optional
class Address(BaseModel):
line1: Optional[str] = None
city: Optional[str] = None
state: Optional[str] = None
postal_code: Optional[str] = None
country: Optional[str] = None
9) Respect legal/technical constraints
Important:
- Check the site’s Terms of Service
- Respect
robots.txt - Rate-limit requests
- Use caching and retries
- Avoid scraping personal data you don’t need
- Prefer official APIs or data providers when possible
Real estate sites may block scraping aggressively, so robust crawling often requires:
- rotating user agents carefully
- backoff/retry
- session management
- headless browser for JS-heavy pages
10) Recommended stack
For a small-to-medium scraper:
- Requests + BeautifulSoup for static pages
- Playwright for dynamic pages
- Pydantic for validation
- PostgreSQL or SQLite for storage
- pandas if you need CSV/analysis
- libpostal/usaddress for addresses
A simple practical pattern
If you’re scraping multiple sites, keep three layers:
1. Site-specific extractors
Return raw dictionaries from each site.
2. Common normalization layer
Convert raw dicts into your canonical schema.
3. Storage/export layer
Write normalized records to DB/CSV/Parquet.
That keeps site quirks isolated and makes maintenance much easier.
If you want, I can give you:
- a Python scraper + normalizer template,
- a Pydantic schema for real estate listings, or
- a site-by-site field mapping strategy.