Prompt

How do I scrape real estate listings and normalize the fields?

Technology / API Platforms1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

Scraping real estate listings and normalizing the fields is usually a 2-part job:

  1. Collect the raw listing data
  2. Map/clean it into a consistent schema

A good approach depends on whether you’re scraping from HTML pages, JSON APIs, or both.


1) Define a canonical schema first

Before scraping, decide what your normalized record should look like. For example:

{
  "source": "zillow",
  "listing_id": "12345",
  "url": "https://...",
  "title": "2 Bed Condo in Downtown",
  "property_type": "condo",
  "status": "for_sale",
  "price": 525000,
  "currency": "USD",
  "beds": 2,
  "baths": 2,
  "sqft": 1100,
  "lot_sqft": null,
  "year_built": 2008,
  "address": {
    "line1": "123 Main St",
    "city": "Austin",
    "state": "TX",
    "postal_code": "78701",
    "country": "US"
  },
  "latitude": 30.2672,
  "longitude": -97.7431,
  "description": "..."
}

Typical normalized fields:

  • source, source_listing_id, url
  • title, description
  • status (for_sale, for_rent, sold, pending)
  • property_type (house, condo, townhouse, land, multi_family)
  • price, currency
  • beds, baths, sqft, lot_sqft
  • address components
  • lat, lon
  • year_built
  • hoa_fee, taxes, parking, etc. if available

2) Scrape raw data

Prefer APIs/JSON when available

Many listing sites render data in embedded JSON or XHR calls. That’s often easier and more reliable than parsing HTML.

Common places to look:

  • <script type="application/ld+json">
  • embedded page state like __NEXT_DATA__, window.__INITIAL_STATE__
  • network requests returning JSON
  • structured data in HTML attributes

Use HTML scraping only when needed

If you scrape HTML:

  • use requests + BeautifulSoup for static pages
  • use Playwright or Selenium for JS-rendered sites

Basic example:

import requests
from bs4 import BeautifulSoup

url = "https://example.com/listing/123"
html = requests.get(url, headers={"User-Agent": "Mozilla/5.0"}).text
soup = BeautifulSoup(html, "html.parser")

title = soup.select_one("h1").get_text(strip=True)
price = soup.select_one(".price").get_text(strip=True)

3) Normalize field names and values

This is the key part: different sites use different labels and formats.

Example mapping

Raw source fields might be:

  • beds, bedrooms, bd
  • baths, bathrooms, ba
  • sqft, livingArea, homeSize
  • price, list_price, $
  • propertyType, type

Normalize them by creating a mapping table:

FIELD_MAP = {
    "bedrooms": "beds",
    "beds": "beds",
    "bd": "beds",
    "bathrooms": "baths",
    "ba": "baths",
    "livingArea": "sqft",
    "homeSize": "sqft",
    "list_price": "price"
}

4) Clean and convert values

Numeric parsing

Real estate sites often format numbers as strings:

  • "$525,000" → 525000
  • "1,100 sqft" → 1100
  • "2.5 baths" → 2.5

Example helper:

import re

def parse_number(text):
    if text is None:
        return None
    m = re.search(r"[\d,.]+", str(text))
    if not m:
        return None
    return float(m.group(0).replace(",", ""))

Then cast to int where appropriate:

price = int(parse_number("$525,000")) if parse_number("$525,000") else None
sqft = int(parse_number("1,100 sqft")) if parse_number("1,100 sqft") else None
beds = parse_number("2 beds")
baths = parse_number("2.5 baths")

Normalize property type

Sites may say:

  • Single Family
  • SFH
  • House
  • Detached

Map them to canonical values:

PROPERTY_TYPE_MAP = {
    "single family": "house",
    "sfh": "house",
    "house": "house",
    "condo": "condo",
    "townhouse": "townhouse",
    "townhome": "townhouse",
    "multi family": "multi_family",
    "land": "land"
}

Normalize status

Map to standard statuses:

  • active, for sale, available → for_sale
  • pending, under contract → pending
  • sold, closed → sold
  • for rent, rental → for_rent

5) Standardize addresses

Address normalization often needs separate parsing:

  • street line
  • unit/suite
  • city
  • state
  • postal code
  • country

Example tools:

  • usaddress for US addresses
  • libpostal for international addresses

You can also keep the original formatted address and store parsed components separately.


6) Handle missing or inconsistent data

Expect:

  • missing bedrooms/baths
  • multiple price values (list price, estimated value, reduced price)
  • area in different units (sqft, m²)
  • mixed types ("2" vs 2 vs 2.0)
  • ranges ("2-3 beds")

Rules to define:

  • Use null for unknowns
  • Prefer listing price over estimate if both exist
  • Convert units consistently
  • Preserve original raw values for traceability

A useful pattern is to store both:

  • raw_* fields
  • normalized fields

Example:

{
  "raw_price": "$525,000",
  "price": 525000
}

7) Build a transformation pipeline

A practical pipeline looks like this:

  1. Fetch page/API
  2. Extract raw listing object
  3. Map raw keys to canonical keys
  4. Parse/clean values
  5. Validate schema
  6. Save to database/CSV

Example in Python:

def normalize_listing(raw):
    return {
        "source": raw.get("source"),
        "listing_id": raw.get("id") or raw.get("listingId"),
        "url": raw.get("url"),
        "title": raw.get("title"),
        "property_type": normalize_property_type(raw.get("propertyType")),
        "status": normalize_status(raw.get("status")),
        "price": parse_money(raw.get("price")),
        "beds": parse_number(raw.get("beds")),
        "baths": parse_number(raw.get("baths")),
        "sqft": parse_number(raw.get("sqft")),
        "year_built": parse_number(raw.get("yearBuilt")),
        "address": normalize_address(raw.get("address")),
        "latitude": raw.get("lat"),
        "longitude": raw.get("lon"),
        "description": raw.get("description"),
    }

8) Validate output

Use a schema validator to catch bad records:

  • pydantic
  • marshmallow
  • JSON Schema

Example with Pydantic:

from pydantic import BaseModel
from typing import Optional

class Address(BaseModel):
    line1: Optional[str] = None
    city: Optional[str] = None
    state: Optional[str] = None
    postal_code: Optional[str] = None
    country: Optional[str] = None

9) Respect legal/technical constraints

Important:

  • Check the site’s Terms of Service
  • Respect robots.txt
  • Rate-limit requests
  • Use caching and retries
  • Avoid scraping personal data you don’t need
  • Prefer official APIs or data providers when possible

Real estate sites may block scraping aggressively, so robust crawling often requires:

  • rotating user agents carefully
  • backoff/retry
  • session management
  • headless browser for JS-heavy pages

10) Recommended stack

For a small-to-medium scraper:

  • Requests + BeautifulSoup for static pages
  • Playwright for dynamic pages
  • Pydantic for validation
  • PostgreSQL or SQLite for storage
  • pandas if you need CSV/analysis
  • libpostal/usaddress for addresses

A simple practical pattern

If you’re scraping multiple sites, keep three layers:

1. Site-specific extractors

Return raw dictionaries from each site.

2. Common normalization layer

Convert raw dicts into your canonical schema.

3. Storage/export layer

Write normalized records to DB/CSV/Parquet.

That keeps site quirks isolated and makes maintenance much easier.


If you want, I can give you:

  1. a Python scraper + normalizer template,
  2. a Pydantic schema for real estate listings, or
  3. a site-by-site field mapping strategy.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.