Prompt
How do I normalize scraped page data into JSON?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
A good way to normalize scraped page data into JSON is to separate the process into three steps:
- Extract raw values from the page
- Clean/standardize those values
- Map them into a consistent JSON structure
1. Define your target JSON schema first
Decide what every record should look like, even if some fields are missing.
Example:
{
"title": "",
"price": null,
"currency": "USD",
"url": "",
"image": "",
"description": "",
"inStock": false,
"source": "example.com",
"scrapedAt": "2026-10-04T12:00:00Z"
}
Having a fixed schema makes normalization much easier.
2. Clean the raw values
Scraped data is usually messy: extra spaces, currency symbols, relative URLs, inconsistent casing, etc.
Typical normalization tasks:
- Trim whitespace
- Convert text to lowercase or title case where appropriate
- Parse numbers from strings
- Convert prices into numeric values
- Resolve relative URLs into absolute URLs
- Normalize dates into ISO 8601
- Turn “yes/no”, “in stock/out of stock” into booleans
- Replace missing or invalid values with
null
Example:
" $1,299.99 "→1299.99" /product/123 "→"https://site.com/product/123""Available"→true
3. Use a consistent transformation function
Here’s a simple Python example:
from urllib.parse import urljoin
from datetime import datetime
import re
def normalize_price(text):
if not text:
return None
cleaned = re.sub(r"[^\d.,]", "", text).replace(",", "")
try:
return float(cleaned)
except ValueError:
return None
def normalize_bool(text):
if not text:
return False
return text.strip().lower() in {"yes", "true", "available", "in stock"}
def normalize_record(raw, base_url, source):
return {
"title": raw.get("title", "").strip() or None,
"price": normalize_price(raw.get("price")),
"currency": raw.get("currency", "USD"),
"url": urljoin(base_url, raw.get("url", "")),
"image": urljoin(base_url, raw.get("image", "")) if raw.get("image") else None,
"description": raw.get("description", "").strip() or None,
"inStock": normalize_bool(raw.get("stock")),
"source": source,
"scrapedAt": datetime.utcnow().isoformat() + "Z"
}
4. Handle missing and inconsistent fields
Scraped pages often vary. Your normalization should:
- Use defaults for optional fields
- Set missing fields to
nullrather than empty strings - Coalesce multiple possible source fields
Example:
title = raw.get("title") or raw.get("h1") or raw.get("name")
5. Deduplicate and validate
Before saving JSON:
- Remove duplicates based on an ID, URL, or normalized title
- Validate types and required fields
- Log or quarantine bad records
If possible, define a schema using:
- Python:
pydantic,jsonschema - JavaScript:
zod,ajv
6. Example normalized output
Raw scraped data:
{
"title": " Blue Sneakers ",
"price": "$89.99",
"url": "/products/blue-sneakers",
"stock": "In Stock"
}
Normalized JSON:
{
"title": "Blue Sneakers",
"price": 89.99,
"currency": "USD",
"url": "https://example.com/products/blue-sneakers",
"image": null,
"description": null,
"inStock": true,
"source": "example.com",
"scrapedAt": "2026-10-04T12:00:00Z"
}
Recommended approach
If you’re scraping at scale, use this structure:
- Parser: extracts raw data from HTML
- Normalizer: converts raw fields into canonical values
- Validator: checks schema and types
- Serializer: writes JSON
This keeps scraping logic separate from data-cleaning logic.
If you want, I can also show:
- a JavaScript/Node.js version
- a Pydantic/JSON Schema-based approach
- or a production-ready pipeline design for large scrapers.