Prompt

How do I normalize scraped page data into JSON?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

A good way to normalize scraped page data into JSON is to separate the process into three steps:

  1. Extract raw values from the page
  2. Clean/standardize those values
  3. Map them into a consistent JSON structure

1. Define your target JSON schema first

Decide what every record should look like, even if some fields are missing.

Example:

{
  "title": "",
  "price": null,
  "currency": "USD",
  "url": "",
  "image": "",
  "description": "",
  "inStock": false,
  "source": "example.com",
  "scrapedAt": "2026-10-04T12:00:00Z"
}

Having a fixed schema makes normalization much easier.


2. Clean the raw values

Scraped data is usually messy: extra spaces, currency symbols, relative URLs, inconsistent casing, etc.

Typical normalization tasks:

  • Trim whitespace
  • Convert text to lowercase or title case where appropriate
  • Parse numbers from strings
  • Convert prices into numeric values
  • Resolve relative URLs into absolute URLs
  • Normalize dates into ISO 8601
  • Turn “yes/no”, “in stock/out of stock” into booleans
  • Replace missing or invalid values with null

Example:

  • " $1,299.99 " → 1299.99
  • " /product/123 " → "https://site.com/product/123"
  • "Available" → true

3. Use a consistent transformation function

Here’s a simple Python example:

from urllib.parse import urljoin
from datetime import datetime
import re

def normalize_price(text):
    if not text:
        return None
    cleaned = re.sub(r"[^\d.,]", "", text).replace(",", "")
    try:
        return float(cleaned)
    except ValueError:
        return None

def normalize_bool(text):
    if not text:
        return False
    return text.strip().lower() in {"yes", "true", "available", "in stock"}

def normalize_record(raw, base_url, source):
    return {
        "title": raw.get("title", "").strip() or None,
        "price": normalize_price(raw.get("price")),
        "currency": raw.get("currency", "USD"),
        "url": urljoin(base_url, raw.get("url", "")),
        "image": urljoin(base_url, raw.get("image", "")) if raw.get("image") else None,
        "description": raw.get("description", "").strip() or None,
        "inStock": normalize_bool(raw.get("stock")),
        "source": source,
        "scrapedAt": datetime.utcnow().isoformat() + "Z"
    }

4. Handle missing and inconsistent fields

Scraped pages often vary. Your normalization should:

  • Use defaults for optional fields
  • Set missing fields to null rather than empty strings
  • Coalesce multiple possible source fields

Example:

title = raw.get("title") or raw.get("h1") or raw.get("name")

5. Deduplicate and validate

Before saving JSON:

  • Remove duplicates based on an ID, URL, or normalized title
  • Validate types and required fields
  • Log or quarantine bad records

If possible, define a schema using:

  • Python: pydantic, jsonschema
  • JavaScript: zod, ajv

6. Example normalized output

Raw scraped data:

{
  "title": "  Blue Sneakers ",
  "price": "$89.99",
  "url": "/products/blue-sneakers",
  "stock": "In Stock"
}

Normalized JSON:

{
  "title": "Blue Sneakers",
  "price": 89.99,
  "currency": "USD",
  "url": "https://example.com/products/blue-sneakers",
  "image": null,
  "description": null,
  "inStock": true,
  "source": "example.com",
  "scrapedAt": "2026-10-04T12:00:00Z"
}

Recommended approach

If you’re scraping at scale, use this structure:

  • Parser: extracts raw data from HTML
  • Normalizer: converts raw fields into canonical values
  • Validator: checks schema and types
  • Serializer: writes JSON

This keeps scraping logic separate from data-cleaning logic.

If you want, I can also show:

  • a JavaScript/Node.js version
  • a Pydantic/JSON Schema-based approach
  • or a production-ready pipeline design for large scrapers.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.