Prompt

How do I collect public web data from multiple sites reliably?

Technology / Data infrastructure1 observationLast seen Sep 24, 2026

Latest observation

Sep 24, 2026OpenAI APIWeb search: off

To collect public web data from multiple sites reliably, think in terms of a pipeline: discover → fetch → parse → validate → store → monitor. The main challenge isn’t just getting the data once; it’s making the process robust against site changes, rate limits, and partial failures.

1) Start with the right source

Prefer, in this order:

  1. Official APIs
    Most reliable, least brittle, usually easiest to maintain.
  2. Feeds / structured exports
    RSS/Atom, CSV exports, sitemaps, downloadable datasets.
  3. HTML scraping
    Use when no API exists.
  4. Browser automation
    Only when pages are heavily JavaScript-driven or require interaction.

If you can avoid scraping, do.


2) Respect site policies and legal constraints

Before collecting:

  • Check robots.txt
  • Review Terms of Service
  • Don’t bypass authentication, paywalls, CAPTCHAs, or access controls
  • Rate-limit requests to avoid abuse
  • Identify your crawler with a clear User-Agent and contact info if appropriate

If the data is public, that doesn’t automatically mean unrestricted reuse or collection.


3) Design for reliability

Fetching

Use:

  • Timeouts
  • Retries with exponential backoff
  • Connection pooling
  • Caching
  • Concurrency limits per domain

Key idea: do not hammer all sites equally. Each site should have its own fetch policy.

Example retry strategy

  • Retry 3 times
  • Backoff: 1s, 2s, 4s
  • Retry only on transient failures:
    • 429
    • 500, 502, 503, 504
    • network timeouts

Do not blindly retry parsing errors.


4) Separate crawling from parsing

Build two layers:

  • Crawler/fetcher: gets raw HTML/JSON
  • Parser/extractor: turns raw content into structured data

Why:

  • You can re-parse old content when site markup changes
  • You can debug extraction independently
  • You can test parser changes on saved pages

Store raw responses when possible.


5) Normalize the data model

Different sites will name the same thing differently.

Create a canonical schema, for example:

  • source_site
  • source_url
  • scraped_at
  • title
  • price
  • currency
  • posted_at
  • author
  • body
  • external_id

Normalize:

  • dates to UTC
  • currencies to ISO codes
  • URLs to canonical form
  • text to consistent encoding
  • numbers to a standard numeric type

6) Make extraction resilient to markup changes

HTML changes constantly. Use strategies like:

  • Prefer stable attributes (data-*, itemprop, JSON-LD, schema.org)
  • Extract from structured embedded data when available
  • Use multiple selectors with fallbacks
  • Avoid brittle full-path CSS selectors when possible

Good idea:

  • First try JSON-LD / embedded JSON
  • Then semantic HTML
  • Then fallback selectors

7) Handle pagination, deduplication, and incremental updates

Pagination

Support:

  • page numbers
  • cursor-based pagination
  • “load more” endpoints

Deduplication

Deduplicate by:

  • canonical URL
  • source-specific ID
  • content hash
  • combination of fields

Incremental collection

Only fetch:

  • new pages
  • updated records
  • changed timestamps

This saves bandwidth and reduces load on sites.


8) Use concurrency carefully

For multiple sites:

  • Use per-domain rate limits
  • Use a worker queue
  • Keep a global concurrency cap
  • Prioritize smaller or slower domains lower

Example policy:

  • max 2 requests/sec per domain
  • max 20 concurrent requests overall
  • jitter in delays to avoid burst patterns

9) Monitor and alert

Reliability requires observability.

Track:

  • request success rate
  • HTTP status codes
  • latency
  • parsing success rate
  • number of extracted records
  • dedup rate
  • schema validation failures

Alert on:

  • sudden drop in records
  • spike in 403/429/5xx
  • parser exceptions
  • unexpected field null rates

A site changing its HTML often looks like a “successful fetch” but “zero extracted items.”


10) Validate data quality

Before storing or using the data:

  • enforce schema validation
  • check required fields
  • validate dates, URLs, numbers
  • detect obviously bad values
  • flag outliers

Examples:

  • price is negative
  • date is in the future by years
  • title is empty but page is otherwise valid

11) Store raw and structured data

A robust pattern is:

  • Raw layer: HTML/JSON as fetched
  • Processed layer: structured records
  • Curated layer: cleaned and deduplicated data

This helps with:

  • debugging
  • reprocessing after parser changes
  • audits and reproducibility

12) Prefer idempotent jobs

Your crawler should be safe to rerun.

That means:

  • re-fetching the same URL doesn’t duplicate data
  • storage uses upserts or deterministic keys
  • job state is checkpointed
  • failures can resume from the last good point

13) Recommended architecture

A practical setup:

  • Scheduler: decides what to crawl
  • Queue: pending URLs/tasks
  • Fetcher: downloads pages
  • Parser: extracts structured data
  • Validator: checks schema and quality
  • Database / object storage: stores raw + parsed data
  • Monitoring: logs and metrics
  • Config per site: rate limits, selectors, pagination rules, auth if allowed

14) Tools you can use

Depending on your stack:

Python

  • requests / httpx for HTTP
  • BeautifulSoup or lxml for parsing
  • scrapy for full crawling framework
  • playwright for JS-heavy pages
  • pydantic for validation

JavaScript/TypeScript

  • axios / fetch
  • cheerio
  • playwright
  • queue libraries like BullMQ

Infrastructure

  • PostgreSQL for structured storage
  • S3/object storage for raw pages
  • Redis for queues/rate limiting
  • Prometheus/Grafana for monitoring

15) A simple reliability checklist

For each site, ask:

  • Is there an API?
  • What is the rate limit?
  • What is the canonical identifier?
  • How do I detect updates?
  • What is the fallback if markup changes?
  • How will I validate extracted data?
  • How will I know if the site breaks tomorrow?

A minimal practical approach

If you’re starting from scratch:

  1. Make a list of target sites and their access methods
  2. Build one connector per site
  3. Store raw response bodies
  4. Parse into a shared schema
  5. Add retries/backoff and per-domain rate limits
  6. Add validation and deduplication
  7. Monitor record counts and error rates
  8. Revisit broken parsers quickly when pages change

If you want, I can also provide:

  • a Python scraping architecture template
  • a Scrapy project structure
  • a per-domain rate limiting example
  • or a sample schema for multi-site public data collection

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.