Prompt

How do I schedule crawls and store the output cleanly?

Technology / Cloud Infrastructure1 observationLast seen Oct 4, 2026

Latest observation

Oct 4, 2026OpenAI APIWeb search: off

A clean way to do this is to split the problem into two parts:

  1. Scheduling: decide when crawls run
  2. Storage: decide how each crawl’s results are written and versioned

1) Scheduling crawls

Common options:

  • Cron: simplest if you just need periodic runs
    • Example: run every night at 2 AM
    • Good for small/simple jobs
  • Task queue + scheduler:
    • Python: Celery + Celery Beat
    • Node: BullMQ + repeatable jobs
    • Good if crawls are distributed or take variable time
  • Workflow orchestrators:
    • Airflow, Prefect, Dagster
    • Best if you have many crawls, dependencies, retries, monitoring

Good scheduling practices

  • Add timeouts so a crawl can’t hang forever
  • Add retries with backoff for transient failures
  • Use concurrency limits so you don’t overload target sites or your system
  • Track run metadata: start time, end time, status, error message, pages fetched

2) Storing crawl output cleanly

A clean storage pattern is:

A. Separate raw output from normalized data

Store:

  • Raw crawl artifacts
    • HTML, JSON, screenshots, response headers, logs
  • Parsed/cleaned records
    • Structured data extracted from raw pages

This helps with:

  • debugging
  • reprocessing later if your parser changes
  • auditing what was actually fetched

B. Use a run-based folder/object layout

For file/object storage, structure by crawl run:

crawls/
  site_a/
    2026-10-04T02-00-00Z/
      raw/
        page1.html
        page2.html
      parsed/
        items.jsonl
      metadata.json

Or in object storage like S3/GCS:

s3://bucket/crawls/site_a/run_id=20261004T020000Z/raw/...
s3://bucket/crawls/site_a/run_id=20261004T020000Z/parsed/...
s3://bucket/crawls/site_a/run_id=20261004T020000Z/metadata.json

C. Prefer append-only, immutable outputs

Don’t overwrite old runs. Instead:

  • create a new run_id for each crawl
  • store results in a new location
  • keep metadata about the crawl job

This makes backfills and audits much easier.

D. Use a database for indexing, files for bulk data

A very common pattern:

  • Database: crawl runs, URLs, statuses, item identifiers, timestamps
  • File/object storage: raw HTML, JSONL exports, large blobs

For example:

  • PostgreSQL tables:
    • crawl_runs
    • crawl_pages
    • crawl_items
  • S3/local disk for actual page content

Recommended schema idea

crawl_runs

  • id
  • site
  • scheduled_at
  • started_at
  • finished_at
  • status (running, success, failed)
  • items_count
  • error

crawl_pages

  • run_id
  • url
  • status_code
  • fetch_time_ms
  • content_hash
  • raw_path

crawl_items

  • run_id
  • item_id
  • source_url
  • data_json
  • parsed_at

File format recommendations

  • JSONL for records: easy to stream and append
  • Parquet if you need analytics and large-scale processing
  • CSV only for simple exports
  • HTML/har files for raw page preservation

Example clean pipeline

  1. Scheduler triggers crawl at fixed time
  2. Crawl runner creates a new run_id
  3. Each fetched page is saved as raw HTML/JSON
  4. Parser extracts structured records into JSONL or DB
  5. Metadata and statuses are written to DB
  6. A final summary file is written for the run

Minimal best-practice checklist

  • Unique run_id for every crawl
  • Separate raw and parsed outputs
  • Store metadata for each run
  • Use append-only storage
  • Add retries, timeouts, and logging
  • Keep structured records in DB, large files in object storage

If you want, I can also give you:

  • a Python example using cron + SQLite/Postgres
  • a Prefect/Airflow example
  • a folder structure and database schema tailored to your crawler

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.