Prompt
How do I schedule crawls and store the output cleanly?
Latest observation
Oct 4, 2026
OpenAI APIWeb search: off
A clean way to do this is to split the problem into two parts:
- Scheduling: decide when crawls run
- Storage: decide how each crawl’s results are written and versioned
1) Scheduling crawls
Common options:
- Cron: simplest if you just need periodic runs
- Example: run every night at 2 AM
- Good for small/simple jobs
- Task queue + scheduler:
- Python: Celery + Celery Beat
- Node: BullMQ + repeatable jobs
- Good if crawls are distributed or take variable time
- Workflow orchestrators:
- Airflow, Prefect, Dagster
- Best if you have many crawls, dependencies, retries, monitoring
Good scheduling practices
- Add timeouts so a crawl can’t hang forever
- Add retries with backoff for transient failures
- Use concurrency limits so you don’t overload target sites or your system
- Track run metadata: start time, end time, status, error message, pages fetched
2) Storing crawl output cleanly
A clean storage pattern is:
A. Separate raw output from normalized data
Store:
- Raw crawl artifacts
- HTML, JSON, screenshots, response headers, logs
- Parsed/cleaned records
- Structured data extracted from raw pages
This helps with:
- debugging
- reprocessing later if your parser changes
- auditing what was actually fetched
B. Use a run-based folder/object layout
For file/object storage, structure by crawl run:
crawls/
site_a/
2026-10-04T02-00-00Z/
raw/
page1.html
page2.html
parsed/
items.jsonl
metadata.json
Or in object storage like S3/GCS:
s3://bucket/crawls/site_a/run_id=20261004T020000Z/raw/...
s3://bucket/crawls/site_a/run_id=20261004T020000Z/parsed/...
s3://bucket/crawls/site_a/run_id=20261004T020000Z/metadata.json
C. Prefer append-only, immutable outputs
Don’t overwrite old runs. Instead:
- create a new
run_idfor each crawl - store results in a new location
- keep metadata about the crawl job
This makes backfills and audits much easier.
D. Use a database for indexing, files for bulk data
A very common pattern:
- Database: crawl runs, URLs, statuses, item identifiers, timestamps
- File/object storage: raw HTML, JSONL exports, large blobs
For example:
- PostgreSQL tables:
crawl_runscrawl_pagescrawl_items
- S3/local disk for actual page content
Recommended schema idea
crawl_runs
idsitescheduled_atstarted_atfinished_atstatus(running,success,failed)items_counterror
crawl_pages
run_idurlstatus_codefetch_time_mscontent_hashraw_path
crawl_items
run_iditem_idsource_urldata_jsonparsed_at
File format recommendations
- JSONL for records: easy to stream and append
- Parquet if you need analytics and large-scale processing
- CSV only for simple exports
- HTML/har files for raw page preservation
Example clean pipeline
- Scheduler triggers crawl at fixed time
- Crawl runner creates a new
run_id - Each fetched page is saved as raw HTML/JSON
- Parser extracts structured records into JSONL or DB
- Metadata and statuses are written to DB
- A final summary file is written for the run
Minimal best-practice checklist
- Unique
run_idfor every crawl - Separate raw and parsed outputs
- Store metadata for each run
- Use append-only storage
- Add retries, timeouts, and logging
- Keep structured records in DB, large files in object storage
If you want, I can also give you:
- a Python example using cron + SQLite/Postgres
- a Prefect/Airflow example
- a folder structure and database schema tailored to your crawler