Prompt
How do I monitor data freshness and failed syncs in a pipeline?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To monitor data freshness and failed syncs in a pipeline, track both timeliness and execution health.
1) Monitor data freshness
Measure how old the data is compared to when it should have arrived.
Key metrics
- Last successful load time: timestamp of the most recent successful ingestion/sync.
- Data lag / freshness lag:
now() - last_successful_load_time - Expected update interval: how often the source should refresh (hourly, daily, etc.)
- Freshness SLA: maximum acceptable lag before data is considered stale
Alerts
Trigger an alert when:
- The freshness lag exceeds a threshold
- No new data arrives within the expected window
- A downstream table/view is older than allowed
Example
If a source updates every hour and the SLA is 90 minutes:
- Alert if
now - last_successful_load > 90 minutes
2) Monitor failed syncs
Track pipeline job execution status and error patterns.
Key metrics
- Failed sync count
- Failure rate: failed runs / total runs
- Consecutive failures
- Retries exhausted
- Error categories: auth, network, schema mismatch, quota, timeout
Alerts
Trigger alerts on:
- Any failed sync for critical pipelines
- N consecutive failures
- Spike in failures compared to baseline
- Failures in a row after previous success
3) Add pipeline observability
Store run metadata in a log or metrics table for each sync:
- pipeline name
- source/target
- run ID
- start/end time
- status
- rows read/written
- error message
- retry count
- latency
This makes it easy to build dashboards and alerting.
4) Practical checks
A good monitoring setup usually includes:
- Heartbeat checks: source or job sends periodic “I’m alive” signals
- Freshness checks: compare current time to latest data timestamp
- Completion checks: verify every scheduled sync ran and finished
- Volume checks: watch for unusual drops/spikes in row counts
- Schema checks: detect missing/renamed columns that cause sync failure
5) Example alert rules
- Freshness > 2× expected interval
- 3 failed syncs in a row
- No successful sync in 24 hours
- Row count drops by > 80% vs 7-day average
- Schema change detected
6) Tools you can use
Depending on your stack:
- Airflow / Dagster / Prefect: job state and retries
- dbt tests: freshness and schema checks
- Prometheus / Grafana: metrics and alerting
- CloudWatch / Stackdriver / Azure Monitor: managed monitoring
- Great Expectations / Soda: data quality and freshness tests
If you want, I can also give you:
- a SQL example for freshness monitoring,
- a dashboard design, or
- a sample alerting policy for Airflow/dbt/Prefect.