Prompt
How do I monitor pipeline freshness and failures?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To monitor pipeline freshness and failures, track two things:
- Is the pipeline running on time?
- Did it complete successfully?
1) Monitor freshness
Freshness means “how recent is the latest successful data in the destination?”
Common ways to measure
- Last successful run time
- Data lag / staleness
- Age of newest record
- Freshness SLA breaches
Example: “must be updated within 15 minutes”
What to alert on
- No successful run within expected interval
- Freshness exceeds threshold
- Downstream table hasn’t changed in X minutes/hours
- Source data timestamp is older than expected
Example metrics
time_since_last_successlatest_event_timestamp - now()pipeline_delay_minutes
2) Monitor failures
Failures are usually easier: detect when the pipeline job exits unsuccessfully or retries too many times.
What to capture
- Job/pipeline status: success, failed, canceled
- Error codes/messages
- Retry count
- Failed step/stage/task
- Duration anomalies
Alert on
- Any failed run
- Repeated failures in a short window
- Partial failures in critical stages
- Job duration much longer than normal
3) Practical setup
A good monitoring setup usually includes:
- Job scheduler alerts: failed cron/Airflow/Prefect/Dagster jobs
- Logs: centralized logging with searchable errors
- Metrics: run success rate, duration, freshness lag
- Dashboards: recent runs, last success, lag trends
- Paging/notifications: email, Slack, PagerDuty
4) Recommended alerts
For production pipelines, set alerts for:
- Failure immediately
- Freshness SLA breached
- No run detected
- Success but data unchanged unexpectedly (if relevant)
- Run duration exceeds threshold
5) Example rule
If your pipeline runs every 30 minutes:
- Alert if the latest successful run is older than 45 minutes
- Alert if any run fails
- Alert if 3 consecutive retries fail
- Alert if runtime is 2× normal
6) If you use a pipeline tool
Most orchestration tools have built-in monitoring:
- Airflow: DAG/task failures, SLA misses
- Prefect/Dagster: flow/job state monitoring
- dbt: tests, run failures, exposures
- Data warehouse: table freshness checks, query history
If you want, I can give you a tool-specific setup for Airflow, dbt, Prefect, Dagster, or a cloud stack like AWS/GCP/Azure.