Prompt
pipeline freshness monitoring
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
Pipeline freshness monitoring is the practice of checking whether a data pipeline is still producing recent, expected data on time. It helps detect delays, failures, and stalled upstream sources before they affect dashboards, ML models, or downstream jobs.
What it typically tracks
- Last successful run time
- Data arrival time / event time
- End-to-end latency from source to destination
- Row counts or volume anomalies
- Missing partitions / missing files
- SLA breaches for scheduled deliveries
Common freshness checks
- Time since last data update
- Example: “The
orderstable should have data no older than 15 minutes.”
- Example: “The
- Partition completeness
- Example: “Today’s hourly partition must exist by 10:15.”
- Watermark lag
- Example: compare max event timestamp in target vs current time
- Run success + recency
- A successful pipeline run isn’t enough if it produced stale data
- Downstream freshness
- Verify data is fresh where consumers actually read it
Alerting patterns
- Alert when freshness exceeds a threshold
- Warn before hard failure thresholds
- Separate alerts for:
- pipeline failure
- data staleness
- abnormal volume drops/spikes
Metrics and signals
freshness_lag_minutesdata_delay_secondsmax_event_timestamplast_successful_ingestionmissing_partition_count
Good design practices
- Define freshness per dataset, not globally
- Use business-relevant thresholds
- Exclude expected delays like weekends/holidays/batch windows
- Combine freshness with quality checks to avoid false confidence
- Monitor at the consumer-facing layer when possible
Example
If a sales table is expected to update every 30 minutes:
- check whether the newest row has a timestamp within the last 30 minutes
- alert if lag > 45 minutes
- page if lag > 2 hours
If you want, I can also provide:
- a SQL example for freshness checks
- a monitoring architecture
- or a dbt/Great Expectations/Fivetran/Airflow implementation pattern.