Prompt
How do I build a pipeline that retries failed jobs automatically?
Latest observation
Oct 8, 2026
OpenAI APIWeb search: off
To build a pipeline that retries failed jobs automatically, you usually want three pieces:
- Detect failures
- Retry with limits
- Make jobs idempotent so retries don’t cause duplicate side effects
Basic pattern
1) Wrap each job with retry logic
Use a retry policy that includes:
- Max retries: e.g. 3–5 attempts
- Backoff: wait longer between attempts
- Jitter: add randomness to avoid retry storms
- Retry only transient failures: timeouts, rate limits, temporary network errors
Example retry strategy:
- Attempt 1 immediately
- Attempt 2 after 10 seconds
- Attempt 3 after 30 seconds
- Attempt 4 after 2 minutes
- Then fail and send to dead-letter/manual review
2) Track job state
Your pipeline should store:
- job ID
- attempt count
- last error
- status: queued / running / failed / succeeded / dead-lettered
This can be in a database, queue metadata, or orchestrator state.
3) Use a scheduler or queue worker
A worker process should:
- pull jobs from a queue
- execute them
- on failure, decide whether to retry
- requeue the job with delay/backoff
Common implementations
If you’re using a workflow orchestrator
Tools like:
- Airflow
- Prefect
- Dagster
- Argo Workflows
often already support retries per task. For example:
- Airflow:
retries,retry_delay - Prefect: retry policies on tasks
- Argo:
retryStrategy
If you’re using a message queue
With systems like:
- Celery
- RabbitMQ
- SQS
- Kafka (usually with a retry topic pattern)
you can implement:
- retry queues
- delay queues
- dead-letter queues
Important design rules
Make tasks idempotent
If a retry happens, the task should be safe to run twice.
Examples:
- Use unique request IDs
- Check whether output already exists before writing
- Use upserts instead of inserts when possible
Retry only the right errors
Do not retry:
- validation errors
- bad input
- permission errors
- logic bugs
Do retry:
- timeouts
- connection resets
- HTTP 429 / 5xx
- temporary downstream outages
Add a dead-letter queue
After too many failures, move the job to:
- a dead-letter queue
- a quarantine table
- manual inspection
This prevents infinite retry loops.
Simple pseudocode
max_retries = 3
def run_job(job):
for attempt in range(max_retries + 1):
try:
execute(job)
mark_success(job)
return
except TransientError as e:
if attempt == max_retries:
mark_failed(job, e)
send_to_dead_letter(job)
return
sleep(backoff(attempt))
except PermanentError as e:
mark_failed(job, e)
send_to_dead_letter(job)
return
Recommended architecture
- Queue for pending jobs
- Worker service to process jobs
- Retry policy with exponential backoff
- State store for attempt counts and status
- Dead-letter queue for exhausted jobs
- Alerts/monitoring for repeated failures
Monitoring
Track:
- retry rate
- failure rate
- average attempts per job
- dead-letter count
- time to recovery
Alert when retries spike, since that usually means a downstream service is unstable.
If you want, I can also show you:
- a Python/Celery example
- an Airflow DAG with retries
- a cloud-native design using SQS/Lambda or Kafka