Prompt
unhappy with Celery task retries not behaving predictably
Latest observation
Celery retries can feel unpredictable when a few common pieces aren’t aligned. The good news is that retries are usually behaving exactly as configured — just not how we expect at first.
Here are the main things to check:
1) Are you using self.retry() inside a bound task?
For Celery retries to work reliably, the task usually needs to be bound:
@app.task(bind=True, max_retries=3)
def my_task(self, ...):
try:
...
except SomeError as exc:
raise self.retry(exc=exc, countdown=10)
If the task isn’t bind=True, self.retry() won’t work.
2) Retry countdown vs. ETA vs. exponential backoff
If you set both countdown and retry_backoff, or use defaults you didn’t mean to, the timing may look inconsistent.
Example:
@app.task(bind=True, autoretry_for=(Exception,), retry_backoff=True, retry_jitter=True)
def my_task(self):
...
This can produce variable delays because of jitter. If you want predictable timing, disable jitter:
retry_jitter=False
3) autoretry_for can hide the actual flow
With autoretry_for, Celery catches exceptions and retries automatically. That’s convenient, but it can make it harder to see why/when retries happen.
If you need full control, prefer explicit self.retry() and log the exception before retrying.
4) max_retries and task state can be misleading
Retries increment per task execution. Depending on how the task is invoked and whether it’s a new message or true retry, it may look like retries reset or skip.
Check:
self.request.retries- task id consistency
- whether the task is being re-queued externally
5) The worker may be killing the task before retry logic completes
If the task is long-running and the worker hits:
soft_time_limittime_limit- worker shutdown
acks_late/ visibility timeout issues
you may see what looks like “failed retry behavior” when the worker is actually terminating the job.
6) Exceptions matter
If you call self.retry(exc=exc) Celery will re-raise a Retry exception internally. That’s normal.
But if you catch too broadly or swallow exceptions, retries can stop happening:
try:
...
except Exception:
self.retry(countdown=10) # works, but if you later swallow errors, behavior changes
Also make sure you’re not doing:
except Exception as exc:
logger.exception(exc)
return
That would prevent retrying.
7) Broker / result backend semantics
Some “unpredictability” comes from the broker rather than Celery logic:
- RabbitMQ vs Redis behave differently
- Redis visibility timeout can cause duplicate delivery or delayed redelivery
- network interruptions can make a task appear retried when it was actually redelivered
8) Logging the retry state helps a lot
Add explicit logging:
@app.task(bind=True, max_retries=5)
def my_task(self):
logger.info("retry=%s task_id=%s", self.request.retries, self.request.id)
try:
...
except Exception as exc:
logger.warning("Retrying task_id=%s attempt=%s", self.request.id, self.request.retries)
raise self.retry(exc=exc, countdown=10)
This often reveals whether the issue is:
- retries never triggered,
- retries triggered but with unexpected delay,
- a new task was submitted instead of a retry,
- or the worker/broker redelivered the task.
If you want, I can help you debug your exact case
Paste:
- the task code,
- Celery config,
- broker/backend,
- and what “unpredictable” looks like
and I’ll help pinpoint the cause.