Prompt

unhappy with Celery task retries not behaving predictably

Technology / Cloud Infrastructure1 observationLast seen Oct 7, 2026

Latest observation

Oct 7, 2026OpenAI APIWeb search: off

Celery retries can feel unpredictable when a few common pieces aren’t aligned. The good news is that retries are usually behaving exactly as configured — just not how we expect at first.

Here are the main things to check:

1) Are you using self.retry() inside a bound task?

For Celery retries to work reliably, the task usually needs to be bound:

@app.task(bind=True, max_retries=3)
def my_task(self, ...):
    try:
        ...
    except SomeError as exc:
        raise self.retry(exc=exc, countdown=10)

If the task isn’t bind=True, self.retry() won’t work.


2) Retry countdown vs. ETA vs. exponential backoff

If you set both countdown and retry_backoff, or use defaults you didn’t mean to, the timing may look inconsistent.

Example:

@app.task(bind=True, autoretry_for=(Exception,), retry_backoff=True, retry_jitter=True)
def my_task(self):
    ...

This can produce variable delays because of jitter. If you want predictable timing, disable jitter:

retry_jitter=False

3) autoretry_for can hide the actual flow

With autoretry_for, Celery catches exceptions and retries automatically. That’s convenient, but it can make it harder to see why/when retries happen.

If you need full control, prefer explicit self.retry() and log the exception before retrying.


4) max_retries and task state can be misleading

Retries increment per task execution. Depending on how the task is invoked and whether it’s a new message or true retry, it may look like retries reset or skip.

Check:

  • self.request.retries
  • task id consistency
  • whether the task is being re-queued externally

5) The worker may be killing the task before retry logic completes

If the task is long-running and the worker hits:

  • soft_time_limit
  • time_limit
  • worker shutdown
  • acks_late / visibility timeout issues

you may see what looks like “failed retry behavior” when the worker is actually terminating the job.


6) Exceptions matter

If you call self.retry(exc=exc) Celery will re-raise a Retry exception internally. That’s normal.

But if you catch too broadly or swallow exceptions, retries can stop happening:

try:
    ...
except Exception:
    self.retry(countdown=10)   # works, but if you later swallow errors, behavior changes

Also make sure you’re not doing:

except Exception as exc:
    logger.exception(exc)
    return

That would prevent retrying.


7) Broker / result backend semantics

Some “unpredictability” comes from the broker rather than Celery logic:

  • RabbitMQ vs Redis behave differently
  • Redis visibility timeout can cause duplicate delivery or delayed redelivery
  • network interruptions can make a task appear retried when it was actually redelivered

8) Logging the retry state helps a lot

Add explicit logging:

@app.task(bind=True, max_retries=5)
def my_task(self):
    logger.info("retry=%s task_id=%s", self.request.retries, self.request.id)
    try:
        ...
    except Exception as exc:
        logger.warning("Retrying task_id=%s attempt=%s", self.request.id, self.request.retries)
        raise self.retry(exc=exc, countdown=10)

This often reveals whether the issue is:

  • retries never triggered,
  • retries triggered but with unexpected delay,
  • a new task was submitted instead of a retry,
  • or the worker/broker redelivered the task.

If you want, I can help you debug your exact case

Paste:

  • the task code,
  • Celery config,
  • broker/backend,
  • and what “unpredictable” looks like
    and I’ll help pinpoint the cause.

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.