Prompt
How do I fix push notification retries not working?
Latest observation
If push notification retries aren’t working, the fix usually depends on where the retry logic lives:
1) Confirm who is responsible for retries
There are 3 common layers:
- Your app/backend: you queue and resend failed notifications
- Push provider (FCM/APNs/OneSignal/etc.): it may retry transient delivery failures internally
- Your worker/job system: it handles retrying failed API calls to the push service
If retries “don’t happen,” it’s often because the failure is happening outside the retryable layer.
2) Check the failure type
Retries should usually happen only for transient errors, like:
- network timeouts
- 5xx server errors
- rate limiting (
429) - temporary service unavailable
Retries usually should not happen for permanent errors, like:
- invalid device token
- expired token
- unauthorized credentials
- malformed payload
If your system treats all failures as permanent, retries won’t trigger.
3) Verify your retry policy
Make sure you actually have:
- a retry count set
- exponential backoff enabled
- a delay between attempts
- retry conditions mapped correctly
Example retry policy:
- attempt 1 immediately
- retry after 1 min
- retry after 5 min
- retry after 30 min
- then give up / dead-letter
Also ensure the max retry count isn’t set to 0 or 1.
4) Check the job queue / worker configuration
If you use a queue (Sidekiq, Celery, BullMQ, SQS, etc.), confirm:
- failed jobs are not being deleted immediately
- retries are enabled for that queue/worker
- the worker process is running
- visibility timeout / ack behavior isn’t causing premature loss
- jobs are not going to a dead-letter queue too early
A common issue is: the send job fails, but the worker catches the error and logs it without rethrowing it, so the queue thinks it succeeded.
5) Inspect provider response codes
Look at the exact response from APNs/FCM/your provider:
401/403→ auth issue, no retry until fixed400→ bad request, usually no retry404/unregistered token → stop sending to that token429→ retry with backoff500/503→ retry
If you’re not parsing provider errors properly, your retry logic may be skipping valid retry cases.
6) Make sure retries aren’t being deduplicated away
Some systems de-duplicate notification requests by:
- message ID
- idempotency key
- job fingerprint
- message hash
If retries reuse the same ID and your queue treats it as a duplicate, the retry may be suppressed.
7) Check timeouts
Sometimes retries “don’t work” because:
- the request hangs too long
- the worker times out and crashes
- the retry framework never gets the failure callback
Set:
- reasonable HTTP client timeout
- job timeout longer than request timeout
- retry on timeout exceptions
8) Look for swallowed exceptions
A very common bug:
try {
sendPush();
} catch (e) {
console.error(e);
// missing: throw e;
}
If you catch the error and don’t rethrow it, the retry system may think the job succeeded.
9) Validate credentials and environment
Retries won’t help if the root cause is:
- expired APNs auth key
- invalid FCM service account
- wrong bundle ID / package name
- wrong app environment (sandbox vs production)
In those cases, fix the configuration first.
10) Add observability
To debug quickly, log:
- notification ID
- attempt number
- provider response code/body
- exception type
- retry scheduled time
- final outcome
This will show whether retries are:
- never scheduled
- scheduled but not executed
- executed but failing again
- discarded as permanent failures
Quick checklist
- Retry enabled in worker/queue
- Errors are rethrown after catch
- Backoff configured
- Only transient errors retried
- No deduplication blocking retries
- Credentials valid
- Timeouts set correctly
- Worker process healthy
If you want, I can help you troubleshoot this for a specific stack like FCM/APNs, Node/BullMQ, Python/Celery, AWS SNS/SQS, or OneSignal.