Prompt
Troubleshoot workers not picking up jobs from Redis
Latest observation
Here’s a practical checklist to troubleshoot workers not picking up jobs from Redis.
1) Confirm jobs are actually being enqueued
- Check the producer/app logs for enqueue success.
- Inspect Redis to see whether jobs are landing in the expected key/queue:
redis-cliKEYS *or betterSCAN 0LLEN <queue>if it’s a list-based queueZRANGE <queue> 0 -1 WITHSCORESif it uses a sorted set
- Verify the producer is writing to the same Redis instance/database as the workers.
2) Confirm workers are connected to the same Redis
Common issues:
- Wrong host/port
- Wrong Redis DB index
- Wrong password/ACL
- Different environment variables between producer and worker
- Sentinel/Cluster misconfiguration
Check:
- Worker startup logs for connection details
- Redis
INFO clients/CLIENT LIST - Network reachability from the worker machine/container:
redis-cli -h <host> -p <port> ping
3) Check the queue name matches exactly
A very common issue is mismatch in:
- queue name
- prefix
- namespace
- environment-specific naming
Example problems:
- Producer pushes to
queue:email - Worker listens on
emails - One side uses a prefix like
prod:and the other doesn’t
4) Verify workers are actually running and polling
- Confirm process/container is alive
- Look for crashes/restarts
- Check logs for:
- “listening on queue…”
- “subscribed”
- “connected”
- If using a process manager:
- systemd:
systemctl status <service> - PM2:
pm2 status - Docker:
docker ps,docker logs <container>
- systemd:
5) Look for stalled or blocked workers
Some queues require long polling or blocking reads.
- Worker may be stuck in a crash loop
- Worker may be blocked on an external dependency
- Worker may be waiting on a lock or rate limiter
- Worker may be consuming jobs but failing before ack/complete
Check:
- worker logs
- retry/dead-letter queues
- metrics for in-flight jobs
6) Inspect job states
Depending on the queue library, jobs may be:
- waiting
- active
- delayed
- failed
- completed
- stalled
If jobs are in failed or delayed, they may not be immediately picked up.
If jobs are in active forever, the worker may be hanging or not acknowledging completion.
7) Check Redis memory and eviction behavior
If Redis is under memory pressure:
- keys may be evicted
- writes may fail
- jobs may disappear unexpectedly
Check:
INFO memoryCONFIG GET maxmemoryCONFIG GET maxmemory-policy- Redis logs for OOM/eviction
8) Check for blocking commands / connection issues
If workers use blocking operations like BLPOP, BRPOP, or XREAD BLOCK:
- interrupted connections can stop consumption
- idle timeouts or proxies may drop connections
- TLS / proxy / load balancer settings may interfere
Look for:
- reconnect loops
- socket timeout errors
READONLYerrors in clusterMOVED/ASKerrors in cluster
9) Verify serialization / payload format
If the queue stores JSON or encoded payloads:
- malformed payloads can cause worker parse failures
- incompatible versions of producer/worker can break deserialization
- large payloads may exceed limits
Check worker logs for parse/validation errors.
10) Check rate limits, concurrency, and visibility timeouts
Depending on the queue system:
- concurrency may be set to 0 or 1 unexpectedly
- rate limits can throttle pickup
- visibility timeouts too short/long can cause reprocessing or stalls
11) Check ACLs and permissions
With Redis ACLs:
- worker may be able to connect but not read/write the queue keys
- commands like
BLPOP,EVAL,XREAD,ZREMmay be denied
Use:
ACL LIST- worker logs for
NOPERM
12) Compare a known-good minimal consumer
Create a small test consumer to verify the queue is consumable independently of your app logic. If the minimal consumer works, the issue is likely in application code, config, or middleware.
Quick triage commands
redis-cli -h <host> -p <port> ping
redis-cli -h <host> -p <port> INFO clients
redis-cli -h <host> -p <port> INFO memory
redis-cli -h <host> -p <port> SCAN 0
redis-cli -h <host> -p <port> LLEN <queue>
redis-cli -h <host> -p <port> MONITOR
Use
MONITORonly briefly in non-production or low-traffic environments, since it is expensive.
Most common root causes
- Producer and worker point to different Redis instances
- Queue name mismatch
- Worker not running or crashing
- Jobs are failing immediately due to code errors
- Redis ACL/network issues
- Jobs are delayed/blocked/stalled rather than “missing”
If you want, I can help you debug this faster if you share:
- the Redis queue library/framework you use,
- worker logs,
- producer enqueue code,
- and the Redis connection config.