Prompt
Kafka messages not being consumed fast enough, how do I troubleshoot this?
Latest observation
Here’s a practical way to troubleshoot Kafka consumers that aren’t keeping up. The goal is to figure out whether the bottleneck is in producer rate, broker/partition layout, consumer config, consumer code, downstream dependency, or rebalance/lag behavior.
1) Confirm the problem: is it real lag or just delayed processing?
Check:
- Consumer lag per partition
- Message ingestion rate vs consumer processing rate
- Whether lag is steadily increasing or spiky
- Whether only some partitions are behind
Useful commands:
kafka-consumer-groups --bootstrap-server <broker> --describe --group <group>- Or your monitoring stack (Prometheus/Grafana, Confluent Control Center, Burrow, etc.)
What to look for:
- One partition with high lag = likely hot partition or uneven key distribution
- All partitions behind = consumer throughput or broker-side issue
- Lag increases during deployments = rebalances or consumer restarts
2) Check partitioning and parallelism
Kafka consumer throughput is often limited by partition count.
Rules of thumb:
- A consumer group can process a topic with at most one active consumer per partition
- If you have 3 partitions, more than 3 consumers in the group won’t increase throughput
- If a single partition is overloaded, you cannot scale that partition horizontally without changing the keying/partitioning strategy
Questions:
- How many partitions does the topic have?
- How many consumers are in the group?
- Is the workload skewed by key?
If one partition is a hotspot:
- Revisit the partition key
- Consider increasing partitions
- Consider spreading “hot” keys
- Be careful: increasing partitions can change ordering guarantees for keyed messages
3) Check consumer application bottlenecks
Common causes:
- Slow downstream calls (DB, HTTP APIs, caches)
- Heavy message deserialization
- Large batch processing with too much work per poll
- Synchronous processing in the poll thread
- Lock contention / thread pool exhaustion
- GC pauses or CPU saturation
Things to inspect:
- CPU, memory, GC, and thread counts on consumer hosts
- Latency of external dependencies
- Time spent per message or per batch
- Whether processing happens on the same thread that calls
poll()
Best practice:
- Keep the poll loop responsive
- Offload work to a worker pool if needed
- Measure “time to process” separately from “time to fetch”
4) Review consumer configuration
Important settings:
Fetch/poll settings
max.poll.records
Higher values let each poll return more records, but too high can cause long processing delays.fetch.min.bytes/fetch.max.wait.ms
Affects batching and latency.max.partition.fetch.bytes
Too low can limit throughput for large messages.session.timeout.ms/heartbeat.interval.ms
Misconfiguration can cause rebalances.
Liveness/processing settings
max.poll.interval.ms
If processing one batch takes too long, the consumer is considered dead and rebalances happen.enable.auto.commit
Auto-commit can mask processing issues and cause duplicates or lag confusion.auto.offset.reset
Not a throughput issue, but useful to confirm behavior if offsets are missing.
What to check:
- Are you processing fewer records per poll than expected?
- Are you hitting
max.poll.interval.msand triggering rebalances? - Are commits frequent enough and aligned with your processing model?
5) Look for rebalances
Frequent rebalances can destroy throughput.
Symptoms:
- Logs show group rejoin/rebalance messages
- Consumers stop processing briefly and lag grows
- Partitions move between consumers often
Common causes:
- Consumer crashes or slow heartbeats
- Processing takes too long between polls
- Network instability
- Too many consumers starting/stopping
- Static membership not used where appropriate
Check logs for:
RebalanceInProgressRevoking previously assigned partitionsLost connectionMax poll interval exceeded
Mitigations:
- Ensure consumer code calls
poll()regularly - Reduce batch processing time
- Increase
max.poll.interval.msif appropriate - Use cooperative rebalancing if supported and suitable
- Consider static group membership to reduce churn
6) Validate broker-side health
Sometimes the consumer is fine and the cluster is the bottleneck.
Check:
- Broker CPU, disk I/O, network
- Controller stability
- Under-replicated partitions
- ISR shrink/expand events
- Request latency and throttling
- Replication lag
If brokers are overloaded:
- Consumers may fetch slowly
- Latency increases
- Produce/fetch requests get throttled
Also inspect:
- Topic replication factor
- Partition distribution across brokers
- Whether one broker hosts too many “hot” partitions
7) Check message size and compression
Large messages reduce throughput.
Look at:
- Average and max message size
- Compression type
- Deserialization cost
If messages are large:
- Consider smaller payloads
- Store large blobs elsewhere and send references
- Increase
max.partition.fetch.bytesif consumers are fetching large records
8) Measure end-to-end processing time
Add metrics around:
- Time spent in
poll() - Time to process a record/batch
- Time to commit offsets
- Time spent waiting on DB/API calls
- Queue depth in any internal worker pool
Useful metrics:
- records consumed/sec
- records processed/sec
- poll-to-process latency
- consumer lag by partition
- rebalance count
- commit latency
- downstream dependency latency
9) Common patterns and likely fixes
Pattern: Lag on one partition only
Likely:
- Hot key / uneven partitioning
Fix:
- Change partition key strategy
- Increase partitions
- Rebalance workload
Pattern: Lag rises during load spikes
Likely:
- Consumer CPU/thread pool saturated
- Downstream dependency slow
Fix:
- Scale consumers
- Optimize processing
- Add buffering/backpressure
- Increase downstream capacity
Pattern: Consumers keep rebalancing
Likely:
- Poll loop blocked too long
- Consumer crashes
- Misconfigured timeouts
Fix:
- Reduce per-poll work
- Use worker threads
- Tune timeouts
Pattern: Plenty of consumers but no throughput gain
Likely:
- Too few partitions
Fix:
- Increase partition count
- Re-key data if necessary
10) A simple troubleshooting checklist
- Check lag by partition
- Confirm partition count vs consumer count
- Verify consumer CPU/memory/GC
- Check downstream dependency latency
- Review consumer logs for rebalances/timeouts
- Validate broker health and throttling
- Inspect consumer configs (
max.poll.records,max.poll.interval.ms, fetch sizes) - Measure message size and serialization overhead
11) If you want a quick diagnostic starting point
Send me:
- Topic partition count
- Number of consumers in the group
- Consumer lag output
- Consumer config
- Whether processing is sync or async
- Any downstream dependencies
- Broker health/metrics if available
And I can help you narrow it down quickly.
If you want, I can also give you:
- a step-by-step Kafka lag debugging runbook
- a consumer config tuning guide
- or help interpret your consumer group describe output