Prompt

Kafka messages not being consumed fast enough, how do I troubleshoot this?

Technology / Data Infrastructure1 observationLast seen Oct 3, 2026

Latest observation

Oct 3, 2026OpenAI APIWeb search: off

Here’s a practical way to troubleshoot Kafka consumers that aren’t keeping up. The goal is to figure out whether the bottleneck is in producer rate, broker/partition layout, consumer config, consumer code, downstream dependency, or rebalance/lag behavior.

1) Confirm the problem: is it real lag or just delayed processing?

Check:

  • Consumer lag per partition
  • Message ingestion rate vs consumer processing rate
  • Whether lag is steadily increasing or spiky
  • Whether only some partitions are behind

Useful commands:

  • kafka-consumer-groups --bootstrap-server <broker> --describe --group <group>
  • Or your monitoring stack (Prometheus/Grafana, Confluent Control Center, Burrow, etc.)

What to look for:

  • One partition with high lag = likely hot partition or uneven key distribution
  • All partitions behind = consumer throughput or broker-side issue
  • Lag increases during deployments = rebalances or consumer restarts

2) Check partitioning and parallelism

Kafka consumer throughput is often limited by partition count.

Rules of thumb:

  • A consumer group can process a topic with at most one active consumer per partition
  • If you have 3 partitions, more than 3 consumers in the group won’t increase throughput
  • If a single partition is overloaded, you cannot scale that partition horizontally without changing the keying/partitioning strategy

Questions:

  • How many partitions does the topic have?
  • How many consumers are in the group?
  • Is the workload skewed by key?

If one partition is a hotspot:

  • Revisit the partition key
  • Consider increasing partitions
  • Consider spreading “hot” keys
  • Be careful: increasing partitions can change ordering guarantees for keyed messages

3) Check consumer application bottlenecks

Common causes:

  • Slow downstream calls (DB, HTTP APIs, caches)
  • Heavy message deserialization
  • Large batch processing with too much work per poll
  • Synchronous processing in the poll thread
  • Lock contention / thread pool exhaustion
  • GC pauses or CPU saturation

Things to inspect:

  • CPU, memory, GC, and thread counts on consumer hosts
  • Latency of external dependencies
  • Time spent per message or per batch
  • Whether processing happens on the same thread that calls poll()

Best practice:

  • Keep the poll loop responsive
  • Offload work to a worker pool if needed
  • Measure “time to process” separately from “time to fetch”

4) Review consumer configuration

Important settings:

Fetch/poll settings

  • max.poll.records
    Higher values let each poll return more records, but too high can cause long processing delays.
  • fetch.min.bytes / fetch.max.wait.ms
    Affects batching and latency.
  • max.partition.fetch.bytes
    Too low can limit throughput for large messages.
  • session.timeout.ms / heartbeat.interval.ms
    Misconfiguration can cause rebalances.

Liveness/processing settings

  • max.poll.interval.ms
    If processing one batch takes too long, the consumer is considered dead and rebalances happen.
  • enable.auto.commit
    Auto-commit can mask processing issues and cause duplicates or lag confusion.
  • auto.offset.reset
    Not a throughput issue, but useful to confirm behavior if offsets are missing.

What to check:

  • Are you processing fewer records per poll than expected?
  • Are you hitting max.poll.interval.ms and triggering rebalances?
  • Are commits frequent enough and aligned with your processing model?

5) Look for rebalances

Frequent rebalances can destroy throughput.

Symptoms:

  • Logs show group rejoin/rebalance messages
  • Consumers stop processing briefly and lag grows
  • Partitions move between consumers often

Common causes:

  • Consumer crashes or slow heartbeats
  • Processing takes too long between polls
  • Network instability
  • Too many consumers starting/stopping
  • Static membership not used where appropriate

Check logs for:

  • RebalanceInProgress
  • Revoking previously assigned partitions
  • Lost connection
  • Max poll interval exceeded

Mitigations:

  • Ensure consumer code calls poll() regularly
  • Reduce batch processing time
  • Increase max.poll.interval.ms if appropriate
  • Use cooperative rebalancing if supported and suitable
  • Consider static group membership to reduce churn

6) Validate broker-side health

Sometimes the consumer is fine and the cluster is the bottleneck.

Check:

  • Broker CPU, disk I/O, network
  • Controller stability
  • Under-replicated partitions
  • ISR shrink/expand events
  • Request latency and throttling
  • Replication lag

If brokers are overloaded:

  • Consumers may fetch slowly
  • Latency increases
  • Produce/fetch requests get throttled

Also inspect:

  • Topic replication factor
  • Partition distribution across brokers
  • Whether one broker hosts too many “hot” partitions

7) Check message size and compression

Large messages reduce throughput.

Look at:

  • Average and max message size
  • Compression type
  • Deserialization cost

If messages are large:

  • Consider smaller payloads
  • Store large blobs elsewhere and send references
  • Increase max.partition.fetch.bytes if consumers are fetching large records

8) Measure end-to-end processing time

Add metrics around:

  • Time spent in poll()
  • Time to process a record/batch
  • Time to commit offsets
  • Time spent waiting on DB/API calls
  • Queue depth in any internal worker pool

Useful metrics:

  • records consumed/sec
  • records processed/sec
  • poll-to-process latency
  • consumer lag by partition
  • rebalance count
  • commit latency
  • downstream dependency latency

9) Common patterns and likely fixes

Pattern: Lag on one partition only

Likely:

  • Hot key / uneven partitioning

Fix:

  • Change partition key strategy
  • Increase partitions
  • Rebalance workload

Pattern: Lag rises during load spikes

Likely:

  • Consumer CPU/thread pool saturated
  • Downstream dependency slow

Fix:

  • Scale consumers
  • Optimize processing
  • Add buffering/backpressure
  • Increase downstream capacity

Pattern: Consumers keep rebalancing

Likely:

  • Poll loop blocked too long
  • Consumer crashes
  • Misconfigured timeouts

Fix:

  • Reduce per-poll work
  • Use worker threads
  • Tune timeouts

Pattern: Plenty of consumers but no throughput gain

Likely:

  • Too few partitions

Fix:

  • Increase partition count
  • Re-key data if necessary

10) A simple troubleshooting checklist

  1. Check lag by partition
  2. Confirm partition count vs consumer count
  3. Verify consumer CPU/memory/GC
  4. Check downstream dependency latency
  5. Review consumer logs for rebalances/timeouts
  6. Validate broker health and throttling
  7. Inspect consumer configs (max.poll.records, max.poll.interval.ms, fetch sizes)
  8. Measure message size and serialization overhead

11) If you want a quick diagnostic starting point

Send me:

  • Topic partition count
  • Number of consumers in the group
  • Consumer lag output
  • Consumer config
  • Whether processing is sync or async
  • Any downstream dependencies
  • Broker health/metrics if available

And I can help you narrow it down quickly.

If you want, I can also give you:

  • a step-by-step Kafka lag debugging runbook
  • a consumer config tuning guide
  • or help interpret your consumer group describe output

Observation history

How did Obsurfable measure this prompt?

Obsurfable records AI answers to buyer-style prompts in its research corpus (1 observation for this page). Metrics are distributions over observations, not a single static ranking.

Which AI systems does Obsurfable collect answers from?

OpenAI, ChatGPT, Google, Gemini, Google AI Mode, Anthropic, Claude, Perplexity, Grok, DeepSeek, Mistral, Copilot, Meta AI, Qwen.

How can brands improve their AEO?

Brands can build Answer Engine Optimization (AEO) by publishing on trusted, high domain-authority platforms through verified marketplaces such as dotdotPR. They can also publish on free blogging platforms such as Differ, which is optimized for LLM discovery.

Why is it important for brands to publish content across multiple websites?

Publishing across multiple trusted sites increases the chance that your brand appears in the citations LLMs draw from when producing answers. Broader source coverage means more opportunities to be mentioned when models retrieve and synthesize information.

Want this interpreted for your brand?

Explorer is the free public corpus. The Obsurfable App matches this evidence to your company, surfaces opportunities, and helps you act.