13. Monitoring Kafka
13.10 What actually breaks in production — Ch. 13 consolidated
| # | Symptom | Root cause | Fix |
|---|---|---|---|
| 1 | Kafka is down and no alert fired | The monitoring pipeline runs on Kafka | Separate monitoring system for Kafka, or cross-DC metric shipping |
| 2 | A JMX port became an RCE vector | Remote JMX enabled without security — "JMX not only allows a view into the state, it also allows CODE EXECUTION" | In-process agent (Jolokia/MX4J); keep remote JMX disabled |
| 3 | All broker metrics look fine but clients can't connect | Broker metrics are subjective | Synthetic clients (Xinfra Monitor) and client-side metrics |
| 4 | Team ignores alerts | Alert fatigue from too many metric-based alerts with hand-tuned thresholds | SLO/burn-rate alerting; "Check Engine light" principle |
| 5 | URP alerts fire constantly for benign reasons | "the URP metric can frequently be NONZERO for benign reasons" | The book retracts its own prior advice — don't alert on URP; use it for diagnosis |
| 6 | SLO breached before anyone noticed | Alerted on the SLO itself, which is a lagging, week-scale indicator | Burn-rate alerting (0.1% → ticket at 0.4% → page at 2%) |
| 7 | Can't tell how many requests breached the latency target | Used quantile metrics as SLIs | Counters/buckets of events inside vs outside the threshold |
| 8 | Took on an SLO you can't control (data freshness/correctness) | Scope creep from customers | "DO NOT agree to support SLOs that you are not responsible for" |
| 9 | Monitoring system overwhelmed by tens of thousands of series | Collecting debugging metrics continuously | Debugging metrics only need to be available, not collected |
| 10 | Alert can't distinguish "broker down" from "monitoring down" | Relied on stale metrics as the health check | Add an external port-connect health check |
| 11 | One broker's slow disk destroyed cluster-wide producer throughput | "One bad egg" — every producer talks to every broker, so one slow broker back-pressures all of them | Treat single-broker outliers as urgent; SMART + IPMI monitoring |
| 12 | Write latency tripled with nothing "broken" | RAID controller BBU failed → onboard cache silently disabled | Monitor the disk controller and BBU, "whether you are using hardware RAID or not" |
| 13 | Broker running on reduced memory/CPU with no failure | Soft hardware failures — bad memory segment or CPU bypassed by the OS | IPMI health, dmesg kernel ring buffer |
| 14 | Disk filled and the broker died abruptly | "brokers operate properly RIGHT UP UNTIL THE DISK IS FILLED, and then this disk WILL FAIL ABRUPTLY" | Alert on free space and inodes as capacity, not performance |
| 15 | Intermittent network problems | Bad cable/connector, speed/duplex mismatch, undersized OS buffers | "the number of ERRORS on the network interfaces — if the error count is increasing, there is probably an unaddressed issue" |
| 16 | One broker behaves differently from the rest | Config drift | Chef/Puppet-style configuration management |
| 17 | A monitoring agent starved the broker | Another process consuming CPU/memory — "a process that is SUPPOSED to be running, such as a monitoring agent, but is having problems" | top; don't colocate (Ch. 2) |
| 18 | Cluster imbalanced even though partitions are evenly spread | "Kafka does not detect issues such as HOT PARTITIONS", and leadership doesn't return after a restart | Preferred replica election first, then Cruise Control / kafka-assigner |
| 19 | Two brokers both claim to be controller | "a controller thread that should have exited has become STUCK" | Restart both; expect controlled shutdown to fail — force stop |
| 20 | No broker is the controller; topic creation hangs | e.g. network partition from ZooKeeper | Fix the root cause, then restart all brokers to reset controller state |
| 21 | Controller queue size grows and never drops | Controller stuck | Move the controller — but "there will often be problems performing a controlled shutdown of ANY broker"; Ch. 12's znode deletion is the alternative |
| 22 | Broker latency high; requests queueing | Request handler idle ratio < 10% | Threads = CPU count (incl. hyperthreads); then reduce load or add brokers |
| 23 | Request handler threads burning CPU unnecessarily | Old clients force decompress → validate → recompress, historically behind a synchronous lock | "ONE OF THE SINGLE LARGEST PERFORMANCE IMPROVEMENTS" — get all clients and brokers to message format 0.10+ |
| 24 | "Bytes out equals bytes in with no consumers" | Bytes-out includes replica fetch traffic | Expected at RF 2. Formula: in × ((RF−1) + consumer_groups) |
| 25 | Looking for a "messages out" metric | Doesn't exist — the broker never expands batches on the read path | Use fetch request rate |
| 26 | Alerted on MeanRate and never saw a spike | MeanRate is averaged since broker start | OneMinuteRate for spikes; 5/15-minute for balance |
| 27 | A broker leads 0 partitions after recovering | Leadership isn't reclaimed automatically | Alert on leader count / LeaderCount ÷ PartitionCount ≈ 1/RF |
| 28 | Offline partitions metric reads 0 everywhere | "ONLY provided by the broker that is the controller" | Aggregate with max, or scrape the controller specifically |
| 29 | Latency is high; no idea which phase | Only collected TotalTimeMs | Collect all seven timings; map to queue → local → remote → throttle → response |
| 30 | Fetch-latency alerts flap | Governed by fetch.min.bytes / fetch.max.wait.ms | Baseline Produce p99.9 instead |
| 31 | One partition is far bigger than its siblings | Hot key — uneven key distribution | Per-partition Size; then a custom partitioner (Ch. 3 §9.5) |
| 32 | LogEndOffset − LogStartOffset ≠ message count | Log compaction creates offset gaps | Don't compute counts from offsets |
| 33 | Broker hit "too many open files" | FDs for every log segment and every network connection; "a problem closing network connections properly could cause the broker to RAPIDLY EXHAUST" | Monitor OpenFileDescriptorCount vs MaxFileDescriptorCount |
| 34 | Broker dropped out of the cluster with no restart | Long GC pause → ZK session expiry (Ch. 6 §1) | GC CollectionTime/CollectionCount + LastGcInfo.duration |
| 35 | Misread system load as a percentage | Load average is a count of runnable + uninterruptible-sleep threads; 100% == number of CPUs | Divide by CPU count |
| 36 | Compacted topics grew without bound; tombstones never processed | "failure in compaction of a SINGLE partition can HALT the log compaction threads ENTIRELY, AND SILENTLY" — and no metric exposes it | Enable kafka.log.LogCleaner, kafka.log.Cleaner, kafka.log.LogCleanerManager at DEBUG by default |
| 37 | Broker filled its disk with its own logs | Request logger left at DEBUG/TRACE | Enable only while debugging |
| 38 | Producer silently dropping messages | Retries exhausted | record-error-rate should ALWAYS be zero — alert on it |
| 39 | Producer back-pressuring the application | Rising request-latency-avg | Baseline it and alert above |
| 40 | Can't tell which broker a producer is struggling with | Only looked at overall metrics | Per-broker request-latency-avg — the only stable per-broker metric |
| 41 | Consumer lag alert missed a stalled consumer | records-lag-max shows one partition and requires a working consumer | External lag monitoring — Burrow |
| 42 | Lag thresholds unmaintainable at scale | 100,000 partitions × per-partition thresholds | Burrow's threshold-free, progress-based status |
| 43 | False alerts on low consumer throughput | Alerting on records-consumed-rate minimums assumes the producer is healthy | Don't; Kafka deliberately decouples the two |
| 44 | Consumer group pauses repeatedly | Rebalance storms | sync-rate should be ~0; check sync-time-avg |
| 45 | Offset commits slow, consumer throughput drops | Commits are produce requests to one broker (Ch. 7 §5.2) | Baseline and alert on commit-latency-avg |
| 46 | Load uneven inside a consumer group | Assignor imbalance (RangeAssignor, Ch. 4 §6.5) | Compare assigned-partitions across instances |
| 47 | Client mysteriously slow; broker healthy; no errors | Quota throttling — the broker returns NO error code | produce-throttle-time-avg / fetch-throttle-time-avg — monitor them even before enabling quotas |
| 48 | Can't tell whether it's the client, the network, or Kafka | No end-to-end view | Xinfra Monitor — produce+consume across every broker, measuring availability and total latency |