Learn Labs
13. Monitoring Kafka

13.10 What actually breaks in production — Ch. 13 consolidated

Production failure catalog
0 rows
#SymptomRoot causeFix
1Kafka is down and no alert firedThe monitoring pipeline runs on KafkaSeparate monitoring system for Kafka, or cross-DC metric shipping
2A JMX port became an RCE vectorRemote JMX enabled without security — "JMX not only allows a view into the state, it also allows CODE EXECUTION"In-process agent (Jolokia/MX4J); keep remote JMX disabled
3All broker metrics look fine but clients can't connectBroker metrics are subjectiveSynthetic clients (Xinfra Monitor) and client-side metrics
4Team ignores alertsAlert fatigue from too many metric-based alerts with hand-tuned thresholdsSLO/burn-rate alerting; "Check Engine light" principle
5URP alerts fire constantly for benign reasons"the URP metric can frequently be NONZERO for benign reasons"The book retracts its own prior advice — don't alert on URP; use it for diagnosis
6SLO breached before anyone noticedAlerted on the SLO itself, which is a lagging, week-scale indicatorBurn-rate alerting (0.1% → ticket at 0.4% → page at 2%)
7Can't tell how many requests breached the latency targetUsed quantile metrics as SLIsCounters/buckets of events inside vs outside the threshold
8Took on an SLO you can't control (data freshness/correctness)Scope creep from customers"DO NOT agree to support SLOs that you are not responsible for"
9Monitoring system overwhelmed by tens of thousands of seriesCollecting debugging metrics continuouslyDebugging metrics only need to be available, not collected
10Alert can't distinguish "broker down" from "monitoring down"Relied on stale metrics as the health checkAdd an external port-connect health check
11One broker's slow disk destroyed cluster-wide producer throughput"One bad egg" — every producer talks to every broker, so one slow broker back-pressures all of themTreat single-broker outliers as urgent; SMART + IPMI monitoring
12Write latency tripled with nothing "broken"RAID controller BBU failed → onboard cache silently disabledMonitor the disk controller and BBU, "whether you are using hardware RAID or not"
13Broker running on reduced memory/CPU with no failureSoft hardware failures — bad memory segment or CPU bypassed by the OSIPMI health, dmesg kernel ring buffer
14Disk filled and the broker died abruptly"brokers operate properly RIGHT UP UNTIL THE DISK IS FILLED, and then this disk WILL FAIL ABRUPTLY"Alert on free space and inodes as capacity, not performance
15Intermittent network problemsBad cable/connector, speed/duplex mismatch, undersized OS buffers"the number of ERRORS on the network interfaces — if the error count is increasing, there is probably an unaddressed issue"
16One broker behaves differently from the restConfig driftChef/Puppet-style configuration management
17A monitoring agent starved the brokerAnother process consuming CPU/memory — "a process that is SUPPOSED to be running, such as a monitoring agent, but is having problems"top; don't colocate (Ch. 2)
18Cluster imbalanced even though partitions are evenly spread"Kafka does not detect issues such as HOT PARTITIONS", and leadership doesn't return after a restartPreferred replica election first, then Cruise Control / kafka-assigner
19Two brokers both claim to be controller"a controller thread that should have exited has become STUCK"Restart both; expect controlled shutdown to fail — force stop
20No broker is the controller; topic creation hangse.g. network partition from ZooKeeperFix the root cause, then restart all brokers to reset controller state
21Controller queue size grows and never dropsController stuckMove the controller — but "there will often be problems performing a controlled shutdown of ANY broker"; Ch. 12's znode deletion is the alternative
22Broker latency high; requests queueingRequest handler idle ratio < 10%Threads = CPU count (incl. hyperthreads); then reduce load or add brokers
23Request handler threads burning CPU unnecessarilyOld clients force decompress → validate → recompress, historically behind a synchronous lock"ONE OF THE SINGLE LARGEST PERFORMANCE IMPROVEMENTS" — get all clients and brokers to message format 0.10+
24"Bytes out equals bytes in with no consumers"Bytes-out includes replica fetch trafficExpected at RF 2. Formula: in × ((RF−1) + consumer_groups)
25Looking for a "messages out" metricDoesn't exist — the broker never expands batches on the read pathUse fetch request rate
26Alerted on MeanRate and never saw a spikeMeanRate is averaged since broker startOneMinuteRate for spikes; 5/15-minute for balance
27A broker leads 0 partitions after recoveringLeadership isn't reclaimed automaticallyAlert on leader count / LeaderCount ÷ PartitionCount ≈ 1/RF
28Offline partitions metric reads 0 everywhere"ONLY provided by the broker that is the controller"Aggregate with max, or scrape the controller specifically
29Latency is high; no idea which phaseOnly collected TotalTimeMsCollect all seven timings; map to queue → local → remote → throttle → response
30Fetch-latency alerts flapGoverned by fetch.min.bytes / fetch.max.wait.msBaseline Produce p99.9 instead
31One partition is far bigger than its siblingsHot key — uneven key distributionPer-partition Size; then a custom partitioner (Ch. 3 §9.5)
32LogEndOffset − LogStartOffset ≠ message countLog compaction creates offset gapsDon't compute counts from offsets
33Broker hit "too many open files"FDs for every log segment and every network connection; "a problem closing network connections properly could cause the broker to RAPIDLY EXHAUST"Monitor OpenFileDescriptorCount vs MaxFileDescriptorCount
34Broker dropped out of the cluster with no restartLong GC pause → ZK session expiry (Ch. 6 §1)GC CollectionTime/CollectionCount + LastGcInfo.duration
35Misread system load as a percentageLoad average is a count of runnable + uninterruptible-sleep threads; 100% == number of CPUsDivide by CPU count
36Compacted topics grew without bound; tombstones never processed"failure in compaction of a SINGLE partition can HALT the log compaction threads ENTIRELY, AND SILENTLY" — and no metric exposes itEnable kafka.log.LogCleaner, kafka.log.Cleaner, kafka.log.LogCleanerManager at DEBUG by default
37Broker filled its disk with its own logsRequest logger left at DEBUG/TRACEEnable only while debugging
38Producer silently dropping messagesRetries exhaustedrecord-error-rate should ALWAYS be zero — alert on it
39Producer back-pressuring the applicationRising request-latency-avgBaseline it and alert above
40Can't tell which broker a producer is struggling withOnly looked at overall metricsPer-broker request-latency-avg — the only stable per-broker metric
41Consumer lag alert missed a stalled consumerrecords-lag-max shows one partition and requires a working consumerExternal lag monitoring — Burrow
42Lag thresholds unmaintainable at scale100,000 partitions × per-partition thresholdsBurrow's threshold-free, progress-based status
43False alerts on low consumer throughputAlerting on records-consumed-rate minimums assumes the producer is healthyDon't; Kafka deliberately decouples the two
44Consumer group pauses repeatedlyRebalance stormssync-rate should be ~0; check sync-time-avg
45Offset commits slow, consumer throughput dropsCommits are produce requests to one broker (Ch. 7 §5.2)Baseline and alert on commit-latency-avg
46Load uneven inside a consumer groupAssignor imbalance (RangeAssignor, Ch. 4 §6.5)Compare assigned-partitions across instances
47Client mysteriously slow; broker healthy; no errorsQuota throttling — the broker returns NO error codeproduce-throttle-time-avg / fetch-throttle-time-avg — monitor them even before enabling quotas
48Can't tell whether it's the client, the network, or KafkaNo end-to-end viewXinfra Monitor — produce+consume across every broker, measuring availability and total latency