Learn Labs
6. Kafka Internals

6.8 What actually breaks in production — Ch. 6 consolidated

Production failure catalog
0 rows
#SymptomInternals-level causeFix
1Brokers drop out of the cluster with no restartA long GC pause or network partition looks identical to "stopped" — the ZK ephemeral node vanishesG1GC tuning (Ch. 2); raise zookeeper.session.timeout.ms; dedicated ZK ensemble
2New broker won't start: node already existsDuplicate broker.id — the ephemeral node is already registeredUnique IDs
3Replaced a dead broker; it doesn't pick up its old partitionsYou gave it a new broker ID; replica lists still reference the old oneReuse the dead broker's ID — assignments are inherited instantly
4Two brokers behaved as controller; contradictory commandsOld controller resumed after a GC pause as a zombieNothing to fix — controller epoch fencing already discards its messages. Understand it so you don't misdiagnose
5Controller failover takes many seconds; cluster-wide stallController must read the full replica state map from ZooKeeper before it can act; "in clusters with large numbers of partitions... several seconds"Limit partition counts (Ch. 2's 14k/broker, 1M/cluster); ultimately, KRaft
6Metadata inconsistent between brokers, controller, and ZooKeeperZK writes sync, broker pushes async, ZK reads async — "edge cases... challenging to detect"KRaft; until then, avoid direct ZooKeeper access (Ch. 5)
7Client produced to a broker that was no longer the leaderThe broker was too out-of-date to know it lost leadershipKRaft's fenced state; today, rely on NotLeaderForPartition + metadata refresh
8Consumer latency higher than expected after enabling follower fetchThe high-water mark propagates to followers with a delay — followers are always slightly behindExpected trade-off: cheaper cross-AZ reads for slightly staler data
9A partition becomes unavailable on leader failure despite RF 3Followers were out of sync (>replica.lag.time.max.ms) so were ineligible for electionFix replication throughput (NIC, Ch. 2); monitor under-replicated partitions; RF++
10One broker holds most leadership; it's hotThe first replica in the list is the preferred leader; a manual reassignment put the same broker first everywhere"make sure you spread those around different brokers"; auto.leader.rebalance.enable; Cruise Control
11Producers get NotLeaderForPartition in burstsLeader election happened; the client's cached metadata is staleNormal and retriable — the client refreshes and retries; check metadata.max.age.ms if persistent
12Acknowledged data lost after a correlated power failureKafka writes to the filesystem cache and does NOT fsync — durability comes from replication, not diskRF ≥ 3 across racks/AZs so failures aren't correlated; min.insync.replicas=2
13acks=all produce latency spikesRequest sits in purgatory until followers replicateFix replication speed; watch purgatory size + under-replicated partitions
14Consumers see an empty response even though the leader has dataThe data is above the high-water mark — not yet on all ISR, therefore "unsafe"Expected. If persistent, replication is slow — see #9
15New messages take unusually long to reach consumersVisibility waits for ISR replication; bounded by replica.lag.time.max.msSame root cause as #9
16Broker CPU high and zero-copy seemingly not workingMessage format down-conversion for old consumersCheck FetchMessageConversionsPerSec / MessageConversionsTimeMs (KIP-188); upgrade clients
17Clients break after a Kafka upgradeYou upgraded clients before brokers; old brokers can't parse newer request versionsAlways upgrade brokers first — "new brokers know how to handle old requests, but not vice versa"
18Fetch overhead high on consumers with many partitionsFetch session not created or evicted (limited cache space; followers and large-partition consumers are prioritized) → fell back to full fetch requestsExpected degradation; reduce partitions per consumer, or accept it. Monitor
19A partition can't grow past a certain size"Partitions cannot be split between multiple brokers, and not even between multiple disks" — bounded by one mount pointMore partitions; bigger mounts; eventually tiered storage
20One disk fills while others sit emptyDirectory placement counts partitions, not bytes — and "if you add a new disk, ALL new partitions will be created on that disk"Monitor per-mount usage; equal-size disks; manual reassignment
21Brokers with more disk space get no more dataBroker allocation ignores available space and existing load entirelyDon't mix heterogeneous hardware casually; use a balancer
22A backfill/historical read destroys everyone's latencyOld reads evict the hot page cache and compete for disk I/O (21 ms → 60 ms p99 in KIP-405's measurement)Tiered storage (network path, leaves page cache intact); until then, isolate historical consumers
23"7-day retention" keeps far moreThe active segment is never deleted, and only closed segments are eligibleSize segments so retention ÷ roll ≈ several segments; log.roll.ms for low-volume topics
24"Too many open files"Broker keeps an open handle to every segment of every partition, including inactive onesRaise ulimits / vm.max_map_count; don't over-shrink segments
25Rebalancing/expanding the cluster is glacially slowMove time is driven by partition size — "large partitions make the cluster less elastic"Smaller partitions; tiered storage; throttle + plan
26Compaction stops working; error in the logsNot even one full segment fits in the per-thread offset map (total memory ÷ thread count)Allocate more offset-map memory or use FEWER cleaner threads
27Compacted topic keeps growingCompaction only touches inactive segments, and only fires at ~50% dirtyTune the dirty ratio; check cleaner threads are alive; consider delete.and.compact
28Compaction fails outrightThe topic contains null keysCompaction requires a key on every record
29GDPR deletion didn't reach the downstream databaseThe consumer was offline while the tombstone came and went; on restart the key just doesn't exist, so no delete event is ever seenTombstone retention > worst-case consumer downtime; alert on long consumer outages
30A delete request wasn't honored within the legal windowNo max.compaction.lag.ms — the tombstone sat in a segment that wasn't eligibleSet max.compaction.lag.ms below your legal deadline (e.g. GDPR 30 days)
31Records "deleted" via deleteRecords but consumers weren't notifieddeleteRecords moves the low-water mark — it emits no eventUse tombstones when downstream systems must learn about deletions
32Suspected index corruption / weird offset lookupsIndex files have no checksumsDelete the index segments — they regenerate automatically (cost: recovery time)
33Poor compression ratio and high per-message overheadBatches of one; per-batch header amortized over a single record; deltas uselesslinger.ms > 0; write to fewer partitions per producer (sticky partitioner)