6. Kafka Internals
6.8 What actually breaks in production — Ch. 6 consolidated
| # | Symptom | Internals-level cause | Fix |
|---|---|---|---|
| 1 | Brokers drop out of the cluster with no restart | A long GC pause or network partition looks identical to "stopped" — the ZK ephemeral node vanishes | G1GC tuning (Ch. 2); raise zookeeper.session.timeout.ms; dedicated ZK ensemble |
| 2 | New broker won't start: node already exists | Duplicate broker.id — the ephemeral node is already registered | Unique IDs |
| 3 | Replaced a dead broker; it doesn't pick up its old partitions | You gave it a new broker ID; replica lists still reference the old one | Reuse the dead broker's ID — assignments are inherited instantly |
| 4 | Two brokers behaved as controller; contradictory commands | Old controller resumed after a GC pause as a zombie | Nothing to fix — controller epoch fencing already discards its messages. Understand it so you don't misdiagnose |
| 5 | Controller failover takes many seconds; cluster-wide stall | Controller must read the full replica state map from ZooKeeper before it can act; "in clusters with large numbers of partitions... several seconds" | Limit partition counts (Ch. 2's 14k/broker, 1M/cluster); ultimately, KRaft |
| 6 | Metadata inconsistent between brokers, controller, and ZooKeeper | ZK writes sync, broker pushes async, ZK reads async — "edge cases... challenging to detect" | KRaft; until then, avoid direct ZooKeeper access (Ch. 5) |
| 7 | Client produced to a broker that was no longer the leader | The broker was too out-of-date to know it lost leadership | KRaft's fenced state; today, rely on NotLeaderForPartition + metadata refresh |
| 8 | Consumer latency higher than expected after enabling follower fetch | The high-water mark propagates to followers with a delay — followers are always slightly behind | Expected trade-off: cheaper cross-AZ reads for slightly staler data |
| 9 | A partition becomes unavailable on leader failure despite RF 3 | Followers were out of sync (>replica.lag.time.max.ms) so were ineligible for election | Fix replication throughput (NIC, Ch. 2); monitor under-replicated partitions; RF++ |
| 10 | One broker holds most leadership; it's hot | The first replica in the list is the preferred leader; a manual reassignment put the same broker first everywhere | "make sure you spread those around different brokers"; auto.leader.rebalance.enable; Cruise Control |
| 11 | Producers get NotLeaderForPartition in bursts | Leader election happened; the client's cached metadata is stale | Normal and retriable — the client refreshes and retries; check metadata.max.age.ms if persistent |
| 12 | Acknowledged data lost after a correlated power failure | Kafka writes to the filesystem cache and does NOT fsync — durability comes from replication, not disk | RF ≥ 3 across racks/AZs so failures aren't correlated; min.insync.replicas=2 |
| 13 | acks=all produce latency spikes | Request sits in purgatory until followers replicate | Fix replication speed; watch purgatory size + under-replicated partitions |
| 14 | Consumers see an empty response even though the leader has data | The data is above the high-water mark — not yet on all ISR, therefore "unsafe" | Expected. If persistent, replication is slow — see #9 |
| 15 | New messages take unusually long to reach consumers | Visibility waits for ISR replication; bounded by replica.lag.time.max.ms | Same root cause as #9 |
| 16 | Broker CPU high and zero-copy seemingly not working | Message format down-conversion for old consumers | Check FetchMessageConversionsPerSec / MessageConversionsTimeMs (KIP-188); upgrade clients |
| 17 | Clients break after a Kafka upgrade | You upgraded clients before brokers; old brokers can't parse newer request versions | Always upgrade brokers first — "new brokers know how to handle old requests, but not vice versa" |
| 18 | Fetch overhead high on consumers with many partitions | Fetch session not created or evicted (limited cache space; followers and large-partition consumers are prioritized) → fell back to full fetch requests | Expected degradation; reduce partitions per consumer, or accept it. Monitor |
| 19 | A partition can't grow past a certain size | "Partitions cannot be split between multiple brokers, and not even between multiple disks" — bounded by one mount point | More partitions; bigger mounts; eventually tiered storage |
| 20 | One disk fills while others sit empty | Directory placement counts partitions, not bytes — and "if you add a new disk, ALL new partitions will be created on that disk" | Monitor per-mount usage; equal-size disks; manual reassignment |
| 21 | Brokers with more disk space get no more data | Broker allocation ignores available space and existing load entirely | Don't mix heterogeneous hardware casually; use a balancer |
| 22 | A backfill/historical read destroys everyone's latency | Old reads evict the hot page cache and compete for disk I/O (21 ms → 60 ms p99 in KIP-405's measurement) | Tiered storage (network path, leaves page cache intact); until then, isolate historical consumers |
| 23 | "7-day retention" keeps far more | The active segment is never deleted, and only closed segments are eligible | Size segments so retention ÷ roll ≈ several segments; log.roll.ms for low-volume topics |
| 24 | "Too many open files" | Broker keeps an open handle to every segment of every partition, including inactive ones | Raise ulimits / vm.max_map_count; don't over-shrink segments |
| 25 | Rebalancing/expanding the cluster is glacially slow | Move time is driven by partition size — "large partitions make the cluster less elastic" | Smaller partitions; tiered storage; throttle + plan |
| 26 | Compaction stops working; error in the logs | Not even one full segment fits in the per-thread offset map (total memory ÷ thread count) | Allocate more offset-map memory or use FEWER cleaner threads |
| 27 | Compacted topic keeps growing | Compaction only touches inactive segments, and only fires at ~50% dirty | Tune the dirty ratio; check cleaner threads are alive; consider delete.and.compact |
| 28 | Compaction fails outright | The topic contains null keys | Compaction requires a key on every record |
| 29 | GDPR deletion didn't reach the downstream database | The consumer was offline while the tombstone came and went; on restart the key just doesn't exist, so no delete event is ever seen | Tombstone retention > worst-case consumer downtime; alert on long consumer outages |
| 30 | A delete request wasn't honored within the legal window | No max.compaction.lag.ms — the tombstone sat in a segment that wasn't eligible | Set max.compaction.lag.ms below your legal deadline (e.g. GDPR 30 days) |
| 31 | Records "deleted" via deleteRecords but consumers weren't notified | deleteRecords moves the low-water mark — it emits no event | Use tombstones when downstream systems must learn about deletions |
| 32 | Suspected index corruption / weird offset lookups | Index files have no checksums | Delete the index segments — they regenerate automatically (cost: recovery time) |
| 33 | Poor compression ratio and high per-message overhead | Batches of one; per-batch header amortized over a single record; deltas useless | linger.ms > 0; write to fewer partitions per producer (sticky partitioner) |