Learn Labs
2. Installing Kafka

2.9 What actually breaks in production — Ch. 2 consolidated

Production failure catalog
0 rows
#FailureRoot causeFix / detection
1Broker won't start, logs an ID errorDuplicate broker.idUnique IDs derived from hostname
2One disk fills while others are emptylog.dirs placement uses fewest partitions, not least spaceMonitor per-mount usage, not just total; rebalance manually; prefer equal-size disks
3Broker restart takes hours after a crashnum.recovery.threads.per.data.dir = 1 default; unclean shutdown must check and truncate every segmentRaise it (remember: × number of log dirs)
4Topics appear from nowhere; typo'd topic names in prodauto.create.topics.enable=true, and metadata requests alone create topicsSet false; provision topics explicitly
5All leadership piles onto one broker after a restartNo automatic preferred-leader rebalanceauto.leader.rebalance.enable=true; tune imbalance thresholds
6"1 week retention" actually keeps 17 daysLow-volume topic never fills a 1 GB segment; retention only applies to closed segmentsLower log.segment.bytes and/or set log.roll.ms for low-volume topics
7Disk usage jumps after a partition reassignment and never fallsTime retention reads mtime; partition moves reset mtime → excess retentionExpect it; verify after rebalances; consider size-based retention for moved partitions
8Data deleted earlier than expectedBoth log.retention.bytes and log.retention.ms set — either triggers deletionPick one policy; the book explicitly recommends this "to prevent surprises and unwanted data loss"
9Retention silently doubles after adding partitionslog.retention.bytes is per partitionRecompute total retention whenever partition count changes
10I/O latency spikes at regular intervalsTime-based segment roll: the clock starts at broker start, so all low-volume partitions roll simultaneouslyStagger, or use size-based rolling where possible
11Consumer permanently stuck on one partitionfetch.message.max.bytes < broker message.max.bytesRaise consumer fetch size and replica.fetch.max.bytes before raising broker max
12Replication stalls on one partitionreplica.fetch.max.bytes < message.max.bytessame as above
13Producer got an ack, message is gonemin.insync.replicas=1 (default) + leader failed before replicating; leadership moved to a replica lacking the writemin.insync.replicas=2 with producer acks=all; RF ≥ 3 (Ch. 7)
14A rolling upgrade causes an outageRF only 1 above min.insync.replicas → planned outage consumes the entire redundancy budget, then a disk diesRF++ (2 above min.insync.replicas)
15Consumer performance degrades after colocating another servicePage cache contention — Kafka's read path is the page cacheDon't colocate significant applications with brokers
16Replication falls behind under peak load; under-replicated partitions climbSaturated NIC — outbound = consumers × inbound + replication + mirroring10 Gb NICs minimum; count replication as a consumer when sizing
17Whole-broker performance collapse, everything slowSwapping — pages evicted, page cache starvedvm.swappiness=1 (not 0 — semantics changed in kernel 3.5-rc1)
18Periodic multi-second I/O stallsvm.dirty_ratio too high → forced synchronous flushesTune with /proc/vmstat under real load; require replication if running high
19Kernel flushes constantly, throughput poorvm.dirty_background_ratio set to 0Use ~5; never 0
20"Too many open files"FD/map limits below partitions × (partition_size/segment_size) + connectionsvm.max_map_count 400k–600k; raise ulimits
21Filesystem corruption / data loss after a host crashExt4 delayed allocation + unsafe tuning (long commit interval)Prefer XFS; if Ext4, understand you accepted this risk
22Unexplained write amplificationatime updates on every readmount noatime (safe — Kafka doesn't use atime; mtime is preserved)
23Connections dropped under burst / poor large-transfer throughputDefault kernel socket + backlog sizesnet.core.*mem_*, net.ipv4.tcp_*mem, tcp_window_scaling=1, raise tcp_max_syn_backlog and netdev_max_backlog
24Long GC pauses → broker drops out of the ISR / ZK session expiresCMS default (pre-G1GC era compatibility choice)G1GC with MaxGCPauseMillis=20, InitiatingHeapOccupancyPercent=35; fixed heap -Xms == -Xmx
25A rack loses power and a partition goes offline despite RF 3Rack awareness applies to newly created partitions only; reassignments silently break it and nothing monitors itCruise Control or equivalent; audit replica placement after every reassignment
26You lose all broker data when a cloud VM movesEphemeral disksAzure Managed Disks; AWS: understand local-SSD vs EBS tradeoff
27Several brokers go offline simultaneously; controller acts strangely for hours afterwardShared ZooKeeper ensemble with a noisy application → ZK latency/timeout → brokers lose ZK together → offline partitions + controller stress + subtle errors long afterwardDedicated ensemble for Kafka; chroot per cluster; never colocate a busy app
28Cluster hits a wall at high partition counts; produce/consume/controller queues back upExceeded replica-per-broker limits≤ 14,000 replicas/broker, ≤ 1,000,000 replicas/cluster (replicas = partitions × RF)
29/tmp/kafka-logs disappears on rebootThe install example's default pathAlways set log.dirs to real, persistent, monitored mounts
30ZooKeeper maintenance causes an outage3-node ensemble tolerates 1 failure; a rolling node swap uses it all5-node ensemble; never exceed 7; add observers for read-only load