2. Installing Kafka
2.9 What actually breaks in production — Ch. 2 consolidated
| # | Failure | Root cause | Fix / detection |
|---|---|---|---|
| 1 | Broker won't start, logs an ID error | Duplicate broker.id | Unique IDs derived from hostname |
| 2 | One disk fills while others are empty | log.dirs placement uses fewest partitions, not least space | Monitor per-mount usage, not just total; rebalance manually; prefer equal-size disks |
| 3 | Broker restart takes hours after a crash | num.recovery.threads.per.data.dir = 1 default; unclean shutdown must check and truncate every segment | Raise it (remember: × number of log dirs) |
| 4 | Topics appear from nowhere; typo'd topic names in prod | auto.create.topics.enable=true, and metadata requests alone create topics | Set false; provision topics explicitly |
| 5 | All leadership piles onto one broker after a restart | No automatic preferred-leader rebalance | auto.leader.rebalance.enable=true; tune imbalance thresholds |
| 6 | "1 week retention" actually keeps 17 days | Low-volume topic never fills a 1 GB segment; retention only applies to closed segments | Lower log.segment.bytes and/or set log.roll.ms for low-volume topics |
| 7 | Disk usage jumps after a partition reassignment and never falls | Time retention reads mtime; partition moves reset mtime → excess retention | Expect it; verify after rebalances; consider size-based retention for moved partitions |
| 8 | Data deleted earlier than expected | Both log.retention.bytes and log.retention.ms set — either triggers deletion | Pick one policy; the book explicitly recommends this "to prevent surprises and unwanted data loss" |
| 9 | Retention silently doubles after adding partitions | log.retention.bytes is per partition | Recompute total retention whenever partition count changes |
| 10 | I/O latency spikes at regular intervals | Time-based segment roll: the clock starts at broker start, so all low-volume partitions roll simultaneously | Stagger, or use size-based rolling where possible |
| 11 | Consumer permanently stuck on one partition | fetch.message.max.bytes < broker message.max.bytes | Raise consumer fetch size and replica.fetch.max.bytes before raising broker max |
| 12 | Replication stalls on one partition | replica.fetch.max.bytes < message.max.bytes | same as above |
| 13 | Producer got an ack, message is gone | min.insync.replicas=1 (default) + leader failed before replicating; leadership moved to a replica lacking the write | min.insync.replicas=2 with producer acks=all; RF ≥ 3 (Ch. 7) |
| 14 | A rolling upgrade causes an outage | RF only 1 above min.insync.replicas → planned outage consumes the entire redundancy budget, then a disk dies | RF++ (2 above min.insync.replicas) |
| 15 | Consumer performance degrades after colocating another service | Page cache contention — Kafka's read path is the page cache | Don't colocate significant applications with brokers |
| 16 | Replication falls behind under peak load; under-replicated partitions climb | Saturated NIC — outbound = consumers × inbound + replication + mirroring | 10 Gb NICs minimum; count replication as a consumer when sizing |
| 17 | Whole-broker performance collapse, everything slow | Swapping — pages evicted, page cache starved | vm.swappiness=1 (not 0 — semantics changed in kernel 3.5-rc1) |
| 18 | Periodic multi-second I/O stalls | vm.dirty_ratio too high → forced synchronous flushes | Tune with /proc/vmstat under real load; require replication if running high |
| 19 | Kernel flushes constantly, throughput poor | vm.dirty_background_ratio set to 0 | Use ~5; never 0 |
| 20 | "Too many open files" | FD/map limits below partitions × (partition_size/segment_size) + connections | vm.max_map_count 400k–600k; raise ulimits |
| 21 | Filesystem corruption / data loss after a host crash | Ext4 delayed allocation + unsafe tuning (long commit interval) | Prefer XFS; if Ext4, understand you accepted this risk |
| 22 | Unexplained write amplification | atime updates on every read | mount noatime (safe — Kafka doesn't use atime; mtime is preserved) |
| 23 | Connections dropped under burst / poor large-transfer throughput | Default kernel socket + backlog sizes | net.core.*mem_*, net.ipv4.tcp_*mem, tcp_window_scaling=1, raise tcp_max_syn_backlog and netdev_max_backlog |
| 24 | Long GC pauses → broker drops out of the ISR / ZK session expires | CMS default (pre-G1GC era compatibility choice) | G1GC with MaxGCPauseMillis=20, InitiatingHeapOccupancyPercent=35; fixed heap -Xms == -Xmx |
| 25 | A rack loses power and a partition goes offline despite RF 3 | Rack awareness applies to newly created partitions only; reassignments silently break it and nothing monitors it | Cruise Control or equivalent; audit replica placement after every reassignment |
| 26 | You lose all broker data when a cloud VM moves | Ephemeral disks | Azure Managed Disks; AWS: understand local-SSD vs EBS tradeoff |
| 27 | Several brokers go offline simultaneously; controller acts strangely for hours afterward | Shared ZooKeeper ensemble with a noisy application → ZK latency/timeout → brokers lose ZK together → offline partitions + controller stress + subtle errors long afterward | Dedicated ensemble for Kafka; chroot per cluster; never colocate a busy app |
| 28 | Cluster hits a wall at high partition counts; produce/consume/controller queues back up | Exceeded replica-per-broker limits | ≤ 14,000 replicas/broker, ≤ 1,000,000 replicas/cluster (replicas = partitions × RF) |
| 29 | /tmp/kafka-logs disappears on reboot | The install example's default path | Always set log.dirs to real, persistent, monitored mounts |
| 30 | ZooKeeper maintenance causes an outage | 3-node ensemble tolerates 1 failure; a rolling node swap uses it all | 5-node ensemble; never exceed 7; add observers for read-only load |