Learn Labs
2. Installing Kafka

2.10 Deploy / monitor / scale / back up — Chapter 2's contribution

Chapter 2 has no backup section, and that absence is informative.

Deploy checklist (production)

  • Linux; latest patch of JDK 8 or 11 (full JDK)
  • ZooKeeper: 5-node ensemble, odd count, ≤7, separate hosts, myid file per node, all 3 ports open between members
  • Dedicated ZK ensemble for Kafka; chroot path per Kafka cluster
  • broker.id derived from hostname; unique
  • listeners explicit; port ≥ 1024; never run Kafka as root
  • log.dirs on real persistent disks (Not /tmp), equal sizes, monitored per-mount
  • num.recovery.threads.per.data.dir raised (× number of dirs)
  • auto.create.topics.enable=false
  • delete.topic.enable considered
  • auto.leader.rebalance.enable=true
  • min.insync.replicas=2 + producers acks=all + RF 3 (RF++ = 4 if you can)
  • Retention: pick time Or size, not both; log.roll.ms for low-volume topics
  • message.max.bytes coordinated with consumer fetch + replica.fetch.max.bytes
  • broker.rack set to rack / cloud fault domain
  • Dual power (2 circuits), dual switches (bonded), brokers in separate racks
  • XFS, mounted noatime,largeio
  • sysctl: swappiness=1, dirty_background_ratio=5, dirty_ratio 60–80, max_map_count 400–600k, overcommit_memory=0, socket + TCP buffers, backlogs
  • G1GC with MaxGCPauseMillis=20, IHOP=35, -Xms == -Xmx
  • ≥10 Gb NIC
  • Rebalancing tool (Cruise Control) to preserve rack awareness over time

Monitoring implied by this chapter

SignalWhy Ch. 2 makes it matter
Per-log-dir disk usage (not aggregate)placement is by partition count, not bytes
Under-replicated partitionsthe symptom of NIC saturation and of GC/ZK trouble
Offline partitionswhat a ZK interruption produces
Active controller count (== 1)ZK stress destabilizes the controller
GC pause timelong pauses → ISR drops → ZK session expiry
Dirty pages (/proc/vmstat)to tune vm.dirty_* empirically under load
Swap usage (should be ~0)swapping starves the page cache
Open file descriptorssegments + connections grow with partitions
NIC utilizationoutbound = consumers × inbound + replication + mirroring
Replicas per broker / per clusterhard ceilings: 14k / 1M
ZooKeeper latency + outstanding requestsKafka is sensitive to ZK latency
Rack-awareness audit after reassignmentsKafka will not tell you it broke

Scaling levers from this chapter

You wantThe lever
more throughputmore partitions (+ brokers to host them)
more consumer parallelismmore consumers in a group (≤ partition count)
more retentionmore disk, or more brokers
more fault tolerancehigher RF (costs ≥100% storage per extra replica)
more ZooKeeper read capacityobserver nodes (NOT more voting members past 7)
faster recoverynum.recovery.threads.per.data.dir
lower produce latencyfaster disks (SSD), tuned dirty ratios
better consumer latencymore RAM for page cache; don’t colocate

"Backup" — what Ch. 2 actually gives you

Chapter 2 has no backup section, and that absence is informative. What it does give you:

  1. Replication factor — intra-cluster redundancy. Not a backup: it won't protect you from a bad --delete, which is exactly why delete.topic.enable=false exists as a config.
  2. Retention — a bounded replay window, and Ch. 2 shows how easily you can miscalculate it (segment rolling, mtime resets, dual policies). Your "backup window" is only as accurate as your understanding of segment mechanics.
  3. delete.topic.enable=false — the closest thing to a "protect me from myself" control.
  4. Managed disks over ephemeral — cloud-specific durability floor. "If a VM is moved, you run the risk of losing all the data on your Kafka broker."
  5. ZooKeeper's dataDir — genuinely worth backing up separately; it holds cluster metadata, and ZK has its own snapshot/txn-log durability story.

Real archival remains a sink concern (Connect → S3/HDFS, Ch. 9) and DR remains a cross-cluster concern (MirrorMaker, Ch. 10).


On this page