Learn Labs
1. Meet Kafka

1.7 Deployment / monitoring / scaling / backup — what Ch. 1 establishes

These are covered properly in Ch. 2, 7, 10, 12, 13, but Ch. 1 sets the frame:

These are covered properly in Ch. 2, 7, 10, 12, 13, but Ch. 1 sets the frame:

Deployment shape

  • A JVM application — Java 8 or 11; Linux recommended.
  • Needs ZooKeeper for cluster metadata.
  • 1 broker (PoC) → 3 brokers (dev) → tens–hundreds (prod).
  • Expansion is ONLINE — no availability impact.

Scaling axes

AxisMechanism
Write/read throughput of a topicmore partitions (spread across more brokers)
Consumer processing throughputmore consumers in a group (capped at partition count)
Fault tolerancehigher replication factor
Cluster capacityadd brokers online
Geographic / isolationmore clusters + MirrorMaker

"Backup" in Kafka terms — reframe the question. Kafka is not backed up like a database, and Ch. 1 explains why the question changes shape:

  1. Replication is the intra-cluster durability mechanism (redundancy, not backup — it won't save you from a bad delete or a logic bug).
  2. Retention is a time-bounded replay buffer — real, but it expires.
  3. Log compaction is indefinite per-key state retention.
  4. MirrorMaker to another cluster/DC is the DR story.
  5. The genuinely durable archive is usually a sink: Connect → HDFS/S3, which is also where "replay from the beginning of time" lives.

Kafka's actual durability contract is "the log is the source of truth for a configured window, and replicated within the cluster during that window." Anything longer is a sink's job.

Monitoring, framed by design

  • Consumer lag is the master metric — it exists because consumers pull and offsets are tracked. Lag vs. retention is the data-loss early warning.
  • Under-replicated partitions — replication falling behind is the "you are one failure from data loss" signal. The chapter notes network saturation as a common cause: "should the network interface become saturated, it is not uncommon for cluster replication to fall behind, which can leave the cluster in a vulnerable state."
  • Offline partitions — leaderless partitions = unavailable data.
  • Controller count — exactly one controller should be active.