1. Meet Kafka
1.7 Deployment / monitoring / scaling / backup — what Ch. 1 establishes
These are covered properly in Ch. 2, 7, 10, 12, 13, but Ch. 1 sets the frame:
These are covered properly in Ch. 2, 7, 10, 12, 13, but Ch. 1 sets the frame:
Deployment shape
- A JVM application — Java 8 or 11; Linux recommended.
- Needs ZooKeeper for cluster metadata.
- 1 broker (PoC) → 3 brokers (dev) → tens–hundreds (prod).
- Expansion is ONLINE — no availability impact.
Scaling axes
| Axis | Mechanism |
|---|---|
| Write/read throughput of a topic | more partitions (spread across more brokers) |
| Consumer processing throughput | more consumers in a group (capped at partition count) |
| Fault tolerance | higher replication factor |
| Cluster capacity | add brokers online |
| Geographic / isolation | more clusters + MirrorMaker |
"Backup" in Kafka terms — reframe the question. Kafka is not backed up like a database, and Ch. 1 explains why the question changes shape:
- Replication is the intra-cluster durability mechanism (redundancy, not backup — it won't save you from a bad delete or a logic bug).
- Retention is a time-bounded replay buffer — real, but it expires.
- Log compaction is indefinite per-key state retention.
- MirrorMaker to another cluster/DC is the DR story.
- The genuinely durable archive is usually a sink: Connect → HDFS/S3, which is also where "replay from the beginning of time" lives.
Kafka's actual durability contract is "the log is the source of truth for a configured window, and replicated within the cluster during that window." Anything longer is a sink's job.
Monitoring, framed by design
- Consumer lag is the master metric — it exists because consumers pull and offsets are tracked. Lag vs. retention is the data-loss early warning.
- Under-replicated partitions — replication falling behind is the "you are one failure from data loss" signal. The chapter notes network saturation as a common cause: "should the network interface become saturated, it is not uncommon for cluster replication to fall behind, which can leave the cluster in a vulnerable state."
- Offline partitions — leaderless partitions = unavailable data.
- Controller count — exactly one controller should be active.