10. Cross-Cluster Data Mirroring
Chapter 10 of Kafka: The Definitive Guide — 12 sections.
Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 10 Terminology, established up front: "In most databases, continuously copying data between database servers is called replication. Since we've used replication to describe movement of data between Kafka nodes that are part of the same cluster, we'll call copying of data between Kafka clusters MIRRORING. Apache Kafka's built-in cross-cluster replicator is called MirrorMaker."
| Term | Scope | Mechanism |
|---|---|---|
| REPLICATION | within one cluster (Ch. 6/7) | synchronous-ish, ISR-based |
| MIRRORING | between clusters (this chapter) | ASYNCHRONOUS, consumer + producer |
The easy case, dismissed immediately: "In some cases, the clusters are completely separated... different departments, different use cases, different SLAs/workloads, different security requirements. Those use cases are fairly easy — managing multiple distinct clusters is the same as running a single cluster multiple times." This chapter is about the interdependent case.
Sections
- 10.1Five use cases for cross-cluster mirroring
- 10.23The realities of cross-datacenter communication
- 10.38Multicluster architectures
- 10.46Disaster recovery planningThat is a genuinely expensive operation to discover mid-incident.
- 10.54Stretch clusters — the synchronous option
- 10.65Apache Kafka's MirrorMaker① "Define aliases for the clusters used in replication flows."
- 10.74Deploying MirrorMaker in productionA canary is the only listed monitor that measures the thing you actually care about — end-to-end delivery — rather than a proxy for it.
- 10.83Tuning MirrorMaker
- 10.95Other cross-cluster mirroring solutionsThat last mechanism is elegant: it automatically trades throughput for availability exactly when needed, and trades back when the emergency ends — precisely the manual decision th…
- 10.10Failure catalogWhat actually breaks in production — Ch. 10 consolidated
- 10.111Decision guide
- 10.12Self-testSelf-test