10.9 Other cross-cluster mirroring solutions
That last mechanism is elegant: it automatically trades throughput for availability exactly when needed, and trades back when the emergency ends — precisely the manual decision th…
"MirrorMaker also has some limitations when used in practice. It is worthwhile to look at some of the alternatives and the ways they address MirrorMaker limitations and complexities."
9.1 Uber uReplicator
The problem Uber hit with legacy MirrorMaker:
MM1 used consumers in A SINGLE CONSUMER GROUP. Rebalances were triggered by:
- adding MirrorMaker THREADS
- adding MirrorMaker INSTANCES
- BOUNCING MirrorMaker instances
- even ADDING NEW TOPICS matching the inclusion regex
“With a very large number of topics and partitions, this can take a while. This is especially true when using OLD CONSUMERS like Uber did. In some cases, THIS CAUSED 5–10 MINUTES OF INACTIVITY, causing mirroring to fall behind and ACCUMULATE A LARGE BACKLOG of events to mirror, WHICH CAN TAKE A LONG TIME TO RECOVER FROM.”
Their attempted workaround made it worse: “To avoid rebalances when someone added a topic matching the filter, Uber decided to MAINTAIN A LIST OF EXACT TOPIC NAMES instead of using a regex. But this was HARD TO MAINTAIN as ALL MirrorMaker instances had to be RECONFIGURED AND BOUNCED to add a new topic. IF NOT DONE CORRECTLY, THIS COULD RESULT IN ENDLESS REBALANCES as the consumers won’t be able to agree on the topics they subscribe to.”
The solution: "Uber decided to use Apache Helix as a central (but highly available) controller to manage the topic list and the partitions assigned to each uReplicator instance. Administrators use a REST API to add new topics to the list in Helix... Uber replaced the Kafka consumers with a 'Helix consumer' — this consumer takes its partition assignment FROM THE APACHE HELIX CONTROLLER rather than as a result of an agreement between the consumers. As a result, the Helix consumer can AVOID REBALANCES and instead LISTEN TO CHANGES in the assigned partitions that arrive from Helix."
The verdict: "uReplicator's dependency on Apache Helix introduces A NEW COMPONENT TO LEARN AND MANAGE, adding complexity to any deployment. As we saw earlier, MirrorMaker 2.0 solves many of these scalability and fault-tolerance issues WITHOUT ANY EXTERNAL DEPENDENCIES."
(MM2's "allocate partitions without the consumer group protocol" is the same idea, internalized.)
9.2 LinkedIn Brooklin
"LinkedIn built a mirroring solution on top of its data streaming system called Brooklin. Brooklin is a distributed service that can stream data between different HETEROGENEOUS data source and target systems, including Kafka."
Three use cases:
- “Data bridge to feed data into stream processing systems from different data sources”
- “Stream CHANGE DATA CAPTURE (CDC) events from different data stores”
- “Cross-cluster MIRRORING solution for Kafka”
"designed for high reliability and has been tested with Kafka at scale. It is used to mirror TRILLIONS OF MESSAGES A DAY and has been optimized for stability, performance, and operability. Brooklin comes with a REST API for management operations. It is a SHARED SERVICE that can process a large number of data pipelines, enabling the same service to mirror data across multiple Kafka clusters."
9.3 Confluent's three solutions
"At the same time that Uber developed its uReplicator, Confluent independently developed Confluent Replicator. Despite the similarities in names, the projects have ALMOST NOTHING IN COMMON — they are different solutions to two different sets of MirrorMaker problems."
Confluent Replicator
| MirrorMaker 2.0 | Confluent Replicator | |
|---|---|---|
| Basis | Kafka Connect | Kafka Connect "can run on existing Connect clusters" |
| Data replication + topologies | ✅ | ✅ |
| Consumer offset migration | ✅ | ✅ |
| Topic config migration | ✅ | ✅ |
| ACL migration | ✅ | ❌ "Replicator doesn't migrate ACLs" |
| Offset translation | ✅ "for any client" | ⚠️ "(using timestamp interceptor) only for Java clients" |
| Local/remote topic concept | ✅ | ❌ "but it supports aggregate topics" |
| Cycle prevention | topic-name prefixing | "using provenance headers" |
| Monitoring | Connect + MM metrics | "replication lag... REST API or Control Center UI" |
| Schema migration | ❌ | ✅ "supports schema migration between clusters and can perform schema translation" |
Multi-Region Clusters (MRC) — a smarter stretch cluster
THE KEY NEW CONCEPT — OBSERVERS: “asynchronous replicas that DO NOT JOIN THE ISR and hence HAVE NO IMPACT ON PRODUCERS USING acks=all, BUT ARE ABLE TO DELIVER RECORDS TO CONSUMERS.”
- ► “SYNCHRONOUS replication WITHIN a region and ASYNCHRONOUS replication BETWEEN regions to benefit from BOTH LOW LATENCY AND HIGH DURABILITY AT THE SAME TIME.”
- MRC is “also suitable only for datacenters within a 50 ms latency, but it uses A COMBINATION OF SYNCHRONOUS AND ASYNCHRONOUS REPLICATION to LIMIT IMPACT ON PRODUCER PERFORMANCE and provide HIGHER NETWORK TOLERANCE” — where plain stretch clusters “require datacenters to be CLOSE to each other and provide a STABLE LOW-LATENCY network to enable SYNCHRONOUS replication.”
- PLUS: “Replica placement constraints… allow you to specify A MINIMUM NUMBER OF REPLICAS PER REGION using rack IDs.”
- 💡 AUTOMATIC OBSERVER PROMOTION (Confluent Platform 6.1): “When
min.insync.replicasfalls below a configured minimum, OBSERVERS THAT HAVE CAUGHT UP ARE AUTOMATICALLY PROMOTED to allow them to JOIN ISRs, bringing the number of ISRs back up to the required minimum. The promoted observers use synchronous replication and MAY IMPACT THROUGHPUT, BUT THE CLUSTER REMAINS OPERATIONAL THROUGHOUT WITHOUT DATA LOSS EVEN IF A REGION FAILS. When the failed region recovers, observers are AUTOMATICALLY DEMOTED, getting the cluster back to normal performance levels.”
That last mechanism is elegant: it automatically trades throughput for availability exactly when needed, and trades back when the emergency ends — precisely the manual decision that unclean.leader.election.enable forces on you in open-source Kafka (Ch. 7 §3.2).
Cluster Linking — offset-preserving mirroring
"introduced as a preview feature in Confluent Platform 6.0, builds inter-cluster replication DIRECTLY INTO the Confluent Server. By using THE SAME PROTOCOL AS INTER-BROKER REPLICATION within a cluster, Cluster Linking performs OFFSET-PRESERVING REPLICATION across clusters, enabling SEAMLESS MIGRATION OF CLIENTS WITHOUT ANY NEED FOR OFFSET TRANSLATION."
- Kept synchronized: topic configuration, partitions, consumer offsets, ACLs ► “to enable failover with LOW RTO”.
- ⚠ “MIRROR TOPICS ARE MARKED AS READ-ONLY in the destination TO PREVENT ANY LOCAL PRODUCE, ensuring that mirror topics are LOGICALLY IDENTICAL to their source topic.” ► THIS is what makes offset preservation possible: no local writes can shift the offsets.
- ADVANTAGES: “operational simplicity WITHOUT THE NEED FOR SEPARATE CLUSTERS like Connect clusters”; “MORE PERFORMANT than external tools since IT AVOIDS DECOMPRESSION AND RECOMPRESSION during mirroring” — the §8.2 CPU cost, eliminated.
- LIMITATIONS: “Unlike MRC, THERE IS NO OPTION FOR SYNCHRONOUS REPLICATION”; “CLIENT FAILOVER IS A MANUAL PROCESS THAT REQUIRES CLIENT RESTART.”
- BEST FOR: “DISTANT datacenters with UNRELIABLE HIGH-LATENCY networks”; “reduces cross-datacenter traffic by REPLICATING ONLY ONCE”; “suitable for CLUSTER MIGRATION and TOPIC SHARING use cases.”
Solution selection matrix
| Requirement | Solution |
|---|---|
| RPO = 0 (synchronous, legal) | Stretch cluster (3 DCs) or MRC |
| DCs < 50 ms apart, want both low latency AND high durability | MRC (observers) |
| DCs far apart, unreliable network, want offset preservation | Cluster Linking |
| Open source, any topology | MirrorMaker 2.0 |
| Need ACL migration | MirrorMaker 2.0 (not Replicator) |
| Need schema migration/translation | Confluent Replicator |
| Offset translation for non-Java | MirrorMaker 2.0 |
| Heterogeneous sources + Kafka mirror | Brooklin |