Learn Labs
10. Cross-Cluster Data Mirroring

10.1 Five use cases for cross-cluster mirroring

① Regional and central clusters

"the classic example is a company that modifies prices based on supply and demand. This company can have a datacenter in each city in which it has a presence, collects information about local supply and demand, and adjusts prices accordingly. All this information will then be mirrored to a central cluster where business analysts can run company-wide reports on its revenue."

② High availability (HA) and disaster recovery (DR)

"The applications run on just one Kafka cluster and don't need data from other locations, but you are concerned about the possibility of the ENTIRE CLUSTER becoming unavailable. For redundancy, you'd like a second Kafka cluster with all the data... so in case of emergency you can direct your applications to the second cluster."

③ Regulatory compliance

"Companies operating in different countries may need different configurations and policies to conform to legal and regulatory requirements in each country. For instance, some datasets may be stored in separate clusters with STRICT ACCESS CONTROL, with SUBSETS of data replicated to other clusters with WIDER ACCESS. To comply with regulatory policies that govern retention period in each region, datasets may be stored in clusters in different regions with different configurations."

④ Cloud migrations

"if a new application is deployed in the cloud but requires some data that is updated by applications running on premises and stored in an on-premises database, you can use Kafka Connect to capture database changes to the LOCAL Kafka cluster and then MIRROR these changes to the CLOUD Kafka cluster where the new application can use them. This helps CONTROL THE COSTS of cross-datacenter traffic as well as improve GOVERNANCE AND SECURITY of the traffic."

The pattern: Connect for the source hop, MirrorMaker for the WAN hop. Never let N cloud applications each pull across the WAN.

⑤ Aggregation of data from edge clusters

*"Several industries, including retail, telecommunications, transportation, and healthcare, generate data from small devices with limited connectivity. An aggregate cluster with high availability can support analytics and other use cases for data from a large number of edge clusters.

This REDUCES connectivity, availability, and durability requirements on low-footprint edge clusters, for example, in IoT use cases. A highly available aggregate cluster provides business continuity EVEN WHEN EDGE CLUSTERS ARE OFFLINE and simplifies the development of applications that don't have to directly deal with a large number of edge clusters with unstable networks."*


On this page