Learn Labs
10. Cross-Cluster Data Mirroring

10.10 What actually breaks in production — Ch. 10 consolidated

Production failure catalog
0 rows
#SymptomRoot causeFix
1Brokers spread across DCs behave badly; timeouts everywhereKafka was "designed, developed, tested, and tuned, all within a single datacenter" — defaults assume LANDon't do it, except as a deliberate stretch cluster with 3 DCs
2Data lost during a WAN partitionMirrorMaker was producing remotely; consumed events couldn't be deliveredRun MirrorMaker at the target DC — remote consuming fails safe
3Enormous cross-DC bandwidth billMultiple applications each consuming the same data across the WANOne cluster per DC; mirror once, consume locally
4A user visits another city's branch and their data isn't thereHub-and-spoke gives no cross-regional data accessOnly use it for data that "can be completely separated between regional datacenters"; otherwise active-active
5Events mirrored back and forth foreverActive-active without cycle preventionPer-DC topic namespaces (SF.users / NYC.users); MM2 does this by default via alias prefixing
6User writes to one DC, reads from another, and doesn't see their own writeAsynchronous multi-master"Stick" each user to a datacenter
7Two conflicting orders for the same user; downstream state divergesConcurrent writes in two DCs"You WILL have conflicts" — define consistent resolution rules both DCs will compute identically
8N² mirroring processes to manageActive-active needs a flow per pair per directionTools that share processes per destination cluster; one shared config file
9Custom origin-DC headers break mirroring"None of the existing mirroring tools will support your specific header format"Use the tool's own convention, or accept the extra work
10DR cluster too small to carry production during a real disasterCost-driven undersizing"A risky decision because you can't be sure it will hold up"
11Failover loses ~5,000 messagesAll mirroring is asynchronous: 1M msg/s × 5 ms = 5,000 messagesExpected. Planned failover avoids it (stop primary, let mirroring drain). Monitor DR lag continuously
12A line item arrives with no corresponding sale"Mirroring solutions currently don't support transactions" — related events across topics arrive independentlyApplications must tolerate it (Ch. 8 §3.6)
13Consumers fail on the DR cluster: offset doesn't existMirrored __consumer_offsets but offsets diverge (retention skew, producer retries)Time-based failover, or offset translation (MM2)
14DR consumer finds a committed offset with no matching recordOffset commits and records mirror independently and raceDecide in advance: beginning or end?
15Failover happened but nobody can explain what was reprocessedOffset-based failover is inherently unexplainableTime-based failover — "'We failed back to 4:03 a.m.' sounds better"
16Offset reset tool silently did nothing / conflictedConsumer group still running"The group should be STOPPED while running this type of tool and started immediately after"
17After failing back, the two clusters are permanently inconsistentThe old primary retained events the DR never got — reverse mirroring leaves phantom history"First SCRAPE the original cluster — delete all data and committed offsets"
18Applications can't find the DR clusterBroker hostnames hardcoded in client configsA DNS name (3 brokers is enough) repointed on failover; build failover logic into clients for low RTO
19Consumers on the DR cluster consume from the wrong position after DNS flip"Most failover scenarios DO require BOUNCING consumer applications"Include the bounce in the runbook
20Two-DC "stretch cluster" goes fully down when one DC failsOne DC always holds the ZooKeeper majorityThree datacenters (or 2.5 DC with a tiebreaker ZK node)
21Legacy MirrorMaker stalls for 5–10 minutes when a topic is addedMM1's consumer-group rebalancesMirrorMaker 2.0 — allocates partitions without the group protocol
22Test topics replicated across an expensive WAN linktopics = .*Use prod.* and a test.* exclusion list
23DR topics have weaker durability than productionmin.insync.replicas is NOT migrated by defaultConfigure it explicitly on the target; customize the exclusion list
24After failover, producers can't write to the DR clusterTopic:Write ACLs are deliberately not migrated (so only MM can write)"Appropriate access must be EXPLICITLY GRANTED at the time of failover" — put it in the runbook
25Prefixed/wildcard ACLs missing on the DR cluster"Only LITERAL topic ACLs that match topics being mirrored are migrated"Configure them on the target explicitly
26MirrorMaker overwrote a live consumer group's offsetsIt doesn't — there's an interlock"MirrorMaker does not overwrite offsets if consumers on the target cluster are actively using the target consumer group"
27Mirroring throughput poor; only one task runningtasks.max default is 1Minimum 2; benchmark 1→32 and set just below the tapering point
28MirrorMaker CPU saturatedDecompress + recompress of compressed eventsExpected; watch CPU while scaling tasks. Cluster Linking avoids it entirely
29A noisy topic starves your latency-critical mirrorShared MirrorMaker clusterSeparate MirrorMaker cluster for sensitive topics
30Cross-DC link never reaches available bandwidthDefault TCP buffers/window; slow-start after idleTune client + broker socket buffers, tcp_window_scaling=1, tcp_slow_start_after_idle=0
31Consumer performance collapses after enabling SSLSSL defeats zero-copy on the broker's read pathConsider consume-locally/produce-remotely — and then set acks=all, retries, errors.tolerance=none. Re-measure on modern Java
32Cloud MirrorMaker can't connect to on-prem brokersFirewall blocks inbound connections to on-premRun MirrorMaker on premises (produce remotely)
33Reported lag jumps around by a minute's worthMethod 1 reads committed offsets; MM commits every minute by defaultUse Burrow, or Method 2, understanding both are approximations
34MirrorMaker silently dropped messages and no alert firedBoth lag methods only track the latest offsetMessage counts + checksums (Confluent Control Center), plus a canary
35Ordering broken in the target cluster during retriesmax.in.flight > 1 with retries (Ch. 3 §6)max.in.flight=1 is "currently the only way" for MM to guarantee ordering — at a heavy WAN throughput cost
36Failover plan worked last year, fails nowUntested plan; "a plan that works today may stop working after an upgrade"Practice at least quarterly; Chaos-Monkey-style continuous testing