10. Cross-Cluster Data Mirroring
10.10 What actually breaks in production — Ch. 10 consolidated
| # | Symptom | Root cause | Fix |
|---|---|---|---|
| 1 | Brokers spread across DCs behave badly; timeouts everywhere | Kafka was "designed, developed, tested, and tuned, all within a single datacenter" — defaults assume LAN | Don't do it, except as a deliberate stretch cluster with 3 DCs |
| 2 | Data lost during a WAN partition | MirrorMaker was producing remotely; consumed events couldn't be delivered | Run MirrorMaker at the target DC — remote consuming fails safe |
| 3 | Enormous cross-DC bandwidth bill | Multiple applications each consuming the same data across the WAN | One cluster per DC; mirror once, consume locally |
| 4 | A user visits another city's branch and their data isn't there | Hub-and-spoke gives no cross-regional data access | Only use it for data that "can be completely separated between regional datacenters"; otherwise active-active |
| 5 | Events mirrored back and forth forever | Active-active without cycle prevention | Per-DC topic namespaces (SF.users / NYC.users); MM2 does this by default via alias prefixing |
| 6 | User writes to one DC, reads from another, and doesn't see their own write | Asynchronous multi-master | "Stick" each user to a datacenter |
| 7 | Two conflicting orders for the same user; downstream state diverges | Concurrent writes in two DCs | "You WILL have conflicts" — define consistent resolution rules both DCs will compute identically |
| 8 | N² mirroring processes to manage | Active-active needs a flow per pair per direction | Tools that share processes per destination cluster; one shared config file |
| 9 | Custom origin-DC headers break mirroring | "None of the existing mirroring tools will support your specific header format" | Use the tool's own convention, or accept the extra work |
| 10 | DR cluster too small to carry production during a real disaster | Cost-driven undersizing | "A risky decision because you can't be sure it will hold up" |
| 11 | Failover loses ~5,000 messages | All mirroring is asynchronous: 1M msg/s × 5 ms = 5,000 messages | Expected. Planned failover avoids it (stop primary, let mirroring drain). Monitor DR lag continuously |
| 12 | A line item arrives with no corresponding sale | "Mirroring solutions currently don't support transactions" — related events across topics arrive independently | Applications must tolerate it (Ch. 8 §3.6) |
| 13 | Consumers fail on the DR cluster: offset doesn't exist | Mirrored __consumer_offsets but offsets diverge (retention skew, producer retries) | Time-based failover, or offset translation (MM2) |
| 14 | DR consumer finds a committed offset with no matching record | Offset commits and records mirror independently and race | Decide in advance: beginning or end? |
| 15 | Failover happened but nobody can explain what was reprocessed | Offset-based failover is inherently unexplainable | Time-based failover — "'We failed back to 4:03 a.m.' sounds better" |
| 16 | Offset reset tool silently did nothing / conflicted | Consumer group still running | "The group should be STOPPED while running this type of tool and started immediately after" |
| 17 | After failing back, the two clusters are permanently inconsistent | The old primary retained events the DR never got — reverse mirroring leaves phantom history | "First SCRAPE the original cluster — delete all data and committed offsets" |
| 18 | Applications can't find the DR cluster | Broker hostnames hardcoded in client configs | A DNS name (3 brokers is enough) repointed on failover; build failover logic into clients for low RTO |
| 19 | Consumers on the DR cluster consume from the wrong position after DNS flip | "Most failover scenarios DO require BOUNCING consumer applications" | Include the bounce in the runbook |
| 20 | Two-DC "stretch cluster" goes fully down when one DC fails | One DC always holds the ZooKeeper majority | Three datacenters (or 2.5 DC with a tiebreaker ZK node) |
| 21 | Legacy MirrorMaker stalls for 5–10 minutes when a topic is added | MM1's consumer-group rebalances | MirrorMaker 2.0 — allocates partitions without the group protocol |
| 22 | Test topics replicated across an expensive WAN link | topics = .* | Use prod.* and a test.* exclusion list |
| 23 | DR topics have weaker durability than production | min.insync.replicas is NOT migrated by default | Configure it explicitly on the target; customize the exclusion list |
| 24 | After failover, producers can't write to the DR cluster | Topic:Write ACLs are deliberately not migrated (so only MM can write) | "Appropriate access must be EXPLICITLY GRANTED at the time of failover" — put it in the runbook |
| 25 | Prefixed/wildcard ACLs missing on the DR cluster | "Only LITERAL topic ACLs that match topics being mirrored are migrated" | Configure them on the target explicitly |
| 26 | MirrorMaker overwrote a live consumer group's offsets | It doesn't — there's an interlock | "MirrorMaker does not overwrite offsets if consumers on the target cluster are actively using the target consumer group" |
| 27 | Mirroring throughput poor; only one task running | tasks.max default is 1 | Minimum 2; benchmark 1→32 and set just below the tapering point |
| 28 | MirrorMaker CPU saturated | Decompress + recompress of compressed events | Expected; watch CPU while scaling tasks. Cluster Linking avoids it entirely |
| 29 | A noisy topic starves your latency-critical mirror | Shared MirrorMaker cluster | Separate MirrorMaker cluster for sensitive topics |
| 30 | Cross-DC link never reaches available bandwidth | Default TCP buffers/window; slow-start after idle | Tune client + broker socket buffers, tcp_window_scaling=1, tcp_slow_start_after_idle=0 |
| 31 | Consumer performance collapses after enabling SSL | SSL defeats zero-copy on the broker's read path | Consider consume-locally/produce-remotely — and then set acks=all, retries, errors.tolerance=none. Re-measure on modern Java |
| 32 | Cloud MirrorMaker can't connect to on-prem brokers | Firewall blocks inbound connections to on-prem | Run MirrorMaker on premises (produce remotely) |
| 33 | Reported lag jumps around by a minute's worth | Method 1 reads committed offsets; MM commits every minute by default | Use Burrow, or Method 2, understanding both are approximations |
| 34 | MirrorMaker silently dropped messages and no alert fired | Both lag methods only track the latest offset | Message counts + checksums (Confluent Control Center), plus a canary |
| 35 | Ordering broken in the target cluster during retries | max.in.flight > 1 with retries (Ch. 3 §6) | max.in.flight=1 is "currently the only way" for MM to guarantee ordering — at a heavy WAN throughput cost |
| 36 | Failover plan worked last year, fails now | Untested plan; "a plan that works today may stop working after an upgrade" | Practice at least quarterly; Chaos-Monkey-style continuous testing |