Learn Labs
10. Cross-Cluster Data Mirroring

10.12 Self-test

Self-test40 questions

—/40
  1. Distinguish "replication" from "mirroring" in Kafka's vocabulary.

  2. Name the five use cases for cross-cluster mirroring. Which one reduces requirements on the source clusters, and how?

  3. Give the three realities of cross-datacenter communication. Why does high latency make bandwidth harder to use?

  4. Why is it not recommended to spread one ordinary cluster's brokers across datacenters?

  5. Explain precisely why remote consuming is safer than remote producing.

  6. State the three guiding principles for multicluster architecture.

  7. Draw hub-and-spoke. What is its single defining limitation, and what does the bank example illustrate?

  8. In hub-and-spoke, where do the mirroring processes run, and why?

  9. What are the two benefits of active-active? Which type of failover does the second one enable?

  10. Describe both conflict types in active-active and the standard mitigation for the first.

  11. Explain the topic-namespace trick that prevents mirroring loops. What must consumers subscribe to?

  12. What's the alternative to namespaces, and what's the catch?

  13. Give both disadvantages of active-standby. Quote the "bottom line" about Kafka failover.

  14. Why is a smaller DR cluster a risky economy?

  15. Define RTO and RPO. What does each one force on your architecture at its extreme?

  16. Do the unplanned-failover data-loss arithmetic for 1M msg/s and 5 ms lag. How does planned failover differ?

  17. Why can a line item arrive without its sale after failover?

  18. List the four failover-offset strategies with their trade-offs.

  19. Give all three caveats of mirroring __consumer_offsets, including the retention-skew example.

  20. Why is time-based failover recommended despite being imprecise? Give the social argument.

  21. How does offset translation store mappings efficiently? Walk the 495/500 → 596/600 example.

  22. After a successful failover, why can't you just reverse the mirroring direction? What's the remedy?

  23. Why do most failover scenarios require bouncing consumers, and what would avoid it?

  24. Why does a stretch cluster need three datacenters rather than two? What is 2.5 DC?

  25. What does a stretch cluster protect against — and what does it not protect against?

  26. What was MM1's central flaw? What are MM2's three key design decisions?

  27. What does MM2 migrate besides data records?

  28. What happens to a topic's name when it's mirrored, and what two problems does that solve?

  29. Which topic config is not migrated by default, and which ACL is deliberately not migrated? What's the failover consequence of each?

  30. Give MirrorMaker's full ACL requirements, source and target.

  31. Where should MirrorMaker run by default? Name the two exceptions and the reason for each.

  32. Why do consumers suffer more from SSL than producers? What must you configure if you produce remotely because of it?

  33. Give both lag-monitoring methods, why each is inaccurate, and the gap neither one covers.

  34. Describe the procedure for finding the right tasks.max. What resource must you watch, and why?

  35. How do you determine whether MirrorMaker's producer or consumer is the bottleneck? Give two methods.

  36. Why is max.in.flight=1 recommended for MirrorMaker when Ch. 3 said idempotence solves this? What does it cost over a WAN?

  37. What problem did uReplicator solve, how, and what did MM2 do about it?

  38. Compare MM2 and Confluent Replicator on ACL migration, offset translation, cycle prevention, and schema handling.

  39. What is an "observer" in MRC, and what does automatic observer promotion accomplish?

  40. How does Cluster Linking preserve offsets? Why must mirror topics be read-only, and what performance cost does it avoid?

    Previous: Chapter 9 — Building Data Pipelines Next: Chapter 11 — Securing Kafka