Learn Labs
6. Replication

6.10 Self-test

Self-test23 questions

—/23
  1. Why doesn't replication remove the need for backups? Give the scenario where replication actively hurts.

  2. Why is it impracticable for all followers to be synchronous? What is the semisynchronous compromise, and what exactly does it guarantee?

  3. Describe the four steps of adding a new follower without downtime. What must the snapshot be associated with, and what is that called in Postgres and MySQL?

  4. A follower has been offline for a week. State the leader's dilemma and both bad outcomes.

  5. List the three steps of automatic failover, and the four things that can go wrong. Which one caused the GitHub incident, and how?

  6. Why does a short failover timeout make an overloaded system worse?

  7. Compare the three replication-log formats on: coupling to the storage engine, ability to run different versions on leader and follower, and external parseability.

  8. Why does WAL shipping prevent zero-downtime upgrades? Why doesn't logical replication?

  9. Define eventual consistency. Why is "eventually" deliberately vague?

  10. For each of read-after-write, monotonic reads, and consistent prefix reads: state the anomaly in one sentence and give one implementation technique.

  11. Why does cross-device read-after-write consistency break the "remember the client's last write timestamp" approach?

  12. Why does the book treat synchronous multi-leader replication as equivalent to single-leader?

  13. Give the precise reason multi-leader replication cannot enforce username uniqueness.

  14. Compare circular, star, and all-to-all topologies on failure tolerance. What problem is unique to all-to-all, and why don't timestamps fix it?

  15. Explain how an app with no offline mode can still be multi-leader.

  16. State the "real meaning" of LWW. Under what single condition is LWW harmless?

  17. Reproduce the Amazon shopping-cart anomaly and explain what a proper collection CRDT does differently.

  18. Write the quorum condition. Why is it a probability adjustment rather than a guarantee? Give three of the six edge cases.

  19. Why is staleness easy to monitor in leader-based replication and hard in leaderless?

  20. Define happens-before. Why can two operations be concurrent even when the speed of light would have permitted causality?

  21. Walk through the shopping-cart version-number trace and explain why no writes are lost despite the clients never being up to date.

  22. Why is a single version number insufficient with multiple replicas? What replaces it, and what does it let the database distinguish?

  23. Design question

    you run a SaaS with users in the US, EU, and APAC. Requirements: (a) usernames globally unique, (b) users' own documents editable offline on mobile, (c) survive the loss of an entire region, (d) p99 write latency < 100 ms in-region. Design the replication topology — you will need more than one model. State exactly which data uses which, and what you give up.