6.8 Decision cheat sheet
Automatic if your failover manager does real fencing and your timeout is tuned.
Which replication model?
- Do you need to enforce invariants (uniqueness, non-negative balance)?
- Yes → single-leader, or consensus. Full stop: multi-leader and leaderless cannot do this. It is “a fundamental limitation of distributed systems.”
- No → do writes need to succeed while disconnected from the leader?
- Yes, and it is end-user devices or offline editing → multi-leader with a sync engine and CRDTs.
- Yes, but not end-user devices → multi-leader (geo) or leaderless.
- No → single-leader: simplest and strongest.
Sync, semisync, or async? Never all-synchronous (one outage halts everything). Semisynchronous — one sync follower + async others — is the default correct answer: durability on two nodes, availability preserved. Pure async only when you can genuinely afford to lose the last few seconds of writes.
Automatic or manual failover? Automatic if your failover manager does real fencing and your timeout is tuned. Manual if not — some operations teams prefer manual even when automatic is supported, and that is a defensible engineering position, not laziness.
Which replication log format? Logical/row-based, unless you specifically need byte-identical standbys. It decouples you from the storage format, enables cross-version upgrades, and gives you CDC for free.
How do I fix a replication-lag anomaly?
- read-your-writes → read from the leader for that user's own data, or track a write timestamp/LSN
- monotonic reads → pin each user to one replica by hash of user ID
- consistent prefix → co-locate causally related writes in one shard, or track causality explicitly
Choosing n, w, r?
Start n=3, w=2, r=2 (or n=5, w=3, r=3). Read-heavy with rare writes → w=n, r=1 only if you accept that one dead node stops all writes. Multi-region → LOCAL_QUORUM by default, EACH_QUORUM only where correctness demands it. And treat w+r>n as a probability adjustment, not a guarantee.
LWW, manual, or CRDT? LWW only if you never update existing records (insert-only with unique keys). Manual if conflicts are rare and the domain genuinely needs a human. CRDT/OT for anything collaborative — and accept that invariants over the merged state are not enforceable.