6. Replication
6.7 Production failure catalog for this chapter
| Symptom | Underlying mechanism |
|---|---|
| Data deleted by mistake is gone from every replica | Replication is not backup |
| Writes acknowledged to the client vanish after failover | Async replication + promotion of a lagging follower |
| Primary keys reused; private data shown to wrong users | Promoted follower's autoincrement counter behind; cross-system (Redis) inconsistency — the GitHub incident |
| Two nodes both accepting writes | Split brain; missing or broken fencing |
| Cluster failed over during a load spike, making it worse | Timeout too short → unnecessary failover |
| Both nodes shut down during split-brain protection | Poorly designed mutual-shutdown mechanism |
| Cannot upgrade the database without downtime | WAL shipping couples leader and follower to one storage format |
| User submits a form and it "didn't save" | No read-after-write consistency |
| A comment appears, then disappears on refresh | No monotonic reads — reads hitting different-lag replicas |
| An answer appears before the question | No consistent prefix reads across shards |
| Primary's disk fills with WAL | Inactive replication slot / retained log for a dead follower |
| Standby queries cancelled with "conflict with recovery" | Replay vs long-running read on the standby |
| Removed cart items reappear | Set-union merge of siblings — the Amazon anomaly |
Merged value is B/C/C/B | Concurrent conflict resolution creating a new conflict |
| A username was registered twice in two regions | Multi-leader cannot enforce cross-leader uniqueness |
| An UPDATE arrives before its INSERT | All-to-all topology overtaking; timestamps insufficient |
| Writes silently dropped in Cassandra | LWW + clock skew |
| Deleted rows come back | Repair not run within gc_grace_seconds |
| A write returned an error but the data is there | Partial write below w is not rolled back |
| Quorum reads still return stale data | Rebalancing, restored-from-old-replica, or concurrent read/write (§4.4) |
| Cannot tell how stale a leaderless cluster is | No fixed write order ⇒ no lag metric |