Learn Labs
6. Replication

6.7 Production failure catalog for this chapter

Production failure catalog
0 rows
SymptomUnderlying mechanism
Data deleted by mistake is gone from every replicaReplication is not backup
Writes acknowledged to the client vanish after failoverAsync replication + promotion of a lagging follower
Primary keys reused; private data shown to wrong usersPromoted follower's autoincrement counter behind; cross-system (Redis) inconsistency — the GitHub incident
Two nodes both accepting writesSplit brain; missing or broken fencing
Cluster failed over during a load spike, making it worseTimeout too short → unnecessary failover
Both nodes shut down during split-brain protectionPoorly designed mutual-shutdown mechanism
Cannot upgrade the database without downtimeWAL shipping couples leader and follower to one storage format
User submits a form and it "didn't save"No read-after-write consistency
A comment appears, then disappears on refreshNo monotonic reads — reads hitting different-lag replicas
An answer appears before the questionNo consistent prefix reads across shards
Primary's disk fills with WALInactive replication slot / retained log for a dead follower
Standby queries cancelled with "conflict with recovery"Replay vs long-running read on the standby
Removed cart items reappearSet-union merge of siblings — the Amazon anomaly
Merged value is B/C/C/BConcurrent conflict resolution creating a new conflict
A username was registered twice in two regionsMulti-leader cannot enforce cross-leader uniqueness
An UPDATE arrives before its INSERTAll-to-all topology overtaking; timestamps insufficient
Writes silently dropped in CassandraLWW + clock skew
Deleted rows come backRepair not run within gc_grace_seconds
A write returned an error but the data is therePartial write below w is not rolled back
Quorum reads still return stale dataRebalancing, restored-from-old-replica, or concurrent read/write (§4.4)
Cannot tell how stale a leaderless cluster isNo fixed write order ⇒ no lag metric