Learn Labs
10. Consistency and Consensus

10.6 Production failure catalog for this chapter

Production failure catalog
0 rows
SymptomUnderlying mechanism
Two users see different "current" values seconds apartNot linearizable — stale replica read
A transcoder processes an old version of a fileCross-channel race: message queue faster than storage replication
Push notification arrives before the data it referencesSame — two channels, no recency guarantee
Two nodes both act as leaderSplit brain — leader election without consensus
A username was registered twiceUniqueness enforced without linearizable CAS
Committed writes lost after failoverAsynchronous replication + promotion; or unclean leader election
Quorum reads still return stale valuesw + r > n does NOT imply linearizability
Cassandra "strong consistency" isn'tTime-of-day-clock LWW breaks linearizability
Reads from a consensus cluster are staleRead served without a quorum check that the leader is still current
ZooKeeper read returned old dataZooKeeper reads are not linearizable without sync
Photo visible despite a prior privacy changeNon-linearizable ID generator + MVCC snapshot
Cluster spends all its time electing leadersElection timeout too small vs real jitter/GC
Leadership bounces between two nodes foreverRaft edge case on one bad link (fixed by pre-vote)
Adding nodes made the cluster slowerConsensus needs a quorum per operation — more nodes = slower
Minority partition frozenConsensus requires a strict majority — by design
etcd went read-only, Kubernetes stoppedStorage quota exceeded
Duplicate primary keys across the fleetColliding machine IDs in a Snowflake-style generator
ID generator refuses to issue IDsClock stepped backward
Every write got slower after a GPS faultTrueTime ε widened → longer commit wait
Coordination service melted under loadUsed as a database for fast-changing data