10. Consistency and Consensus
10.6 Production failure catalog for this chapter
| Symptom | Underlying mechanism |
|---|---|
| Two users see different "current" values seconds apart | Not linearizable — stale replica read |
| A transcoder processes an old version of a file | Cross-channel race: message queue faster than storage replication |
| Push notification arrives before the data it references | Same — two channels, no recency guarantee |
| Two nodes both act as leader | Split brain — leader election without consensus |
| A username was registered twice | Uniqueness enforced without linearizable CAS |
| Committed writes lost after failover | Asynchronous replication + promotion; or unclean leader election |
| Quorum reads still return stale values | w + r > n does NOT imply linearizability |
| Cassandra "strong consistency" isn't | Time-of-day-clock LWW breaks linearizability |
| Reads from a consensus cluster are stale | Read served without a quorum check that the leader is still current |
| ZooKeeper read returned old data | ZooKeeper reads are not linearizable without sync |
| Photo visible despite a prior privacy change | Non-linearizable ID generator + MVCC snapshot |
| Cluster spends all its time electing leaders | Election timeout too small vs real jitter/GC |
| Leadership bounces between two nodes forever | Raft edge case on one bad link (fixed by pre-vote) |
| Adding nodes made the cluster slower | Consensus needs a quorum per operation — more nodes = slower |
| Minority partition frozen | Consensus requires a strict majority — by design |
| etcd went read-only, Kubernetes stopped | Storage quota exceeded |
| Duplicate primary keys across the fleet | Colliding machine IDs in a Snowflake-style generator |
| ID generator refuses to issue IDs | Clock stepped backward |
| Every write got slower after a GPS fault | TrueTime ε widened → longer commit wait |
| Coordination service melted under load | Used as a database for fast-changing data |