Learn Labs
10. Consistency and Consensus

10.7 Decision cheat sheet

Decision cheat sheet
0 rows

Do I need linearizability?

Are you doing any of these?

  • leader election / distributed locking
  • hard uniqueness constraints — usernames, filenames, seat booking
  • anything with a second communication channel — queue + storage, push + fetch

Yes ⇒ you need linearizability for that operation. Use consensus (etcd / ZooKeeper) or a single-leader database with safe failover.

No ⇒ probably not. Prefer a weaker model and pay less latency — remember linearizable systems are slower all the time, not just during faults, and the lower bound is proportional to network delay uncertainty.

Linearizability or serializability — which do I actually want? Serializability if the problem is multi-object transactional correctness (write skew, invariants across rows). Linearizability if the problem is recency of a single object (is this the latest leader? is this username taken?). Both (strict serializability) if you need transactions and recency — and know it costs coordination.

Which ID scheme?

NeedUse
Compact, ordered, single-nodeAutoincrement
Distributed, no coordination, don't care about orderUUIDv4
Distributed, roughly time-ordered, good index localityUUIDv7 / Snowflake / ULID
Causally-consistent ordering across nodesLamport clock or HLC
Detect concurrencyVector clock (pay the space)
Linearizable orderingTimestamp oracle (single node, batched) or Spanner-style clock + commit wait

How many consensus nodes? 3 tolerates 1 failure; 5 tolerates 2. Beyond that you're paying latency for tolerance you don't need. Never even numbers — 4 nodes tolerate the same 1 failure as 3, with more coordination.

Should I use consensus for service discovery? Usually no — it doesn't need linearizability, and availability and speed matter more. Cache aggressively, use TTLs, use observers/read replicas. Use consensus for the leases and leadership; use caching for the lookups.

Rule of thumb for automatic failover:

Any system that provides automatic failover but does not use a proven consensus algorithm is likely to be unsafe. If you're writing your own failover logic, you are writing a consensus algorithm — badly.