9.11 Worked examples
① Why you can't distinguish the six cases. Client sends request at t=0, timeout at t=5 s, no response. Enumerate the posterior: request lost (no effect), request queued (effect may still happen at t=90 s), node dead before processing (no effect), node paused (effect happens when it resumes), response lost (effect already happened), response delayed (effect already happened). Three of six cases mean the write already applied; two mean it may apply later. Hence: every non-idempotent operation over a network needs a deduplication key (Ch 8 §5.6).
② Drift budget. 200 ppm = 200 µs per second.
- Resync every 30 s → 6 ms max drift
- Resync every 60 s → 12 ms
- Resync once per hour → 720 ms
- Resync once per day → 17.3 s If your application assumes clocks agree within 100 ms, you need to resync at least every ~8 minutes — and that's the drift term alone, before NTP's own 35 ms+ error.
③ Why LWW loses data proportionally to skew. Node A's clock is 200 ms ahead. A writes x at true time T (stamped T+200 ms). B writes x at true time T+150 ms (stamped T+150 ms). B's write is genuinely later but has the smaller timestamp, so it loses. Every write B makes in the 200 ms window after any A write is silently discarded — and at 1,000 writes/s that's 200 lost writes per A-write, with no error anywhere.
④ Spanner's commit-wait cost. Uncertainty ε = 7 ms ⇒ commit-wait ≈ 2ε ≈ 14 ms added to every read/write transaction's commit. Halve ε to 3.5 ms (better clock hardware) and you halve the wait. This is the direct, quantifiable reason Google puts atomic clocks in datacenters — it's not for accuracy per se, it's to shrink a latency tax on every write.
⑤ Lease safety margin. Lease duration L = 30 s, renewed when < 10 s remain, so the renewal deadline gives 10 s of slack. Worst realistic pause: a 15 s GC or VM suspend. 10 s slack < 15 s pause ⇒ the zombie window is real. Options: raise the margin above the worst pause (slower failover), reduce the worst pause (GC tuning, §4.3), or — the only sound answer — make the pause harmless with fencing tokens, since you cannot bound the pause.
⑥ Quorum sizes and majorities. n=5, majority=3. Two disjoint majorities would need ≥6 nodes, so there can only ever be one — that's the entire safety argument. Now suppose one node suffers amnesia after a disk wipe and rejoins claiming to have data it lost: the majority intersection guarantee still holds numerically but the value it returns is wrong, which is §6.4's point that the model's stable-storage assumption is doing real work.