9. The Trouble with Distributed Systems
9. The Trouble with Distributed Systems
Chapter 9 of Designing Data-Intensive Applications — 14 sections.
"They're funny things, Accidents. You never have them till you're having them." — A.A. Milne
The mindset shift this chapter demands:
If you want your system to be reliable in the presence of faults, you have to RADICALLY CHANGE YOUR MINDSET and focus on what could go wrong, even though it may be unlikely. It doesn't matter whether there is only a one-in-a-million chance; IN A LARGE ENOUGH SYSTEM, ONE-IN-A-MILLION EVENTS HAPPEN EVERY DAY. Experienced systems operators will tell you that ANYTHING THAT CAN GO WRONG WILL GO WRONG.
In distributed systems, SUSPICION, PESSIMISM, AND PARANOIA PAY OFF.
The chapter is deliberately all problems. Ch 10 is the solutions.
- Unreliable networks — packets are lost and delayed arbitrarily, and you cannot tell why you got no response.
- Unreliable clocks — drift, jumps, and no confidence interval.
- Process pauses — a thread can be frozen for minutes, mid-function, without noticing.
- Knowledge and truth — quorums, leases, fencing, Byzantine faults.
- System models — formalizing what we assume; safety versus liveness; and how we test any of this.
Sections
- 9.1Faults and Partial Failures
- 9.25Unreliable Networks
- 9.33Unreliable Clocks
- 9.4Process Pauses
- 9.53Knowledge, Truth, and Lies
- 9.6System Model and Reality
- 9.71Formal Methods and Randomized Testing
- 9.8Deep divesTechnology deep divesφ = −log₁₀(P(heartbeat arrives later than the current elapsed time)).
- 9.9Failure catalogProduction failure catalog for this chapter
- 9.10Decision sheetDecision cheat sheetThere is no correct constant. Measure the RTT distribution across many machines over an extended period, pick a target trade-off between detection delay and false-positive rate, a…
- 9.11Worked examplesWorked examples
- 9.12Self-testSelf-test
- 9.13TerminologyTerminology introduced here
- 9.14Forward linksForward links