2.3 Reliability and Fault Tolerance
Reliability ≈ "continuing to work correctly, even when things go wrong."
Reliability ≈ "continuing to work correctly, even when things go wrong."
"Working correctly" typically expects:
- The application performs the function the user expected
- It tolerates the user making mistakes or using it in unexpected ways
- Performance is good enough for the use case, under expected load and data volume
- It prevents unauthorized access and abuse
3.1 Fault vs failure
| Term | Definition |
|---|---|
| Fault | A particular part of a system stops working correctly — a hard drive malfunctions, a machine crashes, a dependency has an outage |
| Failure | The system as a whole stops providing the required service to the user — i.e., it does not meet the SLO |
They're the same thing at different levels. A drive that stops working has "failed" as a drive; in a system with multiple drives, that's merely a fault, which the bigger system may tolerate by having a copy elsewhere.
Fault-tolerant = continues providing the required service in spite of certain faults. A part whose fault escalates into whole-system failure is a single point of failure (SPOF).
Case-study example: during fan-out, a machine updating materialized timelines crashes. Fault tolerance requires another machine to take over without missing any posts that should have been delivered, and without duplicating any — that's exactly-once semantics (Ch 12).
Fault tolerance is always bounded: "at most 2 drives at once," "at most 1 of 3 nodes." Tolerating any number of faults is meaningless — if all nodes crash, nothing can be done. (If Earth is swallowed by a black hole, you'd need hosting in space; good luck with the budget line item.)
3.2 Fault injection and chaos engineering
Counterintuitively, in fault-tolerant systems it makes sense to increase the rate of faults deliberately — e.g., randomly killing processes without warning.
Why: many critical bugs are due to poor error handling. Deliberately inducing faults ensures the fault-tolerance machinery is continually exercised and tested, raising confidence it will work when faults occur naturally. Chaos engineering is the discipline built around this.
The exception where prevention beats cure: security. If an attacker compromised the system and got sensitive data, that cannot be undone. This book mostly deals with curable faults.
3.3 Hardware faults — the actual numbers
| Component | Failure characteristics |
|---|---|
| Magnetic HDD | 2%–5% fail per year. In a 10,000-disk cluster → on average one disk failure per day. Getting more reliable, but rates remain significant |
| SSD | 0.5%–1% fail per year. Small bit errors auto-corrected, but uncorrectable errors occur ~once per year per drive, even in nearly-new drives — a higher error rate than magnetic HDDs |
| PSUs, RAID controllers, memory modules | Fail too, less often than disks |
| CPU cores | ~1 in 1,000 machines has a core that occasionally computes the WRONG result, likely from manufacturing defects. Sometimes crashes; sometimes just returns a wrong answer |
| RAM | Corruption from cosmic rays or permanent physical defects. Even with ECC, >1% of machines hit an uncorrectable error per year, typically crashing the machine and requiring module replacement. Certain pathological memory access patterns flip bits with high probability (Rowhammer) |
| Datacenter | Power outage, network misconfiguration; or permanent destruction by fire, flood, earthquake. A solar storm could damage power grids and undersea cables. Rare, but catastrophic if you can't tolerate losing a DC |
The "1 in 1,000 CPUs silently computes wrong results" fact is the scariest one here. It breaks the assumption every program makes. It's why large operators run continuous verification and why Byzantine-ish thinking (Ch 9) isn't purely academic.
Small system: these are rare enough to ignore as long as you can replace faulty hardware easily. Large system: hardware faults happen often enough that they become part of normal system operation.
3.4 Redundancy — and its limits
First response: hardware redundancy. RAID across disks in a machine; dual power supplies; hot-swappable CPUs; datacenter batteries and diesel generators. This can keep a machine running for years.
Redundancy is most effective when component faults are INDEPENDENT — when one fault doesn't change the likelihood of another. Experience shows significant correlations between component failures. Whole-rack and whole-datacenter unavailability still happens more often than we'd like.
Hence the cloud posture: focus less on the reliability of individual machines; make services highly available by tolerating faulty nodes at the software level. Cloud providers expose availability zones to tell you which resources are physically co-located — co-located resources are more likely to fail together.
Operational bonus of machine-level fault tolerance: a single-server system needs planned downtime to reboot for OS security patches; a multi-node fault-tolerant system is patched by restarting one node at a time without affecting users — a rolling upgrade (Ch 5).
3.5 Software faults — the correlated, dangerous kind
Hardware faults are weakly correlated but mostly independent. Software faults are often HIGHLY correlated, because many nodes run the same software and therefore have the same bugs. Harder to anticipate, and they cause many more system failures than uncorrelated hardware faults.
Real examples given:
- The 2012 leap second: a Linux kernel bug caused many Java applications to hang simultaneously, taking down several internet services.
- The 32,768-hour SSD firmware bug: all SSDs of certain models fail after precisely 32,768 hours (<4 years) of operation, rendering data unrecoverable. (Note the number: 2¹⁵ — a signed 16-bit counter overflowing.)
- Runaway process consuming a shared limited resource: CPU, memory, disk space, network bandwidth, or threads. A process consuming too much memory on a large request gets OOM-killed; a client-library bug drives far higher request volume than anticipated.
- A dependency slows down, becomes unresponsive, or returns corrupted responses.
- Emergent behavior from interactions between systems that doesn't occur when each is tested in isolation.
- Cascading failures: one component's problem overloads another, slowing it, which brings down a third.
The general shape: these bugs lie dormant for a long time until an unusual set of circumstances triggers them. At that point it's revealed that the software made an assumption about its environment — usually true, and eventually not.
No quick solution. Lots of small things help: carefully thinking about assumptions and interactions; thorough testing; process isolation; allowing processes to crash and restart; avoiding feedback loops like retry storms; measuring, monitoring, and analyzing behavior in production.
3.6 Humans and reliability
One study of large internet services found that CONFIGURATION CHANGES BY OPERATORS were the leading cause of outages, with hardware faults playing a role in only 10%–25% of cases.
But: blaming people for mistakes is counterproductive. "Human error" is not the cause of an incident — it's a symptom of a problem with the sociotechnical system in which people are doing their best. Complex systems have emergent behavior; unexpected interactions between components also cause failures.
Technical measures that reduce the impact of human mistakes:
- Thorough testing — handwritten tests and property testing over lots of random inputs
- Rollback mechanisms for quickly reverting config changes
- Gradual rollouts of new code
- Detailed, clear monitoring; observability tools for diagnosing production issues
- Well-designed interfaces that encourage the right thing and discourage the wrong thing
And the honest organizational point: all of these cost time and money. Given a choice between more features and more testing, many organizations understandably choose features. So when a preventable mistake occurs, blaming the individual makes no sense — the problem is the organization's priorities.
Blameless postmortems: after an incident, people share full details without fear of punishment, so others can learn to prevent similar problems. The process may reveal a need to change business priorities, invest in neglected areas, change incentives, or escalate a systemic issue to management.
Be suspicious of simplistic answers. "Bob should have been more careful deploying that change" is not productive — but neither is "we must rewrite the backend in Haskell." Management should learn how the sociotechnical system actually works from the people who work with it daily, and improve it from that feedback.
3.7 How important is reliability? — the Post Office Horizon scandal
Reliability isn't only for nuclear plants and air traffic control. A few minutes or hours of outage is tolerable in many applications; permanent data loss or corruption is catastrophic. (Consider a parent whose entire photo/video record of their children lives in your app. Would they even know how to restore from a backup?)
The Horizon case (1999–2019): hundreds of British Post Office branch managers were convicted of theft or fraud because the accounting software showed shortfalls in their accounts. Many of those shortfalls were software bugs. Convictions were eventually overturned — probably the largest miscarriage of justice in British history.
The enabling cause is legal, not technical: English law assumed computers operate correctly, and therefore that computer-produced evidence is reliable, unless evidence exists to the contrary. Engineers may laugh at the idea of bug-free software, but that's "little solace to the people who were wrongfully imprisoned, declared bankruptcy, or even committed suicide."
Sacrificing reliability to cut development cost is sometimes right (prototyping for an unproven market) — but be conscious that you're cutting corners, and keep the consequences in mind.
2.2 Describing Performance
In the case study: posts/s and timeline writes/s are throughput; time to load the home timeline and time until a post reaches followers are response times.
2.4 Scalability
For a new product with few users, the overriding engineering goal is keeping the system simple and flexible so you can adapt as you learn what customers need.