2. Defining Nonfunctional Requirements
2.7 Production failure catalog for this chapter
| Symptom | Underlying concept |
|---|---|
| Latency fine at 60% load, catastrophic at 85% | Queueing delays rise sharply near capacity |
| Outage persists after the traffic spike ends | Metastable failure / retry storm |
| One user action produces 81 backend calls | Retries stacked at multiple layers, no retry budget |
| p99 looks fine, users complain constantly | Measured server-side; missing queueing + network latency |
| Dashboard p99 disagrees with reality | Averaged percentiles across time or instances |
| Benchmark says p99 = 12 ms, prod says 900 ms | Coordinated omission in the load generator |
| End-user p50 ≈ backend p99 | Tail latency amplification across many backend calls |
| A "fast" endpoint is slow only for big customers | Per-user data volume skew — the Amazon p999 argument |
| Everything fails at once, same error | Correlated software fault — same bug on every node |
| SSDs all die in the same week, 4 years in | Firmware bug at 32,768 hours (2¹⁵ counter overflow) |
| Wrong results, no crash, no error | Silent data corruption — ~1 in 1,000 machines has a bad CPU core |
| Whole rack unavailable despite redundancy | Correlated component failures; redundancy assumed independence |
| Leading cause of outages overall | Operator configuration changes, not hardware |
| Incident recurs after "we told Bob to be careful" | Blame instead of blameless postmortem; sociotechnical cause untouched |
| Migration can't be rolled back | Irreversibility — the main obstacle to evolvability |
| Autoscaler oscillates and causes incidents | Automation added where load was predictable enough for manual scaling |