Learn Labs
2. Defining Nonfunctional Requirements

2.7 Production failure catalog for this chapter

Production failure catalog
0 rows
SymptomUnderlying concept
Latency fine at 60% load, catastrophic at 85%Queueing delays rise sharply near capacity
Outage persists after the traffic spike endsMetastable failure / retry storm
One user action produces 81 backend callsRetries stacked at multiple layers, no retry budget
p99 looks fine, users complain constantlyMeasured server-side; missing queueing + network latency
Dashboard p99 disagrees with realityAveraged percentiles across time or instances
Benchmark says p99 = 12 ms, prod says 900 msCoordinated omission in the load generator
End-user p50 ≈ backend p99Tail latency amplification across many backend calls
A "fast" endpoint is slow only for big customersPer-user data volume skew — the Amazon p999 argument
Everything fails at once, same errorCorrelated software fault — same bug on every node
SSDs all die in the same week, 4 years inFirmware bug at 32,768 hours (2¹⁵ counter overflow)
Wrong results, no crash, no errorSilent data corruption — ~1 in 1,000 machines has a bad CPU core
Whole rack unavailable despite redundancyCorrelated component failures; redundancy assumed independence
Leading cause of outages overallOperator configuration changes, not hardware
Incident recurs after "we told Bob to be careful"Blame instead of blameless postmortem; sociotechnical cause untouched
Migration can't be rolled backIrreversibility — the main obstacle to evolvability
Autoscaler oscillates and causes incidentsAutomation added where load was predictable enough for manual scaling