2.11 Self-test
Self-test31 questions
Why is a nonfunctional requirement "just as important" as functionality?
Compute the query-on-read load for the social network from first principles. Which two separate problems does the naive design have?
What is fan-out? Why does materializing timelines trade write cost for read cost, and why is that trade favourable here?
Give the two extreme cases in the timeline design. Why is dropping writes acceptable in one and not the other?
Distinguish response time, service time, queueing delay, and latency. Which does the client experience?
Why must response times be measured on the client side?
What is head-of-line blocking, and why does a small number of slow requests hurt so many fast ones?
Why is the mean a bad "typical" latency? What is it good for?
Reproduce the Amazon p999 argument — why that percentile, and why not p9999?
Why is averaging percentiles meaningless, and what is the correct aggregation?
Define tail latency amplification. Why does parallelizing backend calls not help?
Distinguish an SLO from an SLA.
Define a metastable failure. Draw the loop. Why doesn't removing the load fix it?
Name three client-side and two server-side overload protections, and say what each does.
Distinguish a fault from a failure. What is a SPOF?
Why does it make sense to increase the rate of faults deliberately? What class of bug does this find?
Give the annual failure rates for HDDs and SSDs. Which fact in that list is the most unsettling, and why?
Why is redundancy less effective than the arithmetic suggests?
Why are software faults more dangerous than hardware faults? Give two real examples from the chapter.
What was the leading cause of outages in the study cited? Why is "human error" the wrong conclusion?
What is a blameless postmortem for, and what two simplistic answers should you distrust?
What was the Post Office Horizon scandal, and what legal assumption enabled it?
Why is "X is scalable" a meaningless statement? What three questions replace it?
Beyond throughput, name three statistical characteristics of load that change your design.
Compare shared-memory, shared-disk, and shared-nothing on cost curve, scaling limit, and what they demand of you.
How does the cloud-native storage/compute split differ from classic shared-disk, and why does that difference matter?
Why is there no "magic scaling sauce"? How far ahead should you plan?
Why is more automation not always better for operability?
Give three ways abstraction reduces complexity, and one reason "simplicity" resists definition.
Why is irreversibility the main obstacle to evolvability? Name three practices that reduce it.
- Design question
you own a notification service handling 50,000 events/s, fanning out to email, push, and SMS providers. p99 is 400 ms and rising; during provider outages the whole service becomes unavailable for 20 minutes even after the provider recovers. Diagnose the likely cause, propose the specific protections, define the SLO you'd publish, and say exactly which metrics you'd add — including one you'd deliberately count that isn't an error.