Learn Labs
2. Defining Nonfunctional Requirements

2.8 Decision cheat sheet

p50 for "typical user experience," p99 for the SLO, p999 only if your slowest requests correlate with your most valuable customers (the Amazon case).

Which latency number do I put in the SLO? p50 for "typical user experience," p99 for the SLO, p999 only if your slowest requests correlate with your most valuable customers (the Amazon case). Skip p9999 — too expensive, too noisy, diminishing returns.

Where do I measure? Client side. Server-side timers exclude queueing delay, which is where the variance lives. Instrument both and treat the gap as a signal in itself.

Read-heavy feed: query on read, or materialize on write? Materialize when read volume × query cost ≫ write volume × fan-out. Go hybrid the moment the fan-out distribution has a long tail — write-path for the many, read-path merge for the few.

How do I stop a retry storm? Jittered exponential backoff + a retry budget (token bucket capping retries as a fraction of traffic) + circuit breakers + server-side load shedding with bounded queues. Any one alone is insufficient.

Vertical or horizontal? Vertical while it's cheap and simple — it usually is, for longer than people assume. Horizontal when you need multi-DC fault tolerance, elasticity, or you've hit the price cliff. Remember horizontal costs you explicit sharding plus all of Ch 9.

How far ahead should I design for scale? One order of magnitude. Not two. Expect to rethink the architecture at each 10×.

When is more automation wrong? When load is predictable (manual scaling has fewer surprises), and when the automation makes failures harder to troubleshoot than a manual procedure would be.