Learn Labs
2. Defining Nonfunctional Requirements

2.2 Describing Performance

In the case study: posts/s and timeline writes/s are throughput; time to load the home timeline and time until a post reaches followers are response times.

Tail amplification

13 ms
median
20 ms
mean
requests hitting the slow path18%
requests entirely fast82%
Safe

With 20 calls per request, 18% of user requests touch at least one p99-slow backend call. A 1% tail in a single service becomes a 18% problem for users — and the mean (20 ms) reports none of it. This is why the book insists on percentiles of the end-to-end request, measured at the client.

If one user request fans out to several backend calls, the slowest call sets the user-visible latency — so the tail you have to care about is not the tail of one service.

Two metric families:

  • Response time — elapsed time from the user making a request to receiving the answer. Unit: seconds/ms/µs.
  • Throughput — requests per second, or data volume per second, being processed. For a given hardware allocation there is a maximum throughput. Unit: "somethings per second."

In the case study: posts/s and timeline writes/s are throughput; time to load the home timeline and time until a post reaches followers are response times.

2.1 The relationship: queueing

capacityflat while under-loadedqueueing delay explodes near capacityresponse timethroughput
Figure 2.2.12.1 The relationship: queueing

Why: when a request arrives at a highly loaded system, the CPU is likely already handling an earlier request, so the new one waits. As throughput approaches hardware maximum, queueing delays increase sharply.

Which metric matters to whom: response time is what users care about most. Throughput determines the computing resources required (how many servers) and therefore the cost of serving the workload. A system is scalable if its maximum throughput can be significantly increased by adding computing resources.

2.2 Metastable failure — when an overloaded system won't recover

The vicious cycle:

more loadload risesqueue growsresponse times riseclients time outclients resendretry stormThe loop closes, so the system stays down after the original spike passes — a metastable failure.
Figure 2.2.2The vicious cycle

Key property: even when the original load is removed, the system may stay overloaded until it is rebooted or reset. That's what makes it metastable — the failure state is self-sustaining. This causes serious production outages.

Countermeasures — memorize this table, it's the most operationally reusable content in the chapter:

SideTechniqueWhat it does
ClientExponential backoffincrease and randomize the time between successive retries (randomization = jitter, to break synchronized retry waves)
ClientCircuit breakertemporarily stop sending requests to a service that recently errored or timed out
ClientToken bucketrate-limit retries to a fixed budget
ServerLoad sheddingdetect approaching overload and proactively reject requests
ServerBackpressuresend responses asking clients to slow down
BothQueueing & load-balancing algorithm choicecan materially change behavior under saturation

2.3 Latency vs response time — precise definitions

what the client actually waits forresponse time
  • network latency (out)
  • queueing delay — waiting for a CPU
  • service time — actively processing
  • network latency (back)
Figure 2.2.32.3 Latency vs response time — precise definitions
TermMeaning
Response timeWhat the client sees. Includes all delays incurred anywhere in the system.
Service timeDuration the service is actively processing the request.
Queueing delayWaiting, at several possible points: for a CPU to become available; for an outbound network buffer when other tasks on the machine are sending a lot of data.
LatencyCatchall for time when the request is not being actively processed — i.e., it is latent.
Network latency / delayTime the request and response spend traveling through the network.

Sources of random per-request delay: a context switch to a background process; a lost network packet and TCP retransmission; a garbage collection pause; a page fault forcing a disk read; mechanical vibrations in the server rack.

Head-of-line blocking. A server processes only a small number of things in parallel (bounded by CPU cores). It takes only a small number of slow requests to hold up all subsequent ones. Those later requests may have fast service times but the client still sees a slow response.

Therefore: queueing delay is not part of service time, and this is exactly why you must measure response times on the CLIENT side. Server-side timers systematically under-report the thing users actually experience.

2.4 Average, median, percentiles

Response time is a distribution, not a number. Variation in network delay is called jitter.

  • Mean (arithmetic average) — useful for estimating throughput limits. Bad for "typical" response time: it doesn't tell you how many users actually experienced that delay.
  • Median = p50 — sort fastest→slowest, take the halfway point. p50 = 200 ms means half of requests are faster, half slower. Good metric for "how long users typically wait."
  • p95 / p99 / p999 — how bad the outliers are. p95 = 1.5 s means 5 out of 100 requests take ≥ 1.5 s.

Tail latencies (high percentiles) directly affect user experience.

The Amazon argument (important and non-obvious): Amazon specifies internal service response times at the 99.9th percentile, even though it affects only 1 in 1,000 requests — because the customers with the slowest requests are usually those with the most data on their accounts, i.e., the ones who bought the most, i.e., the most valuable customers. Conversely Amazon judged optimizing p99.99 (1 in 10,000) too expensive for insufficient benefit — very high percentiles are easily affected by random events outside your control and returns diminish.

2.5 What the latency-vs-revenue data actually says

The book is unusually careful here, and it's worth copying the skepticism:

StudyClaimStatus
Google 2006400 ms → 900 ms slowdown ≈ 20% drop in traffic and revenueOften cited, unreliable
Google 2009+400 ms latency → only 0.6% fewer searches/dayContradicts the above
Bing 2009+2 s load time → 4.3% less ad revenue—
Akamai (recent)+100 ms → up to 7% lower ecommerce conversionSame study shows very fast page loads also correlate with LOWER conversion — because the fastest pages are often ones with no useful content (404s). No attempt to separate content effects from load-time effects → probably not meaningful
Yahoo (following year)20–30% more clicks on fast searches when the fast/slow difference is ≥1.25 sControlled for search-result quality — the most trustworthy of these

Lesson: be suspicious of the folklore numbers; the causal claim is real but weaker and better-measured than the slide-deck version.

2.6 Tail latency amplification

High percentiles matter most in backend services called multiple times per end-user request.

one end-user request, four parallel backend calls
service A
12 ms
service B
18 ms
service C
96 ms · p99
service D
11 ms

The end user waits for the slowest, even though the calls ran in parallel — so a 1% tail in one service becomes a much larger share of user-visible slow requests.

Figure 2.2.42.6 Tail latency amplification

Even if the calls are parallel, the request waits for the slowest one. One slow call makes the entire end-user request slow. And the more backend calls per end-user request, the higher the probability of hitting at least one slow call — so a larger proportion of end-user requests end up slow than the per-service percentile suggests.

Quick math intuition: if each backend call independently has a 1% chance of exceeding its p99, then with 100 backend calls, 1 − 0.99¹⁰⁰ ≈ 63% of end-user requests hit at least one p99-slow call. Your service's p99 becomes your user's median.

2.7 SLOs and SLAs

  • SLO (service level objective) — a target. Example: median response time < 200 ms, p99 < 1 s, and ≥99.9% of valid requests return non-error responses.
  • SLA (service level agreement) — a contract specifying what happens if the SLO is not met (e.g., customers entitled to a refund).

The book notes plainly: defining good availability metrics for SLOs/SLAs is not straightforward in practice.

2.8 Computing percentiles efficiently

Need: a rolling window (e.g. last 10 minutes), recomputed every minute for a dashboard.

  • Simplest: keep all response times in the window and sort. Works; often too expensive.
  • Approximation libraries at minimal CPU/memory cost: HdrHistogram, t-digest, OpenHistogram, DDSketch.

⚠️ Averaging percentiles is mathematically meaningless. Not "imprecise" — meaningless. You cannot average p99 across machines, or average p99 over time to downsample a graph. The right way to aggregate response-time data is to add the histograms. Almost every hand-rolled dashboard gets this wrong.


On this page