2.2 Describing Performance
In the case study: posts/s and timeline writes/s are throughput; time to load the home timeline and time until a post reaches followers are response times.
Tail amplification
- 13 ms
- median
- 20 ms
- mean
With 20 calls per request, 18% of user requests touch at least one p99-slow backend call. A 1% tail in a single service becomes a 18% problem for users — and the mean (20 ms) reports none of it. This is why the book insists on percentiles of the end-to-end request, measured at the client.
Two metric families:
- Response time — elapsed time from the user making a request to receiving the answer. Unit: seconds/ms/µs.
- Throughput — requests per second, or data volume per second, being processed. For a given hardware allocation there is a maximum throughput. Unit: "somethings per second."
In the case study: posts/s and timeline writes/s are throughput; time to load the home timeline and time until a post reaches followers are response times.
2.1 The relationship: queueing
Why: when a request arrives at a highly loaded system, the CPU is likely already handling an earlier request, so the new one waits. As throughput approaches hardware maximum, queueing delays increase sharply.
Which metric matters to whom: response time is what users care about most. Throughput determines the computing resources required (how many servers) and therefore the cost of serving the workload. A system is scalable if its maximum throughput can be significantly increased by adding computing resources.
2.2 Metastable failure — when an overloaded system won't recover
The vicious cycle:
Key property: even when the original load is removed, the system may stay overloaded until it is rebooted or reset. That's what makes it metastable — the failure state is self-sustaining. This causes serious production outages.
Countermeasures — memorize this table, it's the most operationally reusable content in the chapter:
| Side | Technique | What it does |
|---|---|---|
| Client | Exponential backoff | increase and randomize the time between successive retries (randomization = jitter, to break synchronized retry waves) |
| Client | Circuit breaker | temporarily stop sending requests to a service that recently errored or timed out |
| Client | Token bucket | rate-limit retries to a fixed budget |
| Server | Load shedding | detect approaching overload and proactively reject requests |
| Server | Backpressure | send responses asking clients to slow down |
| Both | Queueing & load-balancing algorithm choice | can materially change behavior under saturation |
2.3 Latency vs response time — precise definitions
- network latency (out)
- queueing delay — waiting for a CPU
- service time — actively processing
- network latency (back)
| Term | Meaning |
|---|---|
| Response time | What the client sees. Includes all delays incurred anywhere in the system. |
| Service time | Duration the service is actively processing the request. |
| Queueing delay | Waiting, at several possible points: for a CPU to become available; for an outbound network buffer when other tasks on the machine are sending a lot of data. |
| Latency | Catchall for time when the request is not being actively processed — i.e., it is latent. |
| Network latency / delay | Time the request and response spend traveling through the network. |
Sources of random per-request delay: a context switch to a background process; a lost network packet and TCP retransmission; a garbage collection pause; a page fault forcing a disk read; mechanical vibrations in the server rack.
Head-of-line blocking. A server processes only a small number of things in parallel (bounded by CPU cores). It takes only a small number of slow requests to hold up all subsequent ones. Those later requests may have fast service times but the client still sees a slow response.
Therefore: queueing delay is not part of service time, and this is exactly why you must measure response times on the CLIENT side. Server-side timers systematically under-report the thing users actually experience.
2.4 Average, median, percentiles
Response time is a distribution, not a number. Variation in network delay is called jitter.
- Mean (arithmetic average) — useful for estimating throughput limits. Bad for "typical" response time: it doesn't tell you how many users actually experienced that delay.
- Median = p50 — sort fastest→slowest, take the halfway point. p50 = 200 ms means half of requests are faster, half slower. Good metric for "how long users typically wait."
- p95 / p99 / p999 — how bad the outliers are. p95 = 1.5 s means 5 out of 100 requests take ≥ 1.5 s.
Tail latencies (high percentiles) directly affect user experience.
The Amazon argument (important and non-obvious): Amazon specifies internal service response times at the 99.9th percentile, even though it affects only 1 in 1,000 requests — because the customers with the slowest requests are usually those with the most data on their accounts, i.e., the ones who bought the most, i.e., the most valuable customers. Conversely Amazon judged optimizing p99.99 (1 in 10,000) too expensive for insufficient benefit — very high percentiles are easily affected by random events outside your control and returns diminish.
2.5 What the latency-vs-revenue data actually says
The book is unusually careful here, and it's worth copying the skepticism:
| Study | Claim | Status |
|---|---|---|
| Google 2006 | 400 ms → 900 ms slowdown ≈ 20% drop in traffic and revenue | Often cited, unreliable |
| Google 2009 | +400 ms latency → only 0.6% fewer searches/day | Contradicts the above |
| Bing 2009 | +2 s load time → 4.3% less ad revenue | — |
| Akamai (recent) | +100 ms → up to 7% lower ecommerce conversion | Same study shows very fast page loads also correlate with LOWER conversion — because the fastest pages are often ones with no useful content (404s). No attempt to separate content effects from load-time effects → probably not meaningful |
| Yahoo (following year) | 20–30% more clicks on fast searches when the fast/slow difference is ≥1.25 s | Controlled for search-result quality — the most trustworthy of these |
Lesson: be suspicious of the folklore numbers; the causal claim is real but weaker and better-measured than the slide-deck version.
2.6 Tail latency amplification
High percentiles matter most in backend services called multiple times per end-user request.
The end user waits for the slowest, even though the calls ran in parallel — so a 1% tail in one service becomes a much larger share of user-visible slow requests.
Even if the calls are parallel, the request waits for the slowest one. One slow call makes the entire end-user request slow. And the more backend calls per end-user request, the higher the probability of hitting at least one slow call — so a larger proportion of end-user requests end up slow than the per-service percentile suggests.
Quick math intuition: if each backend call independently has a 1% chance of exceeding its p99, then with 100 backend calls, 1 − 0.99¹⁰⁰ ≈ 63% of end-user requests hit at least one p99-slow call. Your service's p99 becomes your user's median.
2.7 SLOs and SLAs
- SLO (service level objective) — a target. Example: median response time < 200 ms, p99 < 1 s, and ≥99.9% of valid requests return non-error responses.
- SLA (service level agreement) — a contract specifying what happens if the SLO is not met (e.g., customers entitled to a refund).
The book notes plainly: defining good availability metrics for SLOs/SLAs is not straightforward in practice.
2.8 Computing percentiles efficiently
Need: a rolling window (e.g. last 10 minutes), recomputed every minute for a dashboard.
- Simplest: keep all response times in the window and sort. Works; often too expensive.
- Approximation libraries at minimal CPU/memory cost: HdrHistogram, t-digest, OpenHistogram, DDSketch.
⚠️ Averaging percentiles is mathematically meaningless. Not "imprecise" — meaningless. You cannot average p99 across machines, or average p99 over time to downsample a graph. The right way to aggregate response-time data is to add the histograms. Almost every hand-rolled dashboard gets this wrong.