Learn Labs
13. A Philosophy of Streaming Systems

13.9 Worked examples

This is §1.5's "asynchrony is what makes systems based on event logs robust," as a number.

① Why asynchrony contains faults, quantified. A write must reach 5 systems, each with 99.9% availability.

  • Distributed transaction (all must commit): availability = 0.999⁵ = 99.5% → ~44 hours of downtime/year, and any one system's outage stops all writes.
  • Log-based derived (write to the log only): the write path depends on 1 system = 99.9%; a failing consumer only makes its own view stale, and it catches up when fixed. This is §1.5's "asynchrony is what makes systems based on event logs robust," as a number.

② The apology calculus. An airline with 200 seats. Strict constraint: 0 overbookings, but ~8% no-show rate → 16 empty seats per flight, ~$4,800 of lost revenue. Overbooking by 5%: ~10 extra seats sold (+$3,000), and the probability that all 210 show up is small; when it happens, compensation costs ~$800 per bumped passenger. Expected value strongly favours the loose constraint — and, critically, the compensation process must exist anyway for weather cancellations. This is why §5.4's argument isn't a hack; it's how the business already works.

③ The four-layer duplicate trace. Probability a $11 transfer becomes $22:

  • P(client-side timeout after commit) ≈ 0.1% of requests on a poor mobile connection
  • P(user retries | error shown) ≈ 60% ⇒ ~0.06% of transfers double-charge — 6 in 10,000. At 100,000 transfers/day that's 60 incorrect transfers per day, every one of them a customer complaint. A single request-ID column removes all of them.

④ Where the write/read boundary goes. 10M documents, 1,000 distinct common queries, 50,000 queries/s of which 80% are the common ones.

  • No index: read cost = scan 10M docs × 50,000/s. Impossible.
  • Index only: write cost = update terms per document; read cost = per-query Boolean evaluation × 50,000/s.
  • Index + cache of the 1,000 common queries: 40,000 q/s served from cache at near-zero cost; 10,000 q/s hit the index. Write cost rises by 1,000 materialized-view updates per relevant document change. The right answer depends entirely on write:read ratio — and the celebrity insight is that within one system, different keys may deserve different answers.

⑤ Integrity vs timeliness, priced. A bank's ledger.

  • Timeliness violation: a transaction doesn't appear for 24 h. Cost: a support call. Self-healing.
  • Integrity violation: debits ≠ credits by $1. Cost: a full audit, regulatory exposure, and manual reconstruction — and it does not self-heal. Even a 1-in-10⁶ integrity failure at 10M transactions/day is 10 per day, each requiring human investigation. This is why the reconciliation job is not optional.

⑥ Auditing coverage. 500 TB across 3 replicas; a scrubber reading at 200 MB/s per node. Full pass = 500e12 / (200e6 × 3) ≈ 833,000 s ≈ 9.6 days. So any given block is verified roughly every 10 days; a corruption introduced today is detected in ~5 days on average. If your backup retention is 7 days, you have ~2 days of margin to restore a clean copy. Lengthen retention or speed up scrubbing — and know which you're relying on.