# Glossary (/docs/glossary)
# Start reading (/docs)
Two books, read closely and written up chapter by chapter. Each chapter follows the
same shape: the argument in order, diagrams redrawn from the figures, technology
deep dives, a production failure catalog, a decision cheat sheet, worked
arithmetic, and a self-test.
## How each chapter is laid out [#how-each-chapter-is-laid-out]
| Section | What's in it |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Concepts** | The chapter's argument, in order, with key sentences preserved verbatim where the exact wording matters |
| **Diagrams** | Redrawn from the book's figures, plus extra diagrams for mechanisms described only in prose |
| **Technology deep dives** | Per system: what problem it solves · why the alternative wasn't enough · how it works internally · deployment · monitoring · scaling · backup · what breaks in production |
| **Production failure catalog** | Symptom → underlying mechanism, filterable |
| **Decision cheat sheet** | The choices you'll actually have to make, with the criteria |
| **Worked examples** | The arithmetic behind the trade-offs |
| **Self-test** | 20–49 questions per chapter, ending in an open design question, with progress tracked |
| **Terminology · Forward links** | Chapter vocabulary, and where each idea gets developed next |
The books supply the concepts and the failure *modes*; the
deployment/monitoring/backup material is the practitioner layer they deliberately
leave to you.
# ClickHouse (/docs/clickhouse)
# Designing Data-Intensive Applications (/docs/ddia)
# Kafka: The Definitive Guide (/docs/kafka)
# PostgreSQL (/docs/postgres)
# Data Types (/docs/clickhouse/data-types/data-types)
ClickHouse's type system looks superficially like any SQL database's,
but several choices here have direct, measurable effects on storage
size and query speed — because every type decision is also a
column-storage decision.
## Numeric types: pick the narrowest one that fits [#numeric-types-pick-the-narrowest-one-that-fits]
ClickHouse gives you explicit-width integers — `UInt8`, `UInt16`,
`UInt32`, `UInt64`, up to `UInt256`, plus signed `Int8..Int256` — and
floating point `Float32`/`Float64`. Unlike Postgres, where `INTEGER`
is the path of least resistance regardless of actual range, picking
the narrowest type that fits your data here directly shrinks every
part on disk and the amount of memory a scan has to move through.
For money or anything needing exact arithmetic, use `Decimal(P, S)`
rather than `Float64` — floats accumulate rounding error under
aggregation, which is exactly the operation ClickHouse spends most of
its time doing.
## String vs FixedString [#string-vs-fixedstring]
`String` is a variable-length byte string with no encoding assumption
(not necessarily UTF-8-validated). `FixedString(N)` stores exactly N
bytes, padding with zero bytes — cheaper only when every value is
genuinely the same length (country codes, fixed-width hashes); using
it for anything else wastes space padding shorter values.
## LowCardinality: dictionary-encode repetitive strings [#lowcardinality-dictionary-encode-repetitive-strings]
`LowCardinality(String)` stores a small dictionary of distinct values
plus an integer index per row, instead of repeating the string itself.
It's the right call for columns like status, country, event type, or
plan tier — anything with a bounded, modestly sized set of distinct
values relative to row count. The compression mechanics of the
underlying dictionary are covered in
[Compression](/clickhouse/compression); the thing to internalize here
is *when* to reach for it: low distinct-value count relative to total
rows, used often in `GROUP BY`, `WHERE`, or joins.
## Nullable(T): correct, but not free [#nullablet-correct-but-not-free]
Wrapping a type in `Nullable(...)` doesn't just add a flag to the
existing column — it adds an entirely separate hidden column: a
bitmap recording, per row, whether the value is null. Every read of
that column now involves reading two columns instead of one, and
several optimizations (certain codecs, some vectorized operations)
can't apply as cleanly through the extra indirection.
{/* Not a chain: the logical column is physically two separate
files, and a read has to touch both. A straight Pipeline chain
would wrongly imply the values file produces the null-mask.
Grid columns are 1fr/1fr, so both branches are exactly equal
width and the 25%/75% connector marks land dead-center on each
box regardless of viewport. */}
Prefer a sentinel default (`0`, empty string, `1970-01-01`) over
`Nullable` where the domain allows it. Reserve `Nullable` for
columns where "unknown" is a meaningfully different state from
"zero/empty" and that distinction actually matters to downstream
queries.
## Date and time precision [#date-and-time-precision]
* `Date` — 2 bytes, day resolution, range up to 2149.
* `Date32` — 4 bytes, day resolution, wider range (including
pre-1970 dates).
* `DateTime` — 4 bytes, second resolution, timezone-aware.
* `DateTime64(precision)` — 8 bytes, configurable sub-second precision
(milli/micro/nanosecond) for event timestamps that need it.
Using `DateTime64(9)` everywhere "just in case" doubles storage over
plain `DateTime` for no benefit if you never actually need nanosecond
precision — match the type to the real precision of your source data.
## Enum8 / Enum16 [#enum8--enum16]
An `Enum` stores a single byte (Enum8) or two bytes (Enum16) on disk
while still comparing, sorting, and displaying as the named string —
similar end result to `LowCardinality(String)` for a fixed,
known-in-advance set of values, with the schema itself documenting the
valid values.
# Data Types (/docs/clickhouse/data-types)
Choosing the right shape for your data, including nested structures.
# Arrays, Maps & JSON (/docs/clickhouse/data-types/semi-structured-data)
ClickHouse isn't limited to flat scalar columns. Arrays, tuples, maps,
nested structures, and a dynamic JSON type all exist — each stored in
a way that stays consistent with the column-oriented model rather than
falling back to an opaque blob.
## Array(T): two parallel columns, not a serialized blob [#arrayt-two-parallel-columns-not-a-serialized-blob]
An `Array(T)` column is physically stored as two columns: a flat
**values** array holding every element from every row concatenated
together, and an **offsets** array recording where each row's slice
ends. There is no per-row serialization/deserialization step — reading
the array for row N is just reading a range out of the shared values
array using the offsets.
## Tuple(T1, T2, ...): a fixed-shape group of columns [#tuplet1-t2--a-fixed-shape-group-of-columns]
A `Tuple` groups a fixed number of typed elements — each element is
effectively its own sub-column internally, so
`Tuple(Float64, Float64)` for a coordinate pair stores two
independent, independently compressible columns rather than one
combined structure.
## Map(K, V): syntax sugar over two arrays, not a hash table [#mapk-v-syntax-sugar-over-two-arrays-not-a-hash-table]
`Map(K, V)` is implemented as an `Array(Tuple(K, V))` under the hood
— key lookup with `tags['env']` is a linear scan over that row's
entries, not an O(1) hash lookup. It's a convenient way to model
open-ended key/value attributes on a row, but it is not a substitute
for a real hash index, and it isn't a good fit for maps with
hundreds of entries per row that need fast random access.
## Nested(...): parallel arrays that stay in sync [#nested-parallel-arrays-that-stay-in-sync]
A `Nested` column is syntactic sugar for a set of `Array` columns that
all share the same length per row — useful for modeling a one-to-many
relationship (like line items on an order) without a separate joined
table.
## ARRAY JOIN: turning array elements into rows [#array-join-turning-array-elements-into-rows]
To analyze array elements individually rather than as a group,
`ARRAY JOIN` expands each element into its own row, duplicating the
rest of the row's columns alongside it — useful for computing
per-page-view stats out of the `sessions` table above:
## JSON: a real column type, not just a String [#json-a-real-column-type-not-just-a-string]
The `JSON` type stores semi-structured data as a set of dynamically
discovered subcolumns rather than one opaque text blob — paths that
appear consistently across rows get their own typed, columnar storage
internally, so querying `payload.user.id` doesn't require re-parsing a
string on every read the way `String` + `JSONExtract` would. Reach for
it when incoming event shapes vary and you don't want to pre-define
every field, and reach for a proper typed schema (plain columns, or
`Nested`) whenever the shape is actually known ahead of time — a fixed
schema is still faster and more predictable to query.
# Foundations (/docs/clickhouse/foundations)
Get it running, write your first queries, and understand how a MergeTree table stores what you write.
# MergeTree Engine (/docs/clickhouse/foundations/mergetree)
MergeTree is not a feature of ClickHouse — it's the table engine
almost everything else in this module sits on top of. Partitions,
primary keys, TTL, replication, skip indexes and projections are all
properties *of a MergeTree table*, not separate systems.
## Why does it exist? [#why-does-it-exist]
Row-oriented databases like Postgres are built around fast,
transactional access to individual rows. ClickHouse is built for the
opposite workload: scanning and aggregating billions of rows for
analytical queries, where a single slow write in exchange for very
fast reads is a good trade. MergeTree is the storage engine that makes
that trade deliberately.
## Where the name comes from [#where-the-name-comes-from]
Every `INSERT` doesn't modify existing data — it writes a brand new,
immutable **part**: a small self-contained directory of sorted,
columnar files. A background process continuously **merges** smaller
parts into larger ones. "MergeTree" is literally: a tree of parts,
merged over time.
[Parts & Background Merges](/clickhouse/parts-merges) goes through
that loop in detail, later in the Storage Engine section. This page is
about the engine family sitting around it.
## Declaring a MergeTree table [#declaring-a-mergetree-table]
`ORDER BY` is not cosmetic — it defines the physical sort order every
part is written in, and doubles as the primary key unless you specify
one separately. Get this wrong and every later optimization (sparse
index, skip indexes, projections) inherits the mistake.
## The MergeTree family [#the-mergetree-family]
Plain `MergeTree` keeps every row you insert. Several variants change
what happens *during a merge*, without changing anything else about
how parts are stored or scanned:
* **ReplacingMergeTree** — during a merge, keeps only the newest row
per sort key. Useful for deduplicating upserts, but duplicates can
still exist between merges.
* **SummingMergeTree** — during a merge, sums numeric columns that
share the same sort key. Good for pre-aggregating metrics without a
separate rollup job.
* **AggregatingMergeTree** — merges rows holding partial aggregate
states (from functions like `avgState`), used almost exclusively as
the target of a materialized view.
* **CollapsingMergeTree / VersionedCollapsingMergeTree** — a "sign"
column marks rows for cancellation, letting you emulate
deletes/updates in an append-only engine.
* **ReplicatedMergeTree** — any of the above, plus coordinated
replication via Keeper. Covered in
[Replication & Keeper](/clickhouse/replication).
Same mechanism, five different outcomes — pick an engine below to see
what actually happens when its parts merge, and a realistic table that
would use it:
## When to use which [#when-to-use-which]
Start from what your query needs to be true, not from the feature list
— most of these have a narrower "reach for something else instead"
case than their name suggests:
Engine
Use it when
Reach for something else when
MergeTree
Every row is independently meaningful — raw events, logs,
immutable facts you never need to collapse.
You need dedup, rollups, or update/delete semantics — any
variant below.
ReplacingMergeTree
You only care about the latest row per key — CDC sync,
upserts, slowly-changing dimensions.
You need an accurate row count or full history right now —
duplicates aren't actually gone until a merge runs;{" "}
FINAL or argMax is required for
correctness before then.
SummingMergeTree
You're pre-aggregating a running numeric total per key —
counters, revenue, event counts.
The aggregate isn't a plain sum (avg, uniq, quantile) —
AggregatingMergeTree; or you need per-row detail — plain
MergeTree.
AggregatingMergeTree
A materialized view needs incremental, non-additive
aggregates — uniq, avg,{" "}
quantile, any \*State function.
A plain sum would do — SummingMergeTree is simpler and
doesn't need a \*Merge combinator at query time.
CollapsingMergeTree
Emulating UPDATE/DELETE via a sign column, and one writer
already guarantees the cancel-then-insert pair arrives in
order.
Multiple writers/partitions can't guarantee that order —
VersionedCollapsingMergeTree; or "latest wins" is all you
need — ReplacingMergeTree is simpler.
VersionedCollapsingMergeTree
Same cancel-and-replace pattern, but rows can arrive out of
order — multiple Kafka partitions, multiple producers.
A single writer already guarantees order — plain
CollapsingMergeTree needs one fewer column.
ReplicatedMergeTree (any of the above)
Any of the above, running on more than one node in
production.
Local single-node dev/testing — the plain variant is one less
moving part (no Keeper dependency).
These variants only resolve duplicates/sums **when parts merge**,
and merges are not scheduled by you. Query results can show
duplicate or unsummed rows until a merge happens. If a query needs a
guaranteed-correct view right now, add `FINAL` or aggregate
explicitly — don't rely on merge timing.
## What you get by choosing MergeTree [#what-you-get-by-choosing-mergetree]
* Data physically sorted by your chosen key, enabling range scans.
* Columnar storage per part — queries only read the columns they
touch.
* Immutable parts — safe concurrent reads, no locking on writes.
* A background process that continuously reorganizes storage for you.
# ORDER BY & Primary Keys (/docs/clickhouse/foundations/primary-key)
`ORDER BY` is the single most consequential decision in a MergeTree
schema. It defines the physical sort order of every row on disk, and —
unless you say otherwise — doubles as the table's primary key.
## Not a B-Tree — a sparse index [#not-a-b-tree--a-sparse-index]
Postgres's B-Tree index has one entry per row. ClickHouse tables can
hold tens of billions of rows, so indexing every single one would make
the index itself enormous. Instead, rows are grouped into
**granules** — 8192 rows by default — and the primary index stores
exactly one entry per granule: the primary key value of its *first*
row. That's the "sparse" part.
## How a lookup actually works [#how-a-lookup-actually-works]
For `WHERE user_id = 204`, ClickHouse doesn't scan rows — it
binary-searches the small in-memory array of marks to find which
granule *could* contain `204`, then reads and decompresses only that
granule's columns (and its immediate neighbors, since the target may
straddle a boundary). On a real table that's the difference between
reading three rows and reading nine — at production scale, between
reading megabytes and reading gigabytes.
## Column order in ORDER BY matters [#column-order-in-order-by-matters]
Rows are sorted by the *first* column, then the second breaks ties
within it, and so on — exactly like a compound index. A filter on the
first column can binary-search straight to the right granules. A
filter on only the second column gets no such benefit, because rows
with a matching `event_time` are scattered across every `user_id`
group.
* Put the column your queries filter on **most often** first.
* Prefer lower-cardinality columns before higher-cardinality ones when
both are filtered together — it keeps runs of identical values
longer, which compresses and skips better.
* The primary key doesn't need to be unique — unlike Postgres,
duplicate keys are completely normal.
Changing `ORDER BY` after a table already has data isn't a metadata
change — existing parts were physically written in the old order.
In practice this means creating a new table with the desired order
and re-inserting, not `ALTER TABLE`.
`index_granularity` (default 8192) controls granule size. Smaller
granules mean a more precise index and less wasted scanning, at the
cost of a larger index held in memory — rarely worth tuning unless
you have unusually wide or narrow rows.
# Docker & Setup (/docs/clickhouse/foundations/setup)
Everything in this lab runs in Docker first. Before ClickHouse
internals mean anything, it helps to be precise about what actually
happens between typing `docker compose up` and having a server
listening on your machine.
## From image to running server [#from-image-to-running-server]
A Docker **image** is a read-only template — a frozen filesystem plus
metadata about how to run it. A **container** is a running process
created from that image, with its own writable layer on top. A
**volume** is storage that lives outside the container's lifecycle, so
data survives when the container is removed and recreated.
{/* Same fixed 3-column, S-shaped grid as every other 4-step chain
in this module (compression, mutations, mergetree, ...): two
rows of two, connected top-right to bottom-right to bottom-left.
Every width is intrinsic to its content, so there's no runtime
measurement that can drift out of sync with the alignment. */}
The image is pulled once and cached locally. Every
`docker compose up` after that reuses it — only the container is
created and destroyed. The volume is what makes this safe:
`docker compose down` removes the container but leaves
`clickhouse_data` untouched; only `down -v` deletes the volume, and
with it, your data.
## The compose file [#the-compose-file]
This repo's `clickhouse/compose.yaml` is intentionally minimal — one
service, two ports, one named volume:
Port mapping reads as `host:container`. The container always listens
on 8123 and 9000 internally; the left-hand side is just where that
gets exposed on your machine. Two protocols, two ports:
* `8123` — HTTP interface. Used by the web UI, most client libraries,
and simple `curl` queries.
* `9000` — native TCP protocol. Used by `clickhouse-client` and
drivers that want the faster binary protocol.
## Environment & credentials [#environment--credentials]
`.env` supplies the database name, user, and password ClickHouse
bootstraps on first start — first start only. Changing these after the
volume already has data does nothing until you wipe the volume,
because the users and database were already created on disk.
## Everyday commands [#everyday-commands]
`down -v` is the one command here that is not reversible — it
deletes the named volume, and every table in it, permanently. If you
only want to restart clean without losing data, use `down` (no
`-v`) or just `stop`.
## Verifying it works [#verifying-it-works]
Once the container is up, connect from your host on either port and
run a few sanity queries:
`numbers(10)` is ClickHouse's built-in table function for generating
rows on the fly — useful throughout this module for quick experiments
without needing real data first.
# SQL Basics (/docs/clickhouse/foundations/sql-basics)
ClickHouse's query language is SQL, and if you've used Postgres or
MySQL, most of what follows will look familiar. This page is the
minimum vocabulary — databases, tables, inserting rows, selecting rows
— plus the handful of places where ClickHouse's dialect quietly
disagrees with what you'd expect from an OLTP database. Everything
after this page assumes you're comfortable with what's here.
## Databases and tables [#databases-and-tables]
A ClickHouse server holds multiple **databases**, each holding
multiple **tables** — same nesting as Postgres or MySQL.
The `.env` file from [Docker & Setup](/clickhouse/setup) already
creates a `learning` database for you on first start, so in practice
you'll mostly skip straight to creating tables inside it.
## Creating a table [#creating-a-table]
Every `CREATE TABLE` needs a column list, a `ENGINE`, and — for the
engine family this whole module is built around — an `ORDER BY`:
`UInt64`, `String`, and `DateTime` are three of ClickHouse's basic
types — the full type system (signed/unsigned widths,
`LowCardinality`, `Nullable`, arrays) is its own page:
[Data Types](/clickhouse/data-types). What `ENGINE` and `ORDER BY`
actually do is covered next, in
[MergeTree Engine](/clickhouse/mergetree) and
[ORDER BY & Primary Keys](/clickhouse/primary-key).
## Inserting rows [#inserting-rows]
`VALUES` is the simplest form. Real pipelines almost never use it
row-by-row like this — see
[Insert & Export Formats](/clickhouse/insert-formats) for the formats
you'd actually use, and
[Async Inserts & Batching](/clickhouse/async-inserts) for why
single-row inserts specifically are a trap.
## Selecting rows [#selecting-rows]
Filtering, sorting, and limiting all read the way they do everywhere
else:
## Aggregating [#aggregating]
`GROUP BY` plus an aggregate function is the query shape ClickHouse is
actually built to make fast over billions of rows — this is the
workload described in [Why ClickHouse](/clickhouse/why-clickhouse):
`count()` with no argument and no `DISTINCT` counts rows in the group;
ClickHouse also accepts the standard `count(*)` spelling. `sum()`,
`avg()`, `min()`, and `max()` all behave as expected. Far more
specialized aggregate functions exist — that's its own page:
[Aggregate Combinators](/clickhouse/aggregate-combinators).
## Where the dialect disagrees with what you'd expect [#where-the-dialect-disagrees-with-what-youd-expect]
* **No auto-increment primary key.** There's no `SERIAL` /
`AUTO_INCREMENT`. IDs are generated by the application, or with a
function like `generateUUIDv4()`, before the row is inserted.
* **Columns aren't nullable by default.** A plain `String` column can
never hold `NULL` — you have to opt in with `Nullable(String)`, and
doing so has a real storage and performance cost. Details in
[Data Types](/clickhouse/data-types).
* **`UPDATE` and `DELETE` aren't the lightweight statements you're
used to.** They exist as `ALTER TABLE ... UPDATE/DELETE`
*mutations* — async, heavyweight, rewrite-the-part operations, not
row-level edits. Covered in
[Mutations (UPDATE / DELETE)](/clickhouse/mutations).
* **No foreign keys, no multi-table transactions.** Joins work
([Joins](/clickhouse/joins)), but nothing enforces referential
integrity, and there's no `BEGIN/COMMIT/ROLLBACK` spanning tables.
Because `DELETE` and `UPDATE` are expensive mutations rather than
cheap row edits, reaching for them the way you would in Postgres —
to fix a handful of rows here and there — is a common first mistake.
If you find yourself doing it often, that usually means the schema
or the pipeline needs to change, not that you need a faster
mutation.
That's the whole vocabulary this module assumes going forward. Next
up: [MergeTree Engine](/clickhouse/mergetree) — the table engine every
`CREATE TABLE` above was quietly already using.
# Async Inserts & Batching (/docs/clickhouse/ingestion-other-engines/async-inserts)
[Parts & Background Merges](/clickhouse/parts-merges) already covers
why small, frequent inserts are dangerous — each one creates a new
part, and enough of them outrun the background merge process.
Batching client-side is the ideal fix. This page is for when you
don't control the client: many independent application servers, each
with a handful of rows to insert, with no good place to accumulate a
batch before sending it.
## Option one: the Buffer engine [#option-one-the-buffer-engine]
A `Buffer` table sits in front of a real MergeTree table. Inserts
land in memory, and the buffer flushes into the underlying table once
row-count, byte-size, or time thresholds are crossed — turning many
tiny inserts into far fewer, larger ones before they ever become a
part.
It's the older mechanism, still used, but with a real cost: the
buffer lives in server memory and is lost on restart or crash —
anything not yet flushed is gone.
## Option two: async\_insert (the modern default choice) [#option-two-async_insert-the-modern-default-choice]
Rather than a separate table, `async_insert` is a setting that
changes how `INSERT` itself behaves: the server accepts the (possibly
tiny) insert immediately, holds it in an internal buffer alongside
inserts from other connections, and flushes the accumulated buffer as
one real part once a size or time threshold is hit.
`wait_for_async_insert` is the trade you're actually making explicit:
* **= 1 (default)** — the client's `INSERT` blocks until the buffered
data is actually flushed to disk. Safer, but the client still waits
roughly as long as it would have without async insert.
* **= 0** — the server acknowledges the insert as soon as it's in the
in-memory buffer, before it's durable. Much lower client-perceived
latency, at the cost of a real window where an acknowledged insert
can be lost if the server crashes before flushing.
The real difference from client-side batching: `async_insert` lets
ClickHouse do the batching itself, across many separate connections
it doesn't control the contents of. Client-side batching is still
strictly better when it's possible — it costs nothing in durability
— but `async_insert` is the right tool when hundreds of independent
services are each inserting a few rows at a time and can't
coordinate a shared batch.
`async_insert` buffers in server memory, same as `Buffer` tables —
high insert concurrency with `wait_for_async_insert = 0` under a
server crash can lose the most recently buffered, unflushed rows.
Choose it deliberately for high-volume, loss-tolerant telemetry, not
for data you cannot afford to lose a few seconds of.
# Ingestion & Other Engines (/docs/clickhouse/ingestion-other-engines)
Getting data in, and the engines beyond MergeTree.
# Insert & Export Formats (/docs/clickhouse/ingestion-other-engines/insert-formats)
Most databases speak one wire format and one bulk-load format. Every
query in and out of ClickHouse — `INSERT`, `SELECT`, even the results
printed to your terminal — passes through a pluggable **format**, and
ClickHouse supports dozens of them. Picking the right one for the
right job is a real performance decision, not just a convenience.
## The formats you actually reach for [#the-formats-you-actually-reach-for]
* **JSONEachRow** — one JSON object per line, no surrounding array or
commas. The default choice for application-side inserts, because
almost every language can produce it without a ClickHouse-specific
library.
* **CSV** / **TSV** — plain delimited text. Universally producible,
but the slowest to parse and the worst at preserving types
(everything round-trips through text).
* **Native** — ClickHouse's own binary, columnar wire format. This is
what `clickhouse-client` and server-to-server replication use by
default, and it's the fastest option because data is already shaped
the way the engine stores it — no row-to-column transposition needed
on the way in.
* **RowBinary** — binary, but row-oriented: values in declaration
order, no field names, no delimiters. Faster than CSV to parse, more
compact, but you must get column order exactly right since there's
nothing self-describing about it.
* **Parquet** / **Arrow** — columnar formats from the wider data-lake
ecosystem. ClickHouse reads and writes both natively, which makes it
easy to sit next to Spark, Pandas, or a data lake on S3 without a
separate conversion step.
## Using a format [#using-a-format]
Every format is invoked the same way: a `FORMAT` clause on the query.
The same thing over the HTTP interface (port 8123, covered in
[Docker & Setup](/clickhouse/setup)) is just a POST body:
## Reading files directly, without an INSERT [#reading-files-directly-without-an-insert]
Table functions like `file()`, `s3()`, and `url()` let a format double
as a way to query external data as if it were a table, no loading step
required:
These table functions are the query-time counterpart to the
[integration engines](/clickhouse/integration-engines) (S3, MySQL,
PostgreSQL, URL, File) — a table function reads once for a single
query; an engine defines a permanent table you can query repeatedly.
For bulk historical loads, prefer `Native` or `Parquet` over `CSV`
or `JSONEachRow` whenever the source can produce them — parsing
text and inferring/validating types row by row is frequently the
actual bottleneck in a large one-time load, not disk or network.
`RowBinary` has no field names and no self-description — if your
`INSERT` column list and the binary layout don't match exactly, you
silently get garbage values in the wrong columns rather than an
error. Reserve it for pipelines where both ends are code you
control.
# Integration Engines (/docs/clickhouse/ingestion-other-engines/integration-engines)
Integration engines define a table backed by data that lives
*somewhere else* — S3, another database, an HTTP endpoint, a local
file — so you can query it with ClickHouse SQL without an ETL step to
copy it in first. Every one of them shares the same formats covered
in [Insert & Export Formats](/clickhouse/insert-formats) for how the
underlying bytes are read or written.
## S3 [#s3]
Reads and writes objects (Parquet, CSV, JSONEachRow, etc.) directly
in an S3 bucket, either as a permanent table or, more commonly, via
the `s3()` table function for one-off queries against data that
already lives in a lake:
## MySQL / PostgreSQL [#mysql--postgresql]
Proxies queries to a live external database — each `SELECT` against
the ClickHouse table is translated and forwarded to the real MySQL or
Postgres server, row by row over the network.
This is convenient for occasional lookups against operational data
without duplicating it, but it inherits the source database's
performance characteristics for every query — running a billion-row
analytical scan through a `PostgreSQL` engine table just turns it
into a billion-row query against Postgres. For anything queried often
or at volume, prefer a [dictionary](/clickhouse/dictionaries) (for
small reference data) or an actual ingestion pipeline into a real
MergeTree table.
## URL [#url]
Reads from or writes to an arbitrary HTTP(S) endpoint, using the same
format machinery as everything else — useful for pulling a one-off
feed or webhook payload into a query without writing a separate
script.
## File [#file]
Reads local files on the server's filesystem. Mostly a local
development and testing convenience — production data almost never
lives as loose files on a ClickHouse server's disk on purpose.
The common thread: integration engines trade query performance for
skipping a separate copy/ETL step. That trade is worth it for
exploration, bridging, and infrequent lookups. It stops being worth
it the moment a query pattern becomes frequent or
performance-sensitive — at that point, ingest the data into a real
MergeTree table (or a dictionary, for small reference data) instead
of querying the external system live, every time.
A `MySQL` or `PostgreSQL` engine table sends real load to that
external database on every query. Treat it the same as any other
client of that database — rate limits, connection pool exhaustion,
and slow queries there are just as real a production risk as they
would be from any other application.
# Kafka Engine (/docs/clickhouse/ingestion-other-engines/kafka-engine)
Streaming events from Kafka into ClickHouse is one of the most common
production patterns, and it's built from two familiar ideas: a table
engine that stores no data itself (the same shape as
[Distributed](/clickhouse/distributed-tables), covered later) and a
[materialized view](/clickhouse/materialized-views) that reacts to
inserts. Kafka wiring just connects them to a topic instead of a
client.
## The three-part shape [#the-three-part-shape]
## Why not just query the Kafka table directly? [#why-not-just-query-the-kafka-table-directly]
You can — `SELECT` against a `Kafka` engine table actually consumes
messages from the topic as a side effect, advancing the consumer
offset. That makes it fundamentally unlike a normal table: reading it
twice does not give you the same rows twice, and two people querying
it concurrently split the messages between them rather than both
seeing everything. In practice it should only ever be read by exactly
one thing — the materialized view attached to it — which is why the
standard advice is to never query it directly yourself.
## What the materialized view is really doing here [#what-the-materialized-view-is-really-doing-here]
ClickHouse's Kafka integration polls the topic in the background and,
for each batch of messages it reads, inserts them into the `Kafka`
engine table — which immediately fires every materialized view
attached to it, exactly like any other insert. The MV's `SELECT` can
reshape, filter, or aggregate the raw message on the way into the
target table, not just copy it verbatim.
`kafka_group_name` is a real Kafka consumer group — offset
tracking, rebalancing, and at-least-once delivery semantics all
follow standard Kafka consumer behavior. ClickHouse is a consumer
like any other; it does not change how Kafka itself guarantees
delivery.
A crash between Kafka delivering a batch and the materialized
view's insert completing can result in messages being re-delivered
on restart — this pattern gives you at-least-once delivery into
ClickHouse, not exactly-once. If exact deduplication matters, pair
it with [ReplacingMergeTree](/clickhouse/mergetree) keyed on a
unique message ID, or dedupe explicitly downstream.
# Other Table Engines (/docs/clickhouse/ingestion-other-engines/table-engines-overview)
Almost everything in this module assumes
[MergeTree](/clickhouse/mergetree) — and for real analytical tables,
that assumption is correct. But ClickHouse ships several other table
engines for narrower jobs, and knowing they exist saves you from
reinventing them with a MergeTree table that doesn't need any of
MergeTree's machinery.
## Log family — simple, single-writer, no index [#log-family--simple-single-writer-no-index]
`TinyLog`, `Log`, and `StripeLog` store columns with no sorting, no
primary key, and no support for concurrent writes. In exchange, they
have almost no overhead — no merges, no index to maintain. They fit
small, rarely-written tables: staging data mid-pipeline, lookup tables
loaded once at startup, temporary scratch tables in a script.
## Memory — no disk at all [#memory--no-disk-at-all]
`Memory` tables live entirely in RAM and are wiped on restart. Useful
for genuinely temporary data within a session or a script — never for
anything that needs to survive a restart, by design.
## Null — a table that discards everything [#null--a-table-that-discards-everything]
`Null` accepts any `INSERT` and keeps nothing. That sounds useless
until you pair it with a
[materialized view](/clickhouse/materialized-views): the view's
trigger still fires on every row inserted into the `Null` table, so
you get the transform-and-store behavior of the view without ever
paying to store the raw input itself.
## Merge — a query-time union, not a materialized copy [#merge--a-query-time-union-not-a-materialized-copy]
`Merge` defines a virtual table over every existing table whose name
matches a regular expression — typically a family of tables sharded
by month or by source. Querying it is a live `UNION ALL` across the
matching tables; nothing is precomputed or stored, which is the key
difference from a materialized view.
## View — a named query, nothing more [#view--a-named-query-nothing-more]
`View` stores a `SELECT` statement under a name; querying it re-runs
the underlying query every time, identically to just writing that
`SELECT` out by hand. It holds no data of its own.
"View" and "Materialized View" share a name but do opposite things:
a `View` is pure syntactic sugar, computed at query time, zero
storage cost, always fresh. A
[Materialized View](/clickhouse/materialized-views) is an insert
trigger that precomputes and physically stores results — real
storage cost, real staleness bounded only by insert timing, much
faster to read. Confusing which one you created is a common and
costly mistake.
None of these engines replicate, shard, or maintain a primary key
the way MergeTree does. Reach for them only when the job is
genuinely small, temporary, or purely structural (routing/union) —
the moment a table needs to hold real analytical data at scale, it
belongs on a MergeTree variant.
# Production Failure Scenarios (/docs/clickhouse/monitoring-operations/failure-scenarios)
Every concept in this module has a corresponding way to misuse it.
This page collects the failures that actually happen in production
ClickHouse deployments, in symptom → cause → fix form, so they're
recognizable the first time you hit them instead of the third.
## "Too many parts" [#too-many-parts]
**Symptom:** inserts start failing with an error mentioning too many
parts, often during a traffic spike.
**Cause:** every `INSERT` creates at least one new part (see
[Parts & Background Merges](/clickhouse/parts-merges)). If inserts
arrive in small, frequent batches — or one row at a time — parts pile
up faster than background merges can combine them, and ClickHouse
starts rejecting new inserts as a safety valve rather than letting the
part count grow unbounded.
**Fix:** batch inserts client-side into the thousands of rows per
`INSERT`, or route through the async insert mechanism described in
[Async Inserts & Batching](/clickhouse/async-inserts). This is
almost always an ingestion-pattern problem, not a server capacity
problem — adding hardware doesn't fix it.
## Query killed for memory (OOM) [#query-killed-for-memory-oom]
**Symptom:** a query fails with a memory limit exceeded error, often a
large `JOIN` or `GROUP BY` with high cardinality.
**Cause:** the query's intermediate state (hash tables for the join or
aggregation) grew past `max_memory_usage`, covered in
[Server Tuning](/clickhouse/server-tuning). This is ClickHouse
protecting the rest of the server, not a bug.
**Fix:** reduce what the query has to hold in memory — filter
earlier, aggregate in stages, or check whether a
[join](/clickhouse/joins) is unintentionally broadcasting a large
table. Raising the memory limit is sometimes correct, but treat it
as a last resort, not the first thing to try.
## An accidental full-table rewrite [#an-accidental-full-table-rewrite]
**Symptom:** a routine [mutation](/clickhouse/mutations)
(`ALTER TABLE ... UPDATE/DELETE`) or an `OPTIMIZE TABLE ... FINAL` run
on a large table turns into a long-running operation that saturates
disk I/O and is awkward to cancel cleanly.
**Cause:** both operations work by rewriting entire parts, not
patching rows in place — exactly as described in
[Mutations](/clickhouse/mutations) and
[Parts & Background Merges](/clickhouse/parts-merges). On a
multi-terabyte table, that's a multi-terabyte rewrite, however small
the actual change.
**Fix:** scope mutations with the narrowest `WHERE` possible (ideally
aligned to a partition), and treat `OPTIMIZE ... FINAL` as a
deliberate, scheduled maintenance operation on large tables — never
a routine one. If rows genuinely need frequent updates, reconsider
the schema (e.g. a `ReplacingMergeTree`) instead of relying on
mutations.
## A replica falling behind [#a-replica-falling-behind]
**Symptom:** reads from one replica return noticeably older data than
another, or the cluster feels inconsistent depending on which node a
query happens to hit.
**Cause:** as described in
[Replication & Keeper](/clickhouse/replication), a replica catches up
by replaying a log coordinated through Keeper. If that replica is
undersized, network-constrained, or busy with its own merges, its
replication queue grows and it drifts further behind.
**Fix:** watch replication queue size as a first-class metric (see
[Monitoring](/clickhouse/monitoring-prometheus-grafana)), not an
afterthought — a queue that only grows is an early warning, not a
one-time blip. A persistently lagging replica usually needs more
resources or fewer competing responsibilities, not a restart.
## Disk full from data that was never supposed to stick around [#disk-full-from-data-that-was-never-supposed-to-stick-around]
**Symptom:** the server refuses inserts, or merges start failing,
because a data disk is out of space.
**Cause:** raw, high-volume tables with no [TTL](/clickhouse/ttl)
policy and no [partition](/clickhouse/partitions)-based cleanup keep
every row forever by default — ClickHouse won't delete anything on
your behalf unless told to.
**Fix:** add a TTL policy (or a scheduled `DROP PARTITION`) to every
table that's meant to have a lifecycle, before it fills a disk — not
after. Disk-full is one of the very few ClickHouse failure modes
that isn't self-healing once it happens; merges themselves need free
space to write their output.
## First three things to check when something's wrong [#first-three-things-to-check-when-somethings-wrong]
1. `system.query_log` for the slow or failing query itself — duration,
rows read, memory used (see
[System Tables](/clickhouse/system-tables)).
2. `system.parts` and `system.merges` for part counts and merge
pressure on the table involved.
3. Your Prometheus/Grafana dashboards (see
[Monitoring](/clickhouse/monitoring-prometheus-grafana)) for the
trend leading up to the incident — a replication queue or disk
usage graph almost always shows the problem building minutes or
hours before it became visible.
Almost every failure on this page is a consequence of one of two
things: writing to ClickHouse the way you'd write to an OLTP
database (tiny frequent writes, row-by-row updates), or letting data
accumulate with no lifecycle plan. Both are schema and
ingestion-pattern problems, not scaling problems — which is why they
show up on a single local instance just as easily as on a large
cluster.
# Monitoring & Operations (/docs/clickhouse/monitoring-operations)
Watching it run, tuning it, and what breaks in production.
# Monitoring (/docs/clickhouse/monitoring-operations/monitoring-prometheus-grafana)
Querying [system tables](/clickhouse/system-tables) by hand is great
for investigating a specific problem right now. It doesn't give you
history, alerting, or a dashboard someone can glance at during an
incident — that's what a metrics pipeline is for.
## The built-in Prometheus exporter [#the-built-in-prometheus-exporter]
ClickHouse ships a Prometheus-compatible metrics endpoint — enabled in
server config and typically served on its own port (or under
`/metrics` on the HTTP port, depending on how it's configured). It
exposes the same kinds of numbers already visible in `system.metrics`,
`system.asynchronous_metrics`, and `system.events`, in a format
Prometheus can scrape on a schedule and retain over time.
## What's actually worth watching [#whats-actually-worth-watching]
A handful of metrics map directly onto concepts already covered in
this module, which is the fastest way to build intuition for what
"normal" looks like on your own cluster:
* Background merge counts and merge queue size — the health of the
process described in
[Parts & Background Merges](/clickhouse/parts-merges). A queue that
only grows means merges can't keep up with insert volume.
* Replication queue size per replica — directly tied to
[Replication & Keeper](/clickhouse/replication). A replica whose
queue keeps growing is falling behind the others.
* Query count, memory usage, and rejected/failed queries — the live
counterparts to what `system.query_log` and `system.processes` show
after the fact.
* Disk space per volume — the earliest warning for the disk-full
failure mode covered in
[Production Failure Scenarios](/clickhouse/failure-scenarios).
## A second, separate integration path: the Grafana data source [#a-second-separate-integration-path-the-grafana-data-source]
Prometheus/Grafana covers server health metrics. Separately,
ClickHouse publishes an official Grafana data source plugin that lets
Grafana query ClickHouse *directly with SQL* — for building dashboards
over your actual application data (events, business metrics) rather
than server internals. It's easy to conflate the two: one path
monitors the database, the other uses the database as a source for
your own analytics dashboards.
Alert on trends, not just thresholds. A replication queue of 500 is
alarming on a cluster that's usually near zero, and unremarkable on
one that always hovers around a few hundred during normal ingestion
bursts. Static thresholds copied from someone else's cluster tend to
either miss real problems or page you constantly for nothing.
# Server Tuning (/docs/clickhouse/monitoring-operations/server-tuning)
ClickHouse has hundreds of settings. Almost none of them matter for a
working local setup, and most production tuning comes down to a
handful that directly correspond to concepts already covered elsewhere
in this module. This page is about those, not an exhaustive settings
reference.
## max\_memory\_usage — the direct lever against OOM [#max_memory_usage--the-direct-lever-against-oom]
Every query has a memory budget. Once a query's intermediate state
(hash tables for joins/aggregations, sort buffers) exceeds
`max_memory_usage`, ClickHouse kills that query rather than let it
take down the whole server. This is the setting standing directly
between a single expensive query and the OOM scenario in
[Production Failure Scenarios](/clickhouse/failure-scenarios).
Set too low, and legitimate heavy queries (large joins, big
`GROUP BY`s) get killed. Set too high — or unset, deferring entirely
to the server-wide default — and one bad query can starve every other
query on the server of memory.
## max\_threads — parallelism per query [#max_threads--parallelism-per-query]
Ties directly into the "within a server" parallelism described in
[Query Optimization](/clickhouse/query-optimization): this caps how
many CPU cores a single query can use at once. Higher generally means
faster individual queries and worse throughput when many queries run
concurrently, since they now compete for the same cores — the classic
latency-vs-throughput trade every tuning knob here eventually comes
back to.
## The mark cache — keeping the sparse index warm [#the-mark-cache--keeping-the-sparse-index-warm]
The sparse index described in
[ORDER BY & Primary Keys](/clickhouse/primary-key) (the array of
"marks", one per granule) has to be read from disk before it can be
binary-searched. The mark cache keeps recently-used marks in memory
across queries, so a table that gets queried repeatedly doesn't pay
that disk read every time. Its size is one of the few caches worth
increasing deliberately on a server with many actively-queried tables
and enough spare RAM.
## The uncompressed block cache — usually leave it off [#the-uncompressed-block-cache--usually-leave-it-off]
A separate, optional cache for already-decompressed data blocks. It
sounds like a pure win, but it competes with the OS page cache (which
is already caching the compressed files ClickHouse reads) for the same
RAM, and decompression with the default LZ4 codec — see
[Compression](/clickhouse/compression) — is already fast. It's
generally only worth enabling for a small number of tables that are
small, hot, and scanned very frequently, not as a blanket setting.
## Three places a setting can be set — and who wins [#three-places-a-setting-can-be-set--and-who-wins]
The same setting name can be configured at three levels. From widest
to narrowest scope, each level overrides the one before it:
1. **Server config** — a global default in `config.xml`, applied
unless overridden.
2. **Settings profile** — assigned to a user or role (see
[Users & Roles](/clickhouse/users-roles)), letting different teams
or applications get different limits from the same server.
3. **Per-query `SETTINGS` clause** — wins over both, as in the
`max_memory_usage` example above.
A per-query `SETTINGS` override is easy to forget you left in a
saved query or dashboard panel. If a query behaves differently than
expected under load, check whether it's silently overriding a server
or profile default before assuming the server itself is
misconfigured.
# System Tables (/docs/clickhouse/monitoring-operations/system-tables)
Most databases make you reach for a separate CLI tool, an admin panel,
or a vendor dashboard to see what the server itself is doing.
ClickHouse exposes almost everything about its own internal state as
ordinary tables in the `system` database — queryable with the same SQL
you already know.
## system.parts — the ground truth for storage [#systemparts--the-ground-truth-for-storage]
Every part described in
[Parts & Background Merges](/clickhouse/parts-merges) is a row here:
its partition, size on disk, row count, and whether it's still active
or waiting to be cleaned up after a merge.
A quick way to spot the "too many parts" problem before it becomes an
error: count active parts per table and watch for numbers climbing
into the thousands.
## system.merges — merge pressure, live [#systemmerges--merge-pressure-live]
Background merges don't run instantly or invisibly — while one is in
progress, it shows up here with its progress percentage and which
parts it's combining.
An empty result isn't necessarily good news on a busy table — it can
mean merges are keeping up comfortably, or that they've stalled.
Cross-check against `system.parts` counts over time to tell the
difference.
## system.mutations — tracking UPDATE/DELETE progress [#systemmutations--tracking-updatedelete-progress]
Every [mutation](/clickhouse/mutations) you run is recorded here until
it finishes, including whether it failed and why.
If a mutation looks stuck, `latest_fail_reason` is the first place
to look — a common cause is the mutation silently retrying against a
part that a concurrent merge keeps replacing.
## system.query\_log — every query that ran, with a cost breakdown [#systemquery_log--every-query-that-ran-with-a-cost-breakdown]
When `query_log` is enabled (on by default in most setups), every
query is logged with how long it took, how many rows and bytes it
read, and how much memory it used — the natural follow-up to the
`EXPLAIN` advice in
[Query Optimization](/clickhouse/query-optimization): `EXPLAIN` tells
you what a query *plans* to do; `query_log` tells you what it actually
did.
## system.processes — what's running right now [#systemprocesses--whats-running-right-now]
Currently executing queries, with their elapsed time and memory usage
so far — and a query ID you can hand to `KILL QUERY` if one is
misbehaving.
None of these tables require special tooling, agents, or a separate
monitoring stack to read — they're just tables. That said, querying
them by hand doesn't replace continuous monitoring; see
[Monitoring](/clickhouse/monitoring-prometheus-grafana) for turning
the same underlying metrics into dashboards and alerts.
# Dictionaries (/docs/clickhouse/query-speed/dictionaries)
ClickHouse is deliberately weak at exactly the thing OLTP databases
are strong at: joining a huge fact table against a small, frequently
looked-up reference table (country codes, product catalogs, user
tiers) row by row. Dictionaries exist to sidestep the join entirely
for that specific, extremely common case.
## The idea: load it once, look it up in memory [#the-idea-load-it-once-look-it-up-in-memory]
Instead of a `JOIN` that has to find matching rows in a second table
for every row of the fact table, a dictionary is loaded into memory
ahead of time — as a hash table or flat array — and looked up with a
plain function call, roughly as fast as reading a local variable.
## Defining and using one [#defining-and-using-one]
`LIFETIME(MIN 300 MAX 600)` tells ClickHouse to reload the dictionary
from its source somewhere between 5 and 10 minutes after the last
load — not on every query, and not on every change in the source.
That staleness window is the trade you're making for speed.
## Layouts trade memory for lookup shape [#layouts-trade-memory-for-lookup-shape]
* `FLAT` — array-indexed by integer key; fastest, but only sensible
for small, dense key ranges.
* `HASHED` — hash table; the general-purpose default for integer or
string keys.
* `RANGE_HASHED` — for keys that change over time (e.g. exchange
rates valid within a date range).
* `DIRECT` — no local storage; queries the source live, per lookup.
Used when the reference table is too large to hold in memory.
A dictionary and a [materialized view](/clickhouse/materialized-views)
solve opposite directions of the same problem: a dictionary pulls
small, slow-changing reference data *in* on a schedule; a
materialized view pushes derived data *out* on every insert.
Dictionaries are loaded per-server, in full, into memory (except
`DIRECT`). A dictionary sized for "a few hundred thousand
countries/products/tiers" is fine; one sized for "every user who
ever signed up" will quietly consume a lot of RAM on every node in
the cluster.
# Query Speed (/docs/clickhouse/query-speed)
The tricks that make scans over billions of rows fast.
# Materialized Views (/docs/clickhouse/query-speed/materialized-views)
In Postgres, a materialized view is a query result you periodically
`REFRESH` — it goes stale between refreshes. A ClickHouse
materialized view is a completely different mechanism: it's an
**insert trigger** that runs a query against every newly inserted
block of rows and pushes the output into a separate target table,
continuously, forever.
## The mental model: push, not pull [#the-mental-model-push-not-pull]
{/* Not a chain: INSERT fans out into two independent things — the
block lands in the source table AND triggers the MV query,
whose output lands in the target table. A straight Pipeline
chain would wrongly imply the source table produces the MV
trigger. Grid columns are 1fr/1fr so the fork's 25%/75% marks
land dead-center on each branch regardless of viewport. */}
Crucially, the view's `SELECT` never runs against the whole source
table — only against the rows in the block that was just inserted.
The cost of maintaining it is proportional to insert volume, not to
the size of the source table.
## The classic pattern: pre-aggregation [#the-classic-pattern-pre-aggregation]
Reads against `events_per_minute` stay fast forever, because they
scan pre-aggregated per-minute rows instead of raw events — even
after the source table grows into the billions.
A materialized view only sees rows inserted **directly** into its
source table. If another materialized view (or a distributed
insert path) writes into that table indirectly, or if you insert
straight into the target table, the trigger may not fire the way
you expect. Always trace where writes actually land.
The target table uses `AggregateFunction(count)` and
`AggregatingMergeTree`, not a plain integer with
`SummingMergeTree` — this lets partial aggregate states from
different inserts merge correctly later. Reading it back requires
`countMerge()`, not a plain `SUM`.
## Materialized view vs. projection [#materialized-view-vs-projection]
Use a materialized view when you want a differently *shaped* result
— aggregated, filtered, joined — living in its own named table you
query explicitly. Use a [projection](/clickhouse/projections) when
you want the *same* rows just sorted differently, chosen
automatically under the same table name.
# Projections (/docs/clickhouse/query-speed/projections)
A table can only be physically sorted one way. If half your queries
filter by `user_id` and the other half filter by `event_type`, one
of those query patterns will never get the sparse index's help — no
matter how you pick `ORDER BY`. Projections solve this by keeping
*additional* physical copies of the data, sorted differently, that
the query optimizer chooses between automatically.
## Declaring one [#declaring-one]
`ADD PROJECTION` only registers the definition; it applies to new
parts as they're written. `MATERIALIZE PROJECTION` builds it for
existing data. Once built, it's maintained automatically on every
future insert and merge — you never write to it directly.
## How the optimizer uses it [#how-the-optimizer-uses-it]
You keep querying the base table as normal. ClickHouse examines the
query's `WHERE`/`GROUP BY` and picks whichever physical layout — the
base table or one of its projections — would read the least data to
answer it. This is transparent: no query rewriting on your end, no
risk of querying "stale" data, since both copies are updated as part
of the same insert.
## Projections vs. materialized views [#projections-vs-materialized-views]
Both maintain a second copy of data incrementally. The difference is
what that copy is *for*:
* A **projection** stores the same rows (or an aggregate of them),
physically re-sorted, and is chosen *automatically* by the query
planner against the same table name.
* A **materialized view** writes into a separate, explicitly named
target table that you query directly yourself. See
[Materialized Views](/clickhouse/materialized-views).
Every projection roughly doubles (or more) the storage and insert
cost for that table, since every write now maintains N physical
copies. Add projections deliberately for specific, high-value query
patterns — not speculatively for every column combination you might
someday filter on.
# Query Optimization (/docs/clickhouse/query-speed/query-optimization)
Every topic in this module — partitions, the sparse index, skip
indexes, projections — exists to answer one question before a query
even starts executing: *how much data can we avoid reading?* This
page is about what happens to whatever data is left after that
pruning.
## The pruning pipeline [#the-pruning-pipeline]
## Vectorized execution: blocks, not rows [#vectorized-execution-blocks-not-rows]
Row-at-a-time execution (call a function, get a row, call it again)
spends most of its time on function-call and branching overhead, not
actual work. ClickHouse's engine instead processes data in
**blocks** — batches of a few thousand to \~65,536 values from one
column at a time — so a filter or arithmetic operation runs as a
tight loop over contiguous memory, which the CPU can pipeline and
auto-vectorize (SIMD) effectively.
## PREWHERE: filter before you even fetch every column [#prewhere-filter-before-you-even-fetch-every-column]
A normal `WHERE` reads all columns needed for the whole query, then
filters. `PREWHERE` reads a cheap column first (here, `status`),
filters, and only then reads the remaining columns (like `payload`)
for the rows that survived — skipping decompression of expensive
columns for rows that were going to be discarded anyway. In recent
versions ClickHouse often applies this automatically; it's still
worth understanding, and sometimes worth forcing explicitly.
## Parallelism, two ways [#parallelism-two-ways]
* **Within a server** — a single query is split across CPU cores,
each processing a different range of granules concurrently.
* **Across a cluster** — a query against a
[Distributed table](/clickhouse/distributed-tables) fans out to
every shard in parallel and merges partial results.
## Reading the plan [#reading-the-plan]
This shows exactly which partitions and how many granules were
pruned by the primary key and by any skip indexes — the fastest way
to confirm a schema decision is actually paying off, instead of
guessing from query latency alone.
Before reaching for a bigger cluster, run `EXPLAIN indexes = 1` on
your slow query. Very often the fix is a better `ORDER BY`, a
missing skip index, or a partition key that doesn't match the
query pattern — not more hardware.
# Skip Indexes (/docs/clickhouse/query-speed/skip-indexes)
The primary key only helps if your query filters on a column that is
part of `ORDER BY`. Skip indexes give ClickHouse a way to avoid
reading granules based on *other* columns too — without physically
reordering anything.
## What a skip index actually stores [#what-a-skip-index-actually-stores]
Unlike a Postgres index, a skip index doesn't point at rows. It
stores a small summary — per group of granules — and uses that
summary to decide whether a granule *could possibly* contain a
match. If it can't, the granule is skipped entirely: never
decompressed, never scanned.
The first two granules have a min/max range of `200–304` — the index
proves `500` cannot be in there, so they're skipped. The third
granule's range includes 500, so it gets scanned. No sorting was
required for this to work; the index just needed the data to already
be somewhat clustered — which is usually true if `status_code`
correlates at all with time, or with a column earlier in `ORDER BY`.
## The index types [#the-index-types]
* `minmax` — stores min and max per block of granules. Cheap,
effective for numeric/date columns with any locality.
* `set(N)` — stores up to N distinct values per block. Good for
low-cardinality columns where exact-match filtering is common.
* `bloom_filter` — probabilistic membership test; can have false
positives (scans a granule unnecessarily) but never false
negatives (never skips a granule that has a match).
* `tokenbf_v1` / `ngrambf_v1` — bloom filters over tokens or n-grams,
for accelerating `LIKE` / substring search inside text columns.
A skip index only helps when values are **locally clustered** —
not uniformly scattered across every granule. Adding a `minmax`
index on a column with random values spread evenly through the
whole table gives every granule the same wide min/max range, so
nothing ever gets skipped. Check correlation with your existing
sort order before adding one.
`GRANULARITY 4` means one index entry covers 4 granules (so 4 ×
8192 = 32768 rows by default), not one entry per granule. Larger
granularity means a smaller index but coarser skipping.
# Distributed Tables (/docs/clickhouse/scale-ops/distributed-tables)
Everything so far assumed one server. Sharding is how ClickHouse
scales past what one machine's disk and CPU can hold: the table is
split by row across multiple servers, and a special engine
(confusingly also just called a table) knows how to talk to all of
them at once.
## Two tables, two jobs [#two-tables-two-jobs]
On each shard, you create an ordinary `MergeTree` table — the "local"
table, which actually stores rows. Separately, you create a
`Distributed` table, typically on every node, which stores *no data at
all* — it's a router that knows the cluster topology and a sharding
expression.
Query the `Distributed` table and it forwards the query to
`events_local` on every shard in parallel, then merges the partial
results — sums get summed, counts get summed, top-N results get
re-sorted and truncated. From the client's perspective it's one table;
underneath, it's N independent MergeTree tables each holding a slice
of the rows.
## The sharding key decides the slice [#the-sharding-key-decides-the-slice]
`cityHash64(user_id)` above means every row for a given `user_id`
always lands on the same shard — useful when queries frequently filter
or aggregate per user, since that work never needs to cross shards. A
poor sharding key (or none — random distribution) is fine for
full-table scans but forces more cross-shard coordination for anything
that needs to group by the key you didn't shard on.
Inserting into a `Distributed` table by default routes each row to
its shard synchronously as part of the insert, which is slow and
fragile over the network. Production setups almost always insert
directly into the local table on each shard from the ingestion
pipeline, or enable `distributed_foreground_insert` /async settings
deliberately, rather than relying on default distributed inserts at
scale.
Sharding solves a different problem than
[replication](/clickhouse/replication). Sharding is about capacity —
spreading data too big for one machine across many. Replication is
about availability — keeping copies so no single machine failing
loses data. Production clusters combine both: each shard is itself a
replicated group of servers.
# Scale & Ops (/docs/clickhouse/scale-ops)
Running it across machines, safely, over time.
# Replication & Keeper (/docs/clickhouse/scale-ops/replication)
Plain `MergeTree` has no idea other servers exist. `ReplicatedMergeTree`
is the same storage engine with one addition: every meaningful
operation (a new part appearing, a merge completing) is recorded in a
shared log that other replicas watch and replay — coordinated through
**ClickHouse Keeper**.
## Keeper: the thing that makes replicas agree [#keeper-the-thing-that-makes-replicas-agree]
Keeper is ClickHouse's built-in replacement for ZooKeeper — a small,
separate consensus service (Raft-based) that stores the replication
log and coordinates leader election for tasks like deciding which
replica performs a given merge. ClickHouse the database doesn't do
consensus itself; it delegates that entirely to Keeper.
## Declaring a replicated table [#declaring-a-replicated-table]
The Keeper path identifies which replicas belong to the same logical
table — every replica registers itself under the same path, using its
own unique replica name. This is the piece that's easy to get wrong: a
typo in the path silently creates an unrelated, unreplicated table
instead of joining the group.
## What replication actually guarantees [#what-replication-actually-guarantees]
* A write acknowledged by one replica will eventually exist on all of
them — via the shared log, not by re-sending the insert to every
replica.
* Any replica can serve reads independently; ClickHouse doesn't
require quorum reads by default, which is faster but means a read
can briefly lag the very latest write on a different replica.
* If a replica goes down and comes back, it catches up by replaying
the log — it doesn't need a full manual resync unless it was down
long enough to fall outside Keeper's log retention.
Keeper itself needs to run as a quorum (typically 3 nodes) to
tolerate a single node failure — a single Keeper instance is a
single point of failure for the entire cluster's coordination, even
if the data replicas themselves are numerous.
Replication (availability, via `ReplicatedMergeTree`) and
[sharding](/clickhouse/distributed-tables) (capacity, via
`Distributed`) are independent axes. A production cluster's typical
shape is: each shard is a small group of replicas, and a
`Distributed` table routes across shards while each shard tolerates
individual node failure via replication.
# TTL (/docs/clickhouse/scale-ops/ttl)
Analytical data almost always has a shelf life: raw events matter most
in the first days, are worth keeping cheaply for a while longer, and
eventually are worth nothing except storage cost. TTL (time-to-live)
lets you encode that lifecycle directly into the table, instead of
writing a cron job that runs `DELETE` statements.
## Declaring a lifecycle [#declaring-a-lifecycle]
## Enforced by merges, not a scheduler [#enforced-by-merges-not-a-scheduler]
TTL rules aren't evaluated by a separate cron-like process — they're
checked whenever a background merge touches a part. Expired rows are
dropped (or moved) as a side effect of the normal merge cycle. This
means TTL cleanup has no fixed schedule: a part that rarely gets
touched by merges may hold expired rows a bit longer than the TTL
literally states, though ClickHouse also runs periodic housekeeping
merges specifically to catch this.
## What TTL can do besides delete [#what-ttl-can-do-besides-delete]
* `TTL ... DELETE` — drop rows once the expression is in the past.
* `TTL ... TO VOLUME 'name'` / `TO DISK 'name'` — move the containing
part to cheaper storage (e.g. HDD or S3-backed disks) without
deleting anything, given a configured storage policy with multiple
tiers.
* `TTL ... GROUP BY` — instead of deleting expired rows, roll them up
into an aggregate first, then delete the raw rows. Useful for "keep
raw events for 7 days, keep hourly rollups forever."
* Column-level TTL — expire just one column (setting it to its
default) while keeping the rest of the row.
Multi-tier TTL (`TO VOLUME` then `DELETE`) requires a
`storage_policy` with multiple disks/volumes configured in server
config first — the `TTL` clause on the table just references a
policy that must already exist.
TTL expressions are evaluated per-part during merges, which means
forcing immediate cleanup on demand requires forcing a merge:
`OPTIMIZE TABLE events FINAL` — the same tool used to force part
merges in general, covered in
[Parts & Background Merges](/clickhouse/parts-merges).
# Backup & Restore (/docs/clickhouse/security-backup/backup-restore)
Replication protects against a machine dying. It does nothing against
a bad `ALTER TABLE`, an accidental `DROP TABLE`, or a bug that
quietly corrupts data — those mistakes replicate just as faithfully
as good data does. Backups are the only defense against mistakes, not
just hardware failure.
## Native BACKUP / RESTORE [#native-backup--restore]
ClickHouse has built-in SQL statements for this — no external tool
required for the basic case:
Because [parts are immutable](/clickhouse/parts-merges), a backup
taken while a table is under active insert load doesn't need to
freeze writes to get a consistent snapshot — it just references the
parts that existed at that instant, the same property that makes
concurrent reads safe without locking. Repeated backups of a
slowly-changing table can also be meaningfully incremental, since
unchanged parts don't need to be copied again.
## Where most production setups actually land: clickhouse-backup [#where-most-production-setups-actually-land-clickhouse-backup]
`clickhouse-backup` is a widely used community tool (not an official
ClickHouse project) built on top of the same underlying mechanics,
adding what the bare SQL statements don't: scheduled backups,
retention policies, easier full-cluster/multi-table coordination, and
simpler restore workflows across environments. Most teams running
ClickHouse in production reach for it rather than scripting `BACKUP`
statements themselves.
## What backups don't cover on their own [#what-backups-dont-cover-on-their-own]
* Table structure/DDL history — back up your migration scripts too,
not just data, so a restored table's schema isn't a guess.
* Users, roles, and quotas — access control state generally needs its
own backup path, separate from table data.
* Dictionaries and their external sources — a dictionary definition
backs up fine, but the reference data it loads from lives elsewhere
and needs its own plan.
A backup you have never restored is a hypothesis, not a safety net.
This matters more for an analytical store than for most OLTP
systems: the raw events sitting in a ClickHouse table are often the
only copy of that data anywhere — there's no upstream system to
re-derive them from if a restore turns out not to work. Test the
restore path on a schedule, into a throwaway table or environment,
before you need it for real.
# Security & Backup (/docs/clickhouse/security-backup)
Users, permissions, and not losing data.
# Row Policies, Quotas & TLS (/docs/clickhouse/security-backup/row-policies-quotas)
[Users and roles](/clickhouse/users-roles) control what a user can
query. Row policies control *which rows* they see when they run that
query, and quotas control *how much* of the cluster's resources
they're allowed to consume doing it. Both exist to let one ClickHouse
cluster safely serve multiple tenants or teams instead of needing a
cluster each.
## Row policies: filtering rows per user, transparently [#row-policies-filtering-rows-per-user-transparently]
A row policy attaches a filter condition to a table. Any query
against that table from a matching user has the condition silently
ANDed into its `WHERE` clause — the user never sees rows outside the
filter, and doesn't need to remember to add the filter themselves.
Here, a user's own name doubles as their tenant identifier via
`currentUser()` — a common pattern when each tenant maps to a
dedicated ClickHouse user. More typically the condition references a
session setting or a mapping table rather than the username directly,
but the mechanism is the same: the policy is enforced by ClickHouse
itself, not by application code remembering to filter correctly on
every query.
The value of a row policy over "just add `WHERE tenant_id = ?` in
the application" is that it can't be forgotten. One missed `WHERE`
clause in one ad-hoc query or one new dashboard is a data leak; a
row policy makes that class of bug structurally impossible for that
user.
## Quotas: limiting how much a user can consume [#quotas-limiting-how-much-a-user-can-consume]
A quota caps resource usage — queries, errors, rows read, execution
time — over a rolling interval, per user or role. It doesn't make
individual queries faster; it protects everyone else on a shared
cluster from one runaway report, one misbehaving job, or one
accidental `SELECT *` over a trillion-row table.
When a user tied to this quota exceeds it, further queries are
rejected until the interval resets — a blunt but effective circuit
breaker that requires no application-side rate limiting.
## Network security, briefly [#network-security-briefly]
Two settings matter most day to day: `listen_host` controls which
network interfaces the server accepts connections on (default
configs often bind to all interfaces, which is fine inside a private
Docker network but not on an open host), and `tcp_port_secure` /
HTTPS enable TLS for the native and HTTP protocols respectively. In a
replicated or sharded cluster, an `interserver_http_credentials`
secret authenticates traffic between replicas and shards themselves,
separate from any client-facing user — worth knowing exists, mostly a
config-file concern rather than a deep concept.
None of this is on by default in a bare-bones local setup like this
repo's — one open port, no TLS, one admin user. That's the correct
trade for learning. Before anything here talks to the public
internet, TLS and a non-default `listen_host` stop being optional.
# Users & Roles (/docs/clickhouse/security-backup/users-roles)
ClickHouse's access control looks a lot like Postgres's: users
authenticate, roles bundle up privileges, and users get roles
assigned to them rather than having privileges granted to them one by
one. The part that surprises people coming from a single-tenant
analytics mindset is how fine-grained the grants can get — down to
specific columns and row filters, not just whole databases.
## Users, roles, and grants [#users-roles-and-grants]
A **user** is an identity that authenticates (password, certificate,
LDAP, Kerberos, etc.). A **role** is a named bundle of privileges that
gets assigned to one or more users. Privileges are granted with
`GRANT`, either straight to a user or to a role that users then
inherit:
This separation matters in practice: when the ingestion pipeline
changes, you touch `events_writer` once and every user holding that
role picks up the change, instead of hunting down every individual
grant.
## Checking what a user can actually do [#checking-what-a-user-can-actually-do]
This is the first thing to run when a query fails with an access
denied error, or — more worryingly — when you're trying to confirm a
user *can't* do something it shouldn't. `REVOKE` works symmetrically
with `GRANT` to take a privilege back.
## Two ways grants get stored [#two-ways-grants-get-stored]
By default, users, roles, and grants created with SQL are persisted
internally (backed by the same coordination storage used for
replication — ClickHouse Keeper in a clustered setup, local disk
otherwise), which is what the examples above assume. Older
deployments — and some still today — instead define users in XML
configuration files (`users.xml`) loaded at server startup. Both
mechanisms can coexist, but mixing them for the same user is a common
source of confusion about which definition actually won.
This repo's own [Docker setup](/clickhouse/setup) creates exactly
one user — the admin user from `.env` — with full access to
everything. That's the right amount of ceremony for a local
learning environment. It is the wrong shape for anything shared or
production: every application and every human that touches the
cluster should get its own scoped user, with a role that grants
only what that specific workload needs.
Prefer granting roles over granting privileges directly to users,
even for a single user. It costs nothing up front, and it's the
difference between updating one role definition and auditing every
user account by hand when access requirements change six months
later.
# Aggregate Combinators (/docs/clickhouse/sql-querying/aggregate-combinators)
ClickHouse lets you attach a small set of suffixes — **combinators**
— to almost any aggregate function to change its behavior. Once you
know the combinators, you can often replace a subquery, a `CASE`
expression, or a self-join with a single function call.
## -If: conditional aggregation without a subquery [#-if-conditional-aggregation-without-a-subquery]
Appending `If` to any aggregate function adds a condition as its
last argument — the function only considers rows where the condition
is true, computed in a single pass over the data instead of one pass
per condition.
## -Array: aggregating over array columns [#-array-aggregating-over-array-columns]
Appending `Array` makes an aggregate function treat its argument as
an array and aggregate over all the array's elements across all
rows, rather than over one scalar value per row.
## -State / -Merge: partial aggregates that combine later [#-state---merge-partial-aggregates-that-combine-later]
This is the mechanism behind
[materialized views](/clickhouse/materialized-views) that
pre-aggregate incrementally, covered later using `countState()` /
`countMerge()`. The general pattern applies to any aggregate
function:
* `fooState()` — instead of returning a final value, returns an
opaque, mergeable intermediate state (stored via the
`AggregateFunction(foo, ...)` column type).
* `fooMerge()` — combines many stored states back into one, and
produces the actual final value.
This is what lets a table hold, say, one row of pre-aggregated
"average response time per minute" per source server, and later
merge those per-server states into a correct overall average — which
a naive average-of-averages would get wrong.
## Counting distinct values: the uniq family [#counting-distinct-values-the-uniq-family]
ClickHouse doesn't have a `-Distinct` combinator — distinct counting
is its own family of functions, because exact and approximate
distinct counting have very different costs:
* `uniqExact` — exact count, implemented by holding every distinct
value seen. Correct, but memory cost grows with cardinality.
* `uniq` — approximate count using an adaptive sampling algorithm;
small, bounded memory use, small statistical error.
* `uniqCombined` — approximate, tuned to use even less memory than
`uniq` for very large cardinalities, at a similar error rate.
## argMax / argMin: "the row where X was largest" [#argmax--argmin-the-row-where-x-was-largest]
A frequent pattern — find the value of one column at the row where
another column is at its maximum — normally needs a subquery or a
window function. `argMax`/`argMin` do it in one call:
This returns, per user, the `plan` value from whichever row has the
largest `updated_at` — effectively "latest plan per user" without a
self-join or a `ROW_NUMBER() OVER (...) = 1` filter.
Combinators exist because ClickHouse's execution engine is built
around single-pass, vectorized aggregation (see
[Query Optimization](/clickhouse/query-optimization)). Expressing
conditional logic, distinct counts, and latest-value-by-key as
aggregate function variants keeps them inside that single pass,
instead of requiring extra scans, joins, or subqueries.
# SQL & Querying (/docs/clickhouse/sql-querying)
Joins, windows, and the SQL features that behave differently here.
# Joins (/docs/clickhouse/sql-querying/joins)
Joins are the part of SQL ClickHouse is honest about being weaker
at. A mature OLTP engine like Postgres has decades of join-order
planning, persistent indexes on both sides, and statistics-driven
cost estimation. ClickHouse's planner is younger and simpler,
and its storage engine has no concept of a join index — every join
has to build one from scratch, at query time.
## How a join actually runs [#how-a-join-actually-runs]
The default algorithm is a hash join: the **right-hand** table (or
subquery) is read in full and loaded into an in-memory hash table,
keyed on the join columns. Then the left-hand table is streamed row
by row, probing that hash table for matches.
{/* Not a chain: two independent inputs converge on "Matched rows" —
the right side is read once and built into a hash table while
the left side streams straight down to the same probe step. A
straight Pipeline chain would wrongly imply the hash table
produces the left-side stream. The left column's line is a
single row-span-2 div so it lands exactly on the hash table's
bottom edge, whatever that box's rendered height turns out to
be — no hardcoded height to keep in sync. */}
The consequence follows directly from the mechanism: the right-hand
side has to fit comfortably in memory. There is no equivalent of a
Postgres merge join over two pre-sorted, indexed tables that never
materializes either side in full.
## GLOBAL JOIN: the distributed correctness trap [#global-join-the-distributed-correctness-trap]
This is the single most common way to silently get wrong answers
out of a ClickHouse cluster. When a query against a
[Distributed table](/clickhouse/distributed-tables) runs
a join, each shard executes the join **independently, against its
own local data**. If the right-hand table is itself sharded (not
fully present on every node), each shard only ever sees its own
slice of it — rows on other shards silently never match, and you
get fewer results than you should, with no error raised.
A plain `JOIN` against a `Distributed` table can look correct in
testing — on a single-shard cluster, or when the right-hand table
happens to be small enough that someone replicated it identically
everywhere — and then quietly under-count the moment a second shard
is added. If either side of a join involves a `Distributed` table,
default to `GLOBAL JOIN` and only drop it once you've confirmed the
right-hand table is fully present on every shard.
## Keep the right-hand side small [#keep-the-right-hand-side-small]
Because the right-hand table is built into an in-memory hash table
in full before any matching happens, join performance and memory
usage scale with the size of the right-hand side, almost
independently of the left. Practical guidance:
* Put the smaller table on the right — ClickHouse doesn't reliably
reorder this for you the way a cost-based OLTP planner would.
* For small, slow-changing reference data (country codes, plan
tiers, product catalogs), prefer a [dictionary](/clickhouse/dictionaries)
and `dictGet()` over a join entirely — same lookup, no per-query
hash table build, no `GLOBAL` broadcast to worry about.
* Filter both sides down with `WHERE` before the join runs wherever
possible, rather than joining first and filtering after.
`join_algorithm` can be set to alternatives like `partial_merge` or
`full_sorting_merge` when the right-hand table is too large to hash
in memory comfortably — they trade some speed for lower memory
pressure by working off sorted, spillable data instead of an
in-memory hash table. Reach for these when a join fails with a
memory limit error, not by default.
# Mutations (UPDATE / DELETE) (/docs/clickhouse/sql-querying/mutations)
`ALTER TABLE ... UPDATE` and `ALTER TABLE ... DELETE` exist and use
familiar syntax, but they don't work anything like an OLTP
`UPDATE`/`DELETE`. Understanding why comes straight from the same
fact covered in
[Parts & Background Merges](/clickhouse/parts-merges): parts are
immutable. Nothing in a part is ever edited in place — not even by a
mutation.
## What actually happens [#what-actually-happens]
A mutation is applied asynchronously, part by part. For every
existing part that contains at least one row matching the mutation's
condition, ClickHouse rewrites the **entire part from scratch** —
every column, every row, not just the ones that changed — with the
update or deletion applied, then atomically swaps the new part in
for the old one.
`ALTER TABLE ... UPDATE/DELETE` statements return immediately after
being queued — they don't block waiting for the rewrite to finish.
Track progress with:
See [System Tables](/clickhouse/system-tables) for more on querying
operational state like this directly with SQL.
A single `UPDATE` touching even a small fraction of rows can force
a rewrite of an entire large partition's worth of parts, competing
for I/O with normal background merges. Mutations are for occasional
corrections and backfills, not a routine part of your application's
write path — if a workload needs frequent row-level updates, that's
a sign ClickHouse (or at least this table's design) is the wrong
tool for that part of the job.
## Lightweight DELETE: the cheaper alternative [#lightweight-delete-the-cheaper-alternative]
Newer ClickHouse versions support `DELETE FROM table WHERE ...` as a
distinct, lighter-weight operation. Instead of immediately rewriting
affected parts, it marks the matching rows as deleted in a mask;
those rows are then filtered out at query time and physically
dropped later, as a side effect of normal background merges — the
same mechanism that already reclaims space for other reasons.
This is meaningfully cheaper for deletes specifically, but it is
still not free, and it doesn't help with `UPDATE` — there is no
equivalent "lightweight update", because changing a value (as
opposed to hiding a row) has no way to be deferred the same way.
If a table's natural access pattern really is
"replace/deduplicate by key on write," reach for
[ReplacingMergeTree](/clickhouse/mergetree) instead of planning
around frequent mutations — let merges resolve the latest version
the way the engine already does it for free, rather than paying
full part rewrites on every correction.
# Window Functions (/docs/clickhouse/sql-querying/window-functions)
`GROUP BY` collapses rows into one row per group. Sometimes you want
the aggregate *alongside* every original row instead — a running
total next to each transaction, a rank next to each score. That's
what window functions are for.
## The shape of a window function [#the-shape-of-a-window-function]
`PARTITION BY` groups rows the same way `GROUP BY` would, but every
row in the group is kept. `ORDER BY` inside `OVER (...)` defines the
order the window function walks rows in within each partition — it
has nothing to do with the table's `ORDER BY` (the primary key). The
frame clause, `ROWS BETWEEN ... AND ...`, controls exactly which rows
around the current one are included in the calculation; omitting it
defaults to the whole partition for most aggregate functions used
this way.
## Common functions [#common-functions]
* `row_number()` — a unique, gapless sequence per partition.
* `rank()` / `dense_rank()` — ranking with (`rank`) or without
(`dense_rank`) gaps after ties.
* `lag(col, n)` / `lead(col, n)` — the value of a column `n` rows
before/after the current one in the ordered partition — useful for
period-over-period comparisons.
* `sum`, `avg`, `min`, `max` used with `OVER (...)` instead of
`GROUP BY`.
Window functions run as a distinct processing step after
aggregation and filtering but before the final `ORDER BY`/`LIMIT`
of the outer query — you can reference a window function's output
in an outer `WHERE` only via a subquery or CTE, the same
restriction most SQL engines share.
## SAMPLE: trading accuracy for speed [#sample-trading-accuracy-for-speed]
Separately from window functions, but in the same spirit of "SQL
features that behave differently here": the `SAMPLE` clause lets a
query run against a deterministic fraction of a table's rows instead
of all of them, for fast approximate answers on huge tables during
exploration.
`SAMPLE` only works on a table that declares a `SAMPLE BY` expression
(usually a hash of some column, so the sampling is deterministic and
evenly distributed rather than arbitrary). `SAMPLE 0.1` reads roughly
10% of the data; results for aggregates need to be scaled back up
manually, as shown above — ClickHouse doesn't do that scaling for
you.
`SAMPLE` is for exploratory or dashboard queries where an
approximate number returned in milliseconds beats an exact number
returned in seconds — not for anything where correctness matters,
like billing or financial reporting.
# Compression (/docs/clickhouse/storage-engine/compression)
Columnar storage doesn't just mean "read fewer columns" — it also
compresses dramatically better than row storage, because every value
sitting next to another value on disk is now the *same type of
thing*: a column of country codes next to more country codes, a
column of timestamps next to more timestamps. Similar values compress
far better than a shuffled row of unrelated types ever could.
## Two layers of compression, per column [#two-layers-of-compression-per-column]
Each column's data passes through an optional specialized
**encoding**, then a general-purpose **compressor**. Both are
configurable per column.
{/* Fixed 3-column grid (box / connector / box), S-shaped across two
rows — same technique as the Docker setup diagram: every width
is intrinsic to its content, nothing depends on a measured
container width, so there's no runtime layout math that can
drift out of sync with an alignment assumption. */}
### Specialized codecs exploit structure the compressor can't see [#specialized-codecs-exploit-structure-the-compressor-cant-see]
* `Delta` — stores the difference between consecutive values. Great
for slowly-increasing IDs or timestamps, where the deltas are much
smaller numbers than the values themselves.
* `DoubleDelta` — deltas of deltas. Even better for
near-constant-interval timestamps.
* `Gorilla` — designed for floating-point time-series (metrics) where
consecutive values are close together.
* `T64` — transposes bits of fixed-width integers to expose more
redundancy before general compression.
### General compressors trade speed for ratio [#general-compressors-trade-speed-for-ratio]
* `LZ4` (default) — very fast to decompress, which matters more than
raw ratio for most analytical queries that decompress a column, use
it, and move on.
* `ZSTD` — noticeably better compression ratio, more CPU per read.
Common choice for cold/rarely-queried data or when storage cost
dominates.
`LowCardinality(String)` isn't a compression codec, but it belongs
in the same conversation: it dictionary-encodes a string column
(storing small integer IDs instead of repeated text), which shrinks
both storage and the amount of data the query engine has to touch —
a huge win for columns like country, status, or event type.
Sort order affects compression too: a column that is the second or
third key in `ORDER BY` tends to have long runs of repeated or
near-sequential values within each granule, which is exactly what
these codecs are built to exploit. Good schema design (primary key)
and good compression are not separate problems.
# Storage Engine (/docs/clickhouse/storage-engine)
How ClickHouse actually stores rows on disk.
# Partitions (/docs/clickhouse/storage-engine/partitions)
It's easy to confuse partitions with the primary key — both involve
splitting data up. The primary key sorts rows *within* a part. A
partition decides *which parts a row can ever end up in*, at a much
coarser, physically visible level: partitions are real directories on
disk.
## Declaring a partition key [#declaring-a-partition-key]
Every row is assigned a partition by evaluating the partition
expression — here, its year and month. Merges never combine parts from
different partitions, so each partition evolves as an independent set
of parts.
## What it buys you: partition pruning [#what-it-buys-you-partition-pruning]
When a query's `WHERE` clause can be matched against the partition
expression, ClickHouse skips entire partitions before opening a single
file inside them — before the sparse index is even consulted.
## The other reason partitions exist: bulk operations [#the-other-reason-partitions-exist-bulk-operations]
Because a partition is a physical, self-contained set of files, it can
be manipulated as a unit — instantly:
`DROP PARTITION` unlinks files; it doesn't rewrite anything, so it's
effectively instant even on a huge partition — compare that to a
row-by-row `DELETE`, which MergeTree handles far more expensively.
Partitioning by something high-cardinality (per-user, per-hour on a
high-volume table) creates thousands of tiny partitions, each
merging independently — the opposite of what you want. A good
partition key produces a modest number of large partitions (weekly
or monthly is typical), not a huge number of small ones.
A table doesn't need a partition key at all — without one, every
part lives in a single implicit partition. Add one only when you
actually need pruning by date/tenant or bulk drop/move — not by
default on every table.
# Parts & Background Merges (/docs/clickhouse/storage-engine/parts-merges)
Every `INSERT` into a MergeTree table creates a new **part** — never
modifies an old one. Understanding parts and how they merge explains
almost every ClickHouse operational quirk: why small, frequent inserts
are bad, why `SELECT COUNT(*)` is instant, and why `OPTIMIZE TABLE`
exists.
## What a part actually is [#what-a-part-actually-is]
A part is a directory on disk. Inside it: one compressed file per
column, a sparse primary index, checksums, and metadata like row count
and column min/max values. A part is **immutable** — once written, it
is only ever read or deleted, never edited in place.
This is where [compressed blocks](/clickhouse/compression) and
[granules](/clickhouse/primary-key) meet. Each column's `.bin` file is
just its compressed blocks written back to back. The paired `.mrk2`
("marks") file is the missing link between them: one mark per granule,
pointing at the exact byte offset in `.bin` where that granule's block
starts. `primary.idx` is the sparse index itself — one entry per
granule, holding the primary key of its first row. A lookup
binary-searches `primary.idx` for the right granule, reads the
matching mark, and seeks straight to that byte offset — no scanning,
no decompressing anything it doesn't need.
The `2` in `.mrk2` is a format version — ClickHouse has used `.mrk`,
`.mrk2`, and `.mrk3` as the on-disk marks format evolved (e.g. to
support adaptive granule sizes). The role never changes: granule →
byte offset.
## Why merges are necessary [#why-merges-are-necessary]
If every insert became a permanent, separate part, a table that
receives thousands of small inserts a day would end up scanning
thousands of tiny files per query — index overhead per part, open file
handles per part, no sorting across inserts. Background merges
continuously combine smaller sorted parts into fewer, larger sorted
parts, restoring the property that the whole table behaves like one
big sorted structure.
This is the same idea as compaction in an LSM-tree (RocksDB, Cassandra,
LevelDB): accept writes fast by appending, then pay the reorganization
cost later, in the background, off the write path.
## Merges are a suggestion, not a promise [#merges-are-a-suggestion-not-a-promise]
ClickHouse decides when and which parts to merge based on part size
and count — you don't control the schedule. You can force it:
`OPTIMIZE ... FINAL` forces a full merge of every part into one,
rewriting the entire table's data on disk. It is expensive and
I/O-heavy — reasonable for a one-off cleanup on a small/medium table,
dangerous to run routinely on a large one.
## The practical consequence: batch your inserts [#the-practical-consequence-batch-your-inserts]
Because every `INSERT` statement creates at least one new part
regardless of size, sending one row per `INSERT` is one of the most
common ways to misuse ClickHouse — it creates a huge number of tiny
parts faster than the background merge process can keep up, and the
server starts rejecting inserts with `Too many parts`.
* Batch inserts client-side: thousands of rows per `INSERT`, not one.
* Or insert through a `Buffer` table / async insert queue that batches
for you.
* Fewer, larger inserts → fewer parts → less merge pressure → faster
queries.
`SELECT count() FROM table` without a `WHERE` is instant because
each part already stores its row count in metadata — ClickHouse just
sums those numbers without reading a single row of actual data.
# Why ClickHouse (/docs/clickhouse/why-clickhouse)
What problem it solves, and when it's the wrong choice.
# Why ClickHouse (/docs/clickhouse/why-clickhouse/why-clickhouse)
Before any internals, it's worth being precise about what kind of
database this is — and, just as importantly, what kind of database it
deliberately is *not*. Most confusion and misuse of ClickHouse comes
from treating it like a faster Postgres, rather than a different tool
built for a different job.
## OLTP vs OLAP [#oltp-vs-olap]
Databases like Postgres and MySQL are optimized for **OLTP** — Online
Transactional Processing: many concurrent, small operations, each
touching a handful of rows (create an order, update a balance, look up
one user), where correctness and low per-operation latency matter
most.
ClickHouse is built for **OLAP** — Online Analytical Processing:
fewer, much larger operations, each scanning and aggregating millions
or billions of rows (total revenue by country last quarter, p99
latency per endpoint over the last hour), where throughput over huge
volumes matters most and a single query touching a lot of data is the
normal case, not the exception.
## Where it came from [#where-it-came-from]
ClickHouse was built inside Yandex to power Yandex.Metrica, a web
analytics product needing to compute arbitrary aggregate reports — on
demand, not from a fixed set of pre-built dashboards — over clickstream
data arriving at a rate of billions of events a day. It was
open-sourced in 2016. That origin still shapes the engine today:
everything is built around the assumption that queries are
unpredictable in shape but predictable in scale — always "scan a lot,
return a little."
## When it's the wrong choice [#when-its-the-wrong-choice]
Being honest about this matters more than the feature list. Don't
reach for ClickHouse for:
* **Single-row lookups and updates at high frequency** — the whole
architecture (immutable parts, background merges, sparse index) is
optimized for scanning ranges, not for "fetch/update exactly one
row" at OLTP-style request rates. See
[Parts & Background Merges](/clickhouse/parts-merges) and
[Mutations](/clickhouse/mutations) for why.
* **Systems that need real transactions** — there is no multi-statement
`BEGIN/COMMIT/ROLLBACK` with isolation guarantees across tables the
way an OLTP database provides them.
* **Strong referential integrity** — there are no foreign key
constraints. Joins work (see [Joins](/clickhouse/joins)), but
nothing stops an orphaned reference from being inserted.
* **Highly relational, join-heavy schemas** — joins are supported and
can be fast, but the engine and query planner are not as mature at
complex multi-way joins as a decades-old relational database.
Denormalizing toward wide tables is often the idiomatic ClickHouse
answer instead of normalizing further.
* **A queue or a cache** — despite the
[Kafka engine](/clickhouse/kafka-engine) and in-memory engines
existing, ClickHouse is not a substitute for a message broker or a
key-value cache; those have very different latency and consistency
guarantees.
The single most common ClickHouse misuse is treating it as a
drop-in replacement for an OLTP database and inserting one row at a
time from an application's request path. It will work at first and
then fail under load in a way that's confusing if you don't already
know why — covered in
[Async Inserts & Batching](/clickhouse/async-inserts).
## When it's the right choice [#when-its-the-right-choice]
* Event/log/metrics analytics at high ingest volume.
* Ad-hoc aggregate queries over huge, mostly-append-only datasets.
* Real-time dashboards that need sub-second answers over billions of
rows.
* Time-series data with a natural date/time-based access pattern.
Everything else in this module assumes you're building one of those —
and explains, concretely, how ClickHouse makes that fast.
# 11.2 Batch Processing in Distributed Systems (/docs/ddia/batch-processing/batch-processing-distributed-systems)
#### The organizing analogy — the distributed operating system [#the-organizing-analogy--the-distributed-operating-system]
#### 2.1 Distributed filesystems [#21-distributed-filesystems]
**The local filesystem stack, layer by layer:**
**Block sizes — and why they're so much bigger:**
| System | Block size |
| ------------------------------- | --------------- |
| **ext4** | **4,096 bytes** |
| **JuiceFS, many object stores** | **4 MB** |
| **HDFS** | **128 MB** |
> **Larger blocks mean LESS METADATA to keep track of, which MAKES A BIG DIFFERENCE ON PETABYTE-SIZED DATASETS. Larger blocks also LOWER THE OVERHEAD OF SEEKING TO A BLOCK RELATIVE TO READING IT.**
>
> **And unlike physical devices, DFSs DON'T need to write partial blocks: a 900 MB file with 128 MB blocks has SEVEN blocks of 128 MB and ONE BLOCK OF 4 MB.**
**Data nodes:** each machine runs a daemon exposing an API to read/write blocks as files on its local filesystem — **HDFS calls them DataNodes, GlusterFS calls them `glusterfsd`.**
**The protocol as the pluggable interface:**
> **Distributed filesystems must expose a protocol so batch systems can read and write. THIS PROTOCOL ACTS AS A PLUGGABLE INTERFACE; ANY DFS MAY BE USED SO LONG AS IT IMPLEMENTS THE PROTOCOL. For example, AMAZON S3's API HAS BEEN WIDELY ADOPTED by MinIO, Cloudflare R2, Tigris, Backblaze B2, and many others.**
**POSIX compatibility** via **FUSE** or **NFS**. *(NFS was originally developed to let multiple clients read/write on a SINGLE SERVER; more recently **Amazon EFS** and **Archil** provide NFS-compatible implementations that are FAR MORE SCALABLE — **clients still connect to one endpoint, but underneath these systems talk to distributed metadata services and data nodes.**)*
**DFS vs NAS/SAN — the shared-nothing point:**
> **Distributed filesystems are based on the SHARED-NOTHING principle, in contrast to the SHARED-DISK approach of NAS and SAN. Shared-disk storage uses a CENTRALIZED STORAGE APPLIANCE, often with CUSTOM HARDWARE and special network infrastructure such as FIBRE CHANNEL. The shared-nothing approach requires NO SPECIAL HARDWARE, only computers connected by a conventional datacenter network.**
>
> **Many DFSs are built on COMMODITY HARDWARE — less expensive but with HIGHER FAILURE RATES. To tolerate machine and disk failures, file blocks are REPLICATED on multiple machines. THIS ALSO ALLOWS SCHEDULERS TO MORE EVENLY DISTRIBUTE WORKLOADS, since they can execute a task on ANY node holding a replica of the task's input data.**
**Replication:** either **several copies** (Ch 6) or **erasure coding (Reed–Solomon), which allows lost data to be recovered with LOWER STORAGE OVERHEAD than full replication.** *(Similar to RAID; the difference is that here **file access and replication are done over a conventional datacenter network without special hardware.**)*
#### 2.2 Object stores [#22-object-stores]
**The URL anatomy:** `s3://my-photo-bucket/2025/04/01/birthday.png`
**Host = the BUCKET (globally unique name); the rest = the object's KEY (unique within its bucket).**
**The differences that bite:**
| | **Distributed filesystem** | **Object store** |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Mutability** | Files are mutable; `fopen`/`fseek` file handles | **OBJECTS ARE IMMUTABLE ONCE WRITTEN.** To update, **FULLY REWRITE with a `put`.** *(Azure Blob and S3 Express One Zone support appends; most others don't.)* **NO FILE HANDLE APIs** |
| **Directories** | Real | **DO NOT EXIST. The path structure is SIMPLY A CONVENTION — the slashes are PART OF THE KEY** |
| **Listing** | `ls` of one level | **A prefix list behaves like a RECURSIVE `ls -R`** — all objects starting with the prefix, **including subpaths** |
| **Empty directories** | Possible | **NOT POSSIBLE.** Delete everything under `.../2025/04/01` and `01` disappears from the listing of `.../2025/04`. **Common practice: create a ZERO-BYTE OBJECT to represent an empty directory** |
| **Hard/symbolic links, file locking** | Often supported | **Typically NOT supported** |
| **Renames** | **Atomic** | **NONATOMIC — copy to the new key, then delete the old. TO RENAME A "DIRECTORY" YOU MUST INDIVIDUALLY RENAME EVERY OBJECT WITHIN IT** |
| **Data locality** | **HDFS allows tasks to RUN ON THE MACHINE STORING A COPY of the file — reading without sending it over the network, saving bandwidth IF THE TASK'S CODE IS SMALLER THAN THE FILE** | **Storage and computation are SEPARATE. Might use more bandwidth, BUT MODERN DATACENTER NETWORKS ARE VERY FAST, so this is often acceptable — and it lets CPU/memory SCALE INDEPENDENTLY OF STORAGE** |
**Size/latency positioning:**
> **The key-value stores of Ch 4 are optimized for SMALL values (kilobytes) and FREQUENT, LOW-LATENCY reads/writes. Distributed filesystems and object stores are optimized for LARGE objects (megabytes to gigabytes) and LESS FREQUENT, LARGER reads. Recently, though, object stores have begun adding support for frequent, smaller I/O — S3 EXPRESS ONE ZONE now offers SINGLE-MILLISECOND LATENCY and a pricing model more similar to key-value stores.**
> ⚠️ **The line is blurry and dangerous:** FUSE drivers let you treat S3 as a filesystem; JuiceFS and Ceph offer both APIs. **However, their APIs, PERFORMANCE, AND CONSISTENCY GUARANTEES ARE VERY DIFFERENT. CARE MUST BE TAKEN to make sure they behave as expected, EVEN IF THEY SEEM TO IMPLEMENT THE REQUISITE APIs.**
#### 2.3 Distributed job orchestration [#23-distributed-job-orchestration]
**A job-start request carries:** number of tasks · memory/CPU/disk per task · a job identifier · access credentials · job parameters (input and output data) · **required hardware details such as GPUs or disk types** · **the location of the job's executable code.**
**Three components you'll find in nearly every orchestrator:**
| Component | Role | YARN | Kubernetes |
| --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------- | ----------------- |
| **Task executors** | Daemon on each node: **runs tasks, sends HEARTBEATS to signal liveness, tracks task status and resource allocation.** Retrieves the job's executable code, starts the task, monitors until it finishes or fails. **Also works with the OS for SECURITY AND PERFORMANCE ISOLATION — both use Linux CGROUPS** — preventing tasks from accessing data without permission or degrading other tasks | **NodeManager** | **kubelet** |
| **Resource manager** | Metadata about each node: **available hardware, task statuses, network location, node status.** Provides **a GLOBAL VIEW of cluster state.** ⚠️ **Its CENTRALIZED nature CAN LEAD TO BOTH SCALABILITY AND AVAILABILITY BOTTLENECKS** | State in **ZooKeeper** | State in **etcd** |
| **Scheduler** | Receives start/stop/status requests; **uses the request plus resource-manager state to decide WHICH TASKS RUN ON WHICH NODES** | ResourceManager | kube-scheduler |
| *(Application-specific sub-schedulers)* | For requirements the central scheduler can't know — e.g. **autoscaling read replicas at a query threshold.** They **work together with the central scheduler** | **ApplicationMasters** | **operators** |
##### Resource allocation — the genuinely hard part [#resource-allocation--the-genuinely-hard-part]
**The five-node, 160-core example with two jobs each wanting 100 cores:**
> **Now imagine hundreds or even MILLIONS of such requests. Finding an optimal solution seems intractable. IN FACT, THE PROBLEM IS NP-HARD — prohibitively slow to solve optimally for all but the smallest examples.**
>
> **In practice, schedulers therefore use HEURISTICS to make NONOPTIMAL BUT REASONABLE decisions:** FIFO · **dominant resource fairness (DRF)** · priority queues · capacity/quota-based scheduling · **bin-packing algorithms.**
##### Scheduling workflows [#scheduling-workflows]
> **A WORKFLOW (or DAG) of jobs: the output of one job becomes the input to one or more others.**
>
> ⚠️ **Terminology collision with Ch 5:** *"In 'Durable Execution and Workflows' we saw workflow engines offering durable execution of a sequence of steps, typically performing RPCs. In BATCH processing, 'workflow' has a DIFFERENT MEANING: a sequence of BATCH PROCESSES, each taking input data and producing output data, but NORMALLY NOT MAKING RPCs TO EXTERNAL SERVICES. Durable execution engines typically process LESS DATA PER REQUEST, though the line is somewhat fuzzy."*
**Three reasons a workflow is needed:**
1. **The output feeds several jobs MAINTAINED BY DIFFERENT TEAMS** → write it where all can read it, and schedule consumers on data update or their own schedule
2. **Transfer data between processing TOOLS** — a Spark job writes HDFS, a Python script triggers a Trino SQL query, which outputs to S3
3. **Multiple internal stages** — **if one stage needs data sharded by one key and the next by a different key, the first stage can output data sharded the way the second requires**
**Coupling choice — pipe vs file:**
> **Orchestration-framework schedulers (YARN's ResourceManager, Spark's built-in scheduler) DO NOT MANAGE ENTIRE WORKFLOWS; they schedule PER JOB. To handle dependencies BETWEEN job executions, WORKFLOW SCHEDULERS were developed: AIRFLOW, DAGSTER, PREFECT.**
>
> **Workflows of 50 TO 100 JOBS are common in many data pipelines, and in a large organization MANY TEAMS MAY BE RUNNING JOBS THAT READ ONE ANOTHER'S OUTPUT ACROSS MANY SYSTEMS. TOOL SUPPORT IS IMPORTANT FOR MANAGING SUCH COMPLEX DATAFLOWS.**
##### Handling faults — and why batch has it easy [#handling-faults--and-why-batch-has-it-easy]
**Two reasons a task doesn't finish:** hardware faults / network interruptions (Ch 2, Ch 9), **and deliberate PREEMPTION by the scheduler.**
> **Preemption is particularly useful with MULTIPLE PRIORITY LEVELS: low-priority tasks are CHEAPER and run whenever there's spare capacity, but RISK BEING PREEMPTED AT ANY MOMENT.** These are **spot instances** (EC2), **spot virtual machines** (Azure), **preemptible instances** (Google Cloud).
>
> **Batch processing is often not time-sensitive, so it's WELL SUITED to spot instances — using spare resources that would otherwise be idle, INCREASING CLUSTER UTILIZATION. However, THOSE TASKS ARE MORE LIKELY TO BE KILLED, BECAUSE PREEMPTIONS OCCUR MORE FREQUENTLY THAN HARDWARE FAULTS.**
> ### **Since batch jobs REGENERATE THEIR OUTPUT FROM SCRATCH every time, TASK FAILURES ARE EASIER TO HANDLE THAN IN ONLINE SYSTEMS: delete the partial output from the failed execution and reschedule the task on another machine.** [#since-batch-jobs-regenerate-their-output-from-scratch-every-time-task-failures-are-easier-to-handle-than-in-online-systems-delete-the-partial-output-from-the-failed-execution-and-reschedule-the-task-on-another-machine]
>
> **It would be WASTEFUL to rerun the ENTIRE job for one task failure. MapReduce and successors therefore KEEP THE EXECUTION OF PARALLEL TASKS INDEPENDENT, so they can RETRY AT THE GRANULARITY OF AN INDIVIDUAL TASK.**
**Three approaches to intermediate-data fault tolerance:**
| System | Approach | Trade-off |
| ------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------- |
| **MapReduce** | **ALWAYS writes intermediate data back to the DFS, and WAITS for the writing task to complete successfully before others read it** | **Works even where preemption is common — BUT MEANS A LOT OF WRITES TO THE DFS, WHICH CAN BE INEFFICIENT** |
| **Spark** | **Keeps intermediate data IN MEMORY (spilling to local disk if it won't fit) and writes ONLY THE FINAL RESULT to the DFS. TRACKS HOW THE INTERMEDIATE DATA WAS COMPUTED, allowing RECOMPUTATION if lost** (lineage) | Much faster; recomputation cost on loss |
| **Flink** | **Periodic CHECKPOINTING of a snapshot of tasks** | Different trade-off again |
***
# 11.3 Batch Processing Models (/docs/ddia/batch-processing/batch-processing-models)
#### 3.1 MapReduce [#31-mapreduce]
**The four steps — and how they map exactly onto the Unix pipeline:**
**Why the functional-programming heritage matters:**
> **Lisp introduced `map` and `reduce` (or `fold`) as higher-order functions on lists.**
>
> ### **The functional principle of AVOIDING MUTABLE STATE IS WHAT ENABLES PARALLEL EXECUTION. As every call depends ONLY on the data the framework EXPLICITLY PASSES, the framework is free to RUN INDEPENDENT CALLS IN PARALLEL ON DIFFERENT NODES — AND IF A TASK FAILS, TO CALL THE MAPPER OR REDUCER AGAIN WITH THE SAME INPUT ON ANOTHER NODE.** [#the-functional-principle-of-avoiding-mutable-state-is-what-enables-parallel-execution-as-every-call-depends-only-on-the-data-the-framework-explicitly-passes-the-framework-is-free-to-run-independent-calls-in-parallel-on-different-nodes--and-if-a-task-fails-to-call-the-mapper-or-reducer-again-with-the-same-input-on-another-node]
**Two damning limitations:**
* **Implementing a complex job with raw MapReduce APIs is QUITE LABORIOUS — ANY JOIN ALGORITHMS WOULD NEED TO BE IMPLEMENTED FROM SCRATCH**
* **MapReduce is QUITE SLOW compared to modern batch processors. One reason: ITS FILE-BASED I/O PREVENTS JOB PIPELINING** — processing output in a downstream job before the upstream job is complete
#### 3.2 Dataflow engines (Spark, Flink) [#32-dataflow-engines-spark-flink]
> **They handle AN ENTIRE WORKFLOW AS ONE JOB, rather than breaking it into independent subjobs. Since they EXPLICITLY MODEL THE FLOW OF DATA through several processing stages, they are known as DATAFLOW ENGINES.**
>
> **Like MapReduce they support a low-level record-at-a-time API, but they also offer HIGHER-LEVEL OPERATORS such as JOIN and GROUP BY. They parallelize by SHARDING inputs and COPY THE OUTPUT OF ONE TASK OVER THE NETWORK to become another's input. UNLIKE MAPREDUCE, OPERATORS NEED NOT TAKE THE STRICT ROLES OF ALTERNATING MAP AND REDUCE — they can be ASSEMBLED IN MORE FLEXIBLE WAYS.**
**The six concrete advantages over MapReduce:**
#### 3.3 Shuffling — the foundational algorithm [#33-shuffling--the-foundational-algorithm]
> ⚠️ **SHUFFLE IS NOT RANDOM. "When you shuffle a deck of cards, you end up with a RANDOM order. In contrast, the shuffle we're talking about PRODUCES A SORTED ORDER, WITH NO RANDOMNESS."**
>
> **A distributed sorting algorithm where BOTH THE INPUT AND THE OUTPUT ARE SHARDED. Batch processors must sort datasets PETABYTES in size.**
> **Modern dataflow engines and cloud warehouses are more sophisticated: BigQuery has optimized its shuffle to KEEP DATA IN MEMORY and to write to EXTERNAL SORTING SERVICES — speeding up shuffling and REPLICATING SHUFFLED DATA TO PROVIDE RESILIENCE.**
#### 3.4 Joins and grouping — the sort-merge join [#34-joins-and-grouping--the-sort-merge-join]
**The scenario:** activity events (clickstream) on the left, a user database on the right. **In star-schema terms, the event log is the FACT TABLE and the user database is one of the DIMENSIONS.**
#### 3.5 Query languages [#35-query-languages]
> **With the problem of physically operating batch processes at scale CONSIDERED MORE OR LESS SOLVED, attention has turned to IMPROVING THE PROGRAMMING MODEL.**
>
> **MapReduce, dataflow engines, and cloud warehouses have all embraced SQL AS THE LINGUA FRANCA. It's a natural fit: legacy warehouses used SQL, analytics and ETL tools already support it, and ALL DEVELOPERS AND ANALYSTS KNOW IT.**
**Two payoffs:**
1. **Human:** less code, **and INTERACTIVE USE** — an efficient and natural way for **business analysts, product managers, sales and finance teams** to explore data. **SQL support has made distributed batch systems suitable for EXPLORATORY QUERIES.**
2. **Machine:** **the translation from query → syntax tree → physical operators ALLOWS THE ENGINE TO OPTIMIZE. Hive, Trino, Spark, and Flink have COST-BASED OPTIMIZERS that analyze the properties of join inputs and AUTOMATICALLY DECIDE WHICH ALGORITHM IS MOST SUITABLE. Optimizers MIGHT EVEN CHANGE THE ORDER OF JOINS so the amount of INTERMEDIATE STATE IS MINIMIZED.**
**Niche languages:** **Apache Pig** (relational operators, pipelines specified **step by step rather than as one big SQL query**), **Morel** (a modern language influenced by Pig), JSON query languages (**jq, JMESPath, JSONPath**), and graph languages (**Apache TinkerPop's Gremlin**).
#### 3.6 Batch processing and cloud warehouses converge [#36-batch-processing-and-cloud-warehouses-converge]
**But they haven't fully merged. What SQL/warehouses still struggle with:**
* **ITERATIVE GRAPH ALGORITHMS such as PageRank**, complex ML tasks
* **AI DATA PROCESSING — nonrelational and MULTIMODAL data such as images, video, audio**
* **ROW-BY-ROW COMPUTATION is less efficient with column-oriented storage**
* **Cloud warehouses TEND TO BE MORE EXPENSIVE. It can be MORE COST-EFFICIENT to run large jobs in Spark or Flink**
> **The decision often comes down to COST, CONVENIENCE, EASE OF IMPLEMENTATION, AND AVAILABILITY. Most large enterprises have MANY data processing systems, giving them flexibility. SMALLER COMPANIES OFTEN GET BY WITH JUST ONE.**
#### 3.7 DataFrames in a distributed setting [#37-dataframes-in-a-distributed-setting]
**Why they exist here:** *"Data scientists wanted to interact with the LARGE DATASETS found in batch environments USING THE DATAFRAME APIs THEY WERE USED TO, since SQL AND MAPREDUCE ARE NOT WELL SUITED TO THEIR NEEDS."*
> ⚠️ **Two traps:**
>
> **① "LOCAL DATAFRAMES ARE USUALLY INDEXED AND ORDERED, WHILE DISTRIBUTED DATAFRAMES ARE GENERALLY NOT. THIS CAN LEAD TO PERFORMANCE SURPRISES WHEN MIGRATING TO BATCH FRAMEWORKS."**
>
> **② "PANDAS EXECUTES OPERATIONS IMMEDIATELY when DataFrame methods are called; SPARK FIRST TRANSLATES ALL THE API CALLS INTO A QUERY PLAN AND RUNS QUERY OPTIMIZATION before executing."** *(Eager vs lazy.)*
**Hybrid execution:** **Daft supports BOTH client- and server-side computation — smaller in-memory operations on the client, larger datasets on a server. Columnar formats such as APACHE ARROW offer a UNIFIED DATA MODEL that both execution engines can share.**
***
# 11.1 Batch Processing with Unix Tools (/docs/ddia/batch-processing/batch-processing-unix-tools)
**The NGINX access log line** (one line, wrapped for readability):
```txt
216.58.210.78 - - [27/Jun/2025:17:55:11 +0000] "GET /css/typography.css HTTP/1.1"
200 3377 "https://martin.kleppmann.com/" "Mozilla/5.0 (Macintosh; …) Chrome/137.0.0.0 …"
```
**Format:** `$remote_addr - $remote_user [$time_local] "$request" $status $body_bytes_sent "$http_referer" "$http_user_agent"`
> **Though log parsing might seem contrived, it's A CRITICAL PART OF THE OPERATIONS OF MANY MODERN TECHNOLOGY COMPANIES — used for everything from AD PIPELINES to PAYMENT PROCESSING. Indeed, IT WAS A DRIVING FORCE BEHIND THE RAPID ADOPTION OF MAPREDUCE AND THE "BIG DATA" MOVEMENT.**
#### 1.1 The five-command pipeline [#11-the-five-command-pipeline]
```bash
cat /var/log/nginx/access.log | # read the log
awk '{print $7}' | # 7th field = the requested URL
sort | # so identical URLs become ADJACENT
uniq -c | # collapse adjacent duplicates, -c = count them
sort -r -n | # sort by the leading NUMBER, REVERSED
head -n 5 # top 5
```
```txt
4189 /favicon.ico
3631 /2016/02/08/how-to-do-distributed-locking.html
2124 /2020/11/18/distributed-systems-and-elliptic-curves.html
1369 /
915 /css/typography.css
```
> **It will process GIGABYTES of log files IN A MATTER OF SECONDS, and you can easily modify the analysis.** Omit CSS files: change awk to `$7 !~ /\.css$/ {print $7}`. Count top client IPs: `{print $1}`.
>
> **Many data analyses can be done in a few minutes using a combination of `awk`, `sed`, `grep`, `sort`, `uniq`, and `xargs`, and THEY PERFORM SURPRISINGLY WELL.**
#### 1.2 The crucial contrast: sorting vs in-memory aggregation [#12-the-crucial-contrast-sorting-vs-in-memory-aggregation]
**The Python equivalent** keeps an **in-memory hash table** `url → count`. **The Unix pipeline has NO hash table — it relies on SORTING a list in which multiple occurrences are simply repeated.**
> **GNU Coreutils `sort` AUTOMATICALLY handles larger-than-memory datasets by SPILLING TO DISK and AUTOMATICALLY PARALLELIZES sorting across multiple CPU cores. The simple chain of Unix commands EASILY SCALES TO LARGE DATASETS without running out of memory. THE BOTTLENECK IS LIKELY TO BE THE RATE AT WHICH THE INPUT FILE CAN BE READ FROM DISK.**
>
> **A limitation of Unix tools is that THEY RUN ON A SINGLE MACHINE — and that's where distributed batch processing frameworks come in.**
***
# 11.4 Batch Use Cases (/docs/ddia/batch-processing/batch-use-cases)
> **You'll find batch jobs WHEREVER THERE'S A LOT OF DATA AND DATA FRESHNESS ISN'T IMPORTANT. This might sound limiting, but IT TURNS OUT THAT A SIGNIFICANT AMOUNT OF DATA PROCESSING TASKS FIT THIS MODEL:**
>
> * **Accounting and inventory reconciliation** — verifying transactions line up with bank accounts and inventory
> * **Demand forecasting in manufacturing** — a periodic batch job
> * **Ecommerce, media, and social media train their RECOMMENDATION MODELS with batch jobs**
> * ### **"Many financial systems are batch-based; for example, THE US BANKING NETWORK RUNS ALMOST ENTIRELY ON BATCH JOBS."** [#many-financial-systems-are-batch-based-for-example-the-us-banking-network-runs-almost-entirely-on-batch-jobs]
#### 4.1 ETL / ELT [#41-etl--elt]
**Why batch fits so well:**
* **The PARALLEL nature suits transformation, much of which is "EMBARRASSINGLY PARALLEL"** — filtering, projecting fields, and many common warehouse transformations **can all be done in parallel**
* **Robust workflow schedulers** make it easy to schedule, orchestrate, and debug. **Schedulers RETRY jobs to mitigate transient issues; a job that fails repeatedly is MARKED AS FAILED, which helps developers EASILY SEE WHICH JOB STOPPED WORKING.** Airflow ships **built-in source, sink, and query operators for MySQL, PostgreSQL, Snowflake, Spark, Flink, and dozens of others**
* **Easy to troubleshoot** — *"Failed files can be easily inspected to see what went wrong, and ETL batch jobs can be FIXED AND RERUN."* If an input file lacks a field the job needs, **engineers can easily spot it and update either the transformation or the job that produced the input**
**The organizational shift:**
> **"Data pipelines USED TO BE MANAGED BY A SINGLE DATA ENGINEERING TEAM, as it was considered UNFAIR to ask product teams to write and manage complex batch pipelines. Recently, improvements in batch processing models and METADATA MANAGEMENT have made it MUCH EASIER FOR ENGINEERS ACROSS AN ORGANIZATION TO CONTRIBUTE TO AND MANAGE THEIR OWN PIPELINES."** — **data mesh**, **data contracts**, **data fabric** provide **standards and tools to help teams SAFELY PUBLISH THEIR DATA for consumption by anybody in the organization.**
>
> **"Many batch ETL jobs now run ON THE SAME SYSTEMS as the analytical queries that read their output — SparkSQL, Trino, or DuckDB. Such an architecture FURTHER BLURS THE LINE BETWEEN APPLICATION ENGINEERING, DATA ENGINEERING, ANALYTICS ENGINEERING, AND BUSINESS ANALYSIS."**
#### 4.2 Analytics — the lakehouse [#42-analytics--the-lakehouse]
> **Analysts write SQL that executes atop a query engine reading from and writing to a DFS or object store. Table metadata is managed with TABLE FORMATS such as Apache Iceberg and CATALOGS such as Unity. THIS ARCHITECTURE IS KNOWN AS A DATA LAKEHOUSE.**
| Query style | Characteristics |
| ------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Pre-aggregation** | Data rolled up into **OLAP cubes or data marts** (Ch 4 §7.7). **Pre-aggregated data is queried in the warehouse OR PUSHED TO PURPOSE-BUILT REAL-TIME OLAP SYSTEMS such as Druid or Pinot.** Runs at a **scheduled interval**, managed by workflow schedulers |
| **Ad hoc** | **Response times are IMPORTANT here. Analysts run queries ITERATIVELY as they get responses and learn more about the data. Fast query execution REDUCES WAITING TIMES** |
**BI integration:** SQL support enables **Tableau, Power BI, Looker, Apache Superset** — *"Tableau offers SparkSQL and Presto connectors; Superset supports Trino, Hive, Spark SQL, Presto."*
#### 4.3 Machine learning [#43-machine-learning]
| Use | What goes in / comes out |
| ----------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Feature engineering** | **Raw data filtered and transformed into data models can train on. Predictive models often need NUMERIC data, so engineers must transform TEXT OR DISCRETE VALUES into the required format** |
| **Model training** | **Training data is the INPUT; the WEIGHTS of the trained model are the OUTPUT** |
| **Batch inference** | **Predictions in bulk when datasets are large and real-time results aren't required — including EVALUATING THE MODEL'S PREDICTIONS ON A TEST DATASET** |
*(Tooling: **Spark MLlib**, **Flink FlinkML** — feature engineering tools, statistical functions, classifiers.)*
**Graph processing:** recommendation engines and ranking systems use it heavily.
> **Many graph algorithms are expressed by TRAVERSING ONE EDGE AT A TIME, joining one vertex with an adjacent vertex TO PROPAGATE SOME INFORMATION, and REPEATING UNTIL A CERTAIN CONDITION IS MET — until there are no more edges to follow, or until a metric CONVERGES.**
>
> **The BULK SYNCHRONOUS PARALLEL (BSP) model has become popular — implemented by Apache Giraph, Spark's GraphX, Flink's Gelly. Also known as the PREGEL MODEL, after Google's Pregel paper.**
**LLM data preparation — batch's newest big job:**
*(And **notebooks** — Jupyter, Hex — where cells of Markdown/Python/SQL execute sequentially, **many using batch processing via DataFrame APIs or SQL.**)*
#### 4.4 Serving derived data — and the anti-pattern [#44-serving-derived-data--and-the-anti-pattern]
> **Batch jobs build precomputed datasets — product recommendations, user-facing reports, ML features — typically served from a production database, key-value store, or search engine. THE PRECOMPUTED DATA NEEDS TO MAKE ITS WAY FROM THE BATCH PROCESSOR'S STORAGE BACK INTO THE DATABASE SERVING LIVE TRAFFIC.**
**❌ THE ANTI-PATTERN: writing directly to the production database from inside the job.**
**✔ Solution A — push to a stream (Kafka):**
**✔ Solution B — build the database in the job and bulk-load it:**
> **Build a BRAND-NEW DATABASE INSIDE the batch job and BULK-LOAD those files directly.** Tools: **TiDB's Lightning, Apache Pinot's Hadoop import jobs, RocksDB's SST bulk-import API.**
>
> **VERY FAST, and makes it easier for systems to ATOMICALLY SWITCH BETWEEN DATASET VERSIONS. On the other hand, IT CAN BE CHALLENGING TO INCREMENTALLY UPDATE datasets from jobs that build brand-new databases.**
>
> **It's common to take a HYBRID approach when both bootstrapping and incremental loads are needed** — **Venice supports hybrid stores allowing batch row-based updates AND full dataset swaps.**
***
# 11.7 Decision cheat sheet (/docs/ddia/batch-processing/decision-cheat-sheet)
**Batch or stream?**
Batch when **data freshness isn't important** and the input is **bounded**. Stream when you need second-level latency or the input is **unbounded** (Ch 12). Note that "the US banking network runs almost entirely on batch jobs" — **freshness matters far less often than people assume.**
**Single machine or distributed?**
**Try the Unix pipeline / DuckDB / Polars first.** GNU `sort` spills to disk and parallelizes across cores; the bottleneck is disk read rate. Go distributed only when data exceeds one machine's disk or the job exceeds your time budget. (Ch 1: *more nodes are not always faster*.)
**Hash aggregation or sort?**
**Hash** when the number of *distinct keys* fits in memory — note the working set depends on **distinct keys, not record count**. **Sort** when it doesn't; sorting degrades gracefully to disk with sequential I/O.
**HDFS or object store?**
**Object store** by default now — decoupled scaling, no NameNode, better durability economics. **HDFS only if data locality genuinely dominates** your bandwidth budget. **Either way, use a table format (Iceberg/Delta) so you don't depend on atomic rename.**
**MapReduce, dataflow engine, or warehouse SQL?**
**Never raw MapReduce for new work.** **Warehouse SQL** for relational analytics your analysts will touch. **Spark/Flink** for large jobs where warehouse cost bites, for non-relational/multimodal data, for iterative graph and ML work, and for anything awkward in SQL.
**How do I get batch output into production?**
**Spot instances or on-demand?**
**Spot for batch** — it's exactly what batch is good at (not time-sensitive, restartable at task granularity, uses otherwise-idle capacity). But **budget for preemptions being more frequent than hardware faults**, and make sure intermediate-data loss doesn't cascade into whole-stage recomputation.
***
# 11.11 Forward links (/docs/ddia/batch-processing/forward-links)
| Concept here | Where it's developed |
| --------------------------------------------------------------------- | --------------------------------------------- |
| Unbounded inputs; reacting in seconds instead of hours | **Ch 12** — Stream Processing |
| Pushing batch output into Kafka topics | **Ch 12** |
| Composing batch and stream deliberately | **Ch 13** — A Philosophy of Streaming Systems |
| Columnar storage and vectorized execution | **Ch 4** |
| LSM segment merging (the same algorithm as shuffle sorting) | **Ch 4** §2 |
| Sharding by hash of key (how shuffle assigns reducers) | **Ch 7** §3.2 |
| Star schemas — the fact/dimension join | **Ch 3** §1.7 |
| Coordination services holding cluster state | **Ch 10** §4 |
| Avro and Parquet as batch file formats | **Ch 5** |
| Read-committed isolation (the model for hiding incomplete job output) | **Ch 8** §3.1 |
# 11. Batch Processing (/docs/ddia/batch-processing)
> "A system cannot be successful if it is too strongly influenced by a single person. Once the initial design is complete and fairly robust, the real test begins as people with many different viewpoints undertake their own experiments." — Donald Knuth
**The two families of data processing:**
| | **Online systems** | **Offline systems (batch)** |
| -------------- | -------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| Shape | **You ask for something, and the system tries to give you an answer AS QUICKLY AS POSSIBLE** | **A job takes READ-ONLY input data and produces output GENERATED FROM SCRATCH EVERY TIME IT RUNS** |
| Primary metric | **Response time** | **THROUGHPUT — how much data per unit of time** |
| Duration | milliseconds | **Minutes, hours, or even DAYS.** Often scheduled periodically |
| Faults | **Require fault tolerance for high availability** | **Some abort and restart the WHOLE job; others tolerate node crashes** |
> **A batch job typically DOES NOT MUTATE DATA the way a read/write transaction would. The output is DERIVED from the input. IF YOU DON'T LIKE THE OUTPUT, YOU CAN DELETE IT, ADJUST THE JOB'S LOGIC, AND RUN THE JOB AGAIN.**
#### The four benefits of immutable inputs and no side effects [#the-four-benefits-of-immutable-inputs-and-no-side-effects]
**Two honest costs:**
* **With most frameworks, output can be processed by other jobs ONLY AFTER THE WHOLE JOB FINISHES**
* **ANY CHANGE TO THE INPUT DATA — EVEN A SINGLE BYTE — REQUIRES THE JOB TO REPROCESS THE ENTIRE INPUT DATASET**
#### The historical arc — and where MapReduce stands now [#the-historical-arc--and-where-mapreduce-stands-now]
***
# 11.6 Production failure catalog for this chapter (/docs/ddia/batch-processing/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ---------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
| Bad data written to the DB; rolling back the code doesn't fix it | **Online systems lack human fault tolerance** — batch's rerun-from-immutable-input does |
| One byte changed → the whole dataset reprocessed | Batch's fundamental inefficiency (→ stream processing, Ch 12) |
| Job succeeded but produced no/garbage data | **Success ≠ correctness**; no monitoring job comparing to the previous run |
| 199 tasks finish in seconds, one runs for hours | **Data skew** on the shuffle key |
| Shuffle spills to disk and the job takes 10× longer | Working set exceeds executor memory |
| Job restarts from scratch after one node dies | Intermediate data not checkpointed / lineage lost |
| Whole upstream stage recomputed repeatedly on spot instances | **Preemption + in-memory shuffle output** with no external shuffle service |
| Output is 40,000 tiny Parquet files | No repartition before write → **small-file problem** |
| Job "commit" step takes hours on S3 | **Rename-based commit protocol** on a store with nonatomic rename |
| NameNode out of memory | **Millions of small files** — metadata is per-file |
| List-then-process silently misses new files | Object-store listing semantics assumed to be a directory listing |
| Production DB melts when the nightly job runs | **Writing directly to the production database** from parallel tasks |
| Partially-complete job's output visible to users | External side effects break the **all-or-nothing** guarantee |
| Duplicate records after a task retry | Same — a restarted task duplicates external writes |
| Airflow scheduler falls behind | **Heavy top-level code** in DAG files re-executed on every parse |
| A backfill takes down the source database | Hundreds of concurrent DagRuns |
| Job processed yesterday's data — or tomorrow's | **`execution_date` / timezone** semantics |
| Cluster deadlocks with everything half-scheduled | **Gang scheduling** holding partial allocations |
| Large jobs never run | **Starvation** by a stream of small jobs |
| Pandas code is instant, Spark code hangs on `.count()` | **Eager vs lazy** evaluation |
| `collect()` kills the driver | Pulling a distributed dataset into one JVM |
***
# 11.9 Self-test (/docs/ddia/batch-processing/self-test)
Give the four benefits of treating batch inputs as immutable with no side effects. Which one do read/write databases fundamentally lack, and why?
State the two costs of batch processing. Which one motivates stream processing?
Walk through the five-command Unix pipeline and say what each stage contributes. Why is the first `sort` there?
What is a job's "working set"? Why does it depend on distinct keys rather than record count?
When does sorting beat an in-memory hash table, and what property of mergesort makes it work on disk?
Map the three components of a single machine onto the three components of a distributed batch framework.
Why are DFS blocks 128 MB rather than 4 KB? Give two reasons.
What plays the role of the VFS in a distributed filesystem? Give a concrete example of why that matters commercially.
Contrast shared-nothing DFS with NAS/SAN on hardware, cost, and failure handling.
Give five ways object stores differ from filesystems that would break naive code.
Why does HDFS allow computation on the node holding the data, and why is that less compelling now?
Name the three orchestrator components and what each is responsible for. Where is cluster state stored in YARN and in Kubernetes?
Work through the 160-core, two-job scheduling example. Define gang scheduling, starvation, and preemption, and give the downside of each.
Why is optimal scheduling intractable, and what do real schedulers do instead?
In batch processing, what does "workflow" mean, and how does it differ from Ch 5's usage?
Contrast pipe-style coupling with file-style coupling between jobs. Which is more typical, and why?
Why doesn't Spark's own scheduler manage workflows? What tools do?
What are spot instances, and why is batch processing unusually well suited to them? What's the catch?
Why are task failures easier to handle in batch than online? At what granularity is work retried?
Compare MapReduce, Spark, and Flink on how they handle intermediate data and its loss.
Give the four steps of MapReduce and map each onto the Unix pipeline. Which step do you *not* write?
Why does avoiding mutable state enable both parallelism and retries?
Give the six advantages of dataflow engines over MapReduce. Which one directly fixes MapReduce's inability to pipeline?
Why is "shuffle" a misleading name? Describe the full data path from mapper output to reducer output.
What determines the number of map tasks? What determines the number of reduce tasks?
Explain a sort-merge join. What is secondary sort, and what does it buy the reducer?
Why does a reducer in a sort-merge join need only one user record in memory and no network requests?
Give two machine-level reasons SQL improves batch jobs, not just human-level ones.
Name three workloads that are hard for cloud data warehouses and three ways warehouses and batch frameworks have converged.
Give two ways distributed DataFrames surprise people coming from Pandas.
List four real-world domains where batch dominates. Which one is the most surprising?
Give three reasons batch fits ETL particularly well.
What is a data lakehouse? Contrast pre-aggregation queries with ad hoc queries.
Name three ML uses of batch processing and describe the BSP/Pregel model.
List the three reasons you should not write directly to a production database from a batch job. Which one is about correctness rather than performance?
Give the four benefits of pushing batch output through a stream. What problem does streaming *not* solve, and what's the fix?
When is bulk-loading a purpose-built database file better than streaming, and what does it make hard?
you must produce daily personalized recommendations for 50 million users from 2 TB/day of clickstream data, serve them at p99 \< 20 ms, and be able to fix a bad model version within one hour of discovering it. Design the pipeline: storage layer, processing engine, workflow orchestration, fault-tolerance strategy, and the mechanism for getting output into the serving path. For each choice, name the specific trade-off from this chapter you're relying on, and state what happens when a single task fails, when the whole job produces bad output, and when the serving database is at capacity.
# 11.5 Technology deep dives (/docs/ddia/batch-processing/technology-deep-dives)
***
#### 5.1 Apache Spark [#51-apache-spark]
**Problem it solves.** Run an entire multi-stage workflow as one optimized job, keeping intermediate data in memory, with fault tolerance that doesn't require writing every stage to a replicated filesystem.
**Why wasn't MapReduce enough?** §3.1's two limits plus §3.2's six: mandatory sort between every stage, no operator fusion, no locality planning, all intermediate state written to the DFS (replicated, on disk, on every replica), no pipelining, and a new JVM per task.
**How it works internally.** The **RDD/DataFrame** is the unit; transformations build a **lazy DAG**, and only an *action* triggers execution. **Catalyst** optimizes the logical plan (predicate pushdown, column pruning, join reordering) and **Tungsten** generates whole-stage code (Ch 4's query compilation). The DAG is cut into **stages at shuffle boundaries**; within a stage, narrow transformations are **fused into one task** (advantage ②). **Lineage** — the record of *how* each partition was computed — is what lets Spark **recompute lost intermediate data instead of replicating it** (§2.3's fault-tolerance table). Adaptive Query Execution re-plans mid-job using actual shuffle statistics.
**Deployment.** On YARN, Kubernetes, or standalone; executors sized as `cores × memory`; data in an object store via Parquet/Iceberg; driven by Airflow/Dagster.
**Monitoring.**
* **Task-duration skew within a stage** — the single most diagnostic metric. 199 tasks in 10 s and one in 4 hours means a skewed key.
* **Shuffle read/write bytes** and **spill (memory and disk)** — spill volume is your "this doesn't fit" signal
* **GC time as a fraction of task time** (>10% means executor memory is wrong)
* **Stage retry count** and **`FetchFailedException`** rate (lost shuffle files → whole-stage recomputation)
* **Executor loss reason** — distinguishing preemption (spot instances, §2.3) from OOM matters enormously
**Scaling.** More executors helps until shuffle becomes the bottleneck; then the lever is **reducing shuffle volume** (broadcast joins, better partitioning, pre-aggregation) rather than adding machines.
**What actually breaks.**
* **Data skew.** One key holds 40% of the rows; one task runs for hours while the cluster idles. Fixes: salting the key, AQE skew join handling, or a broadcast join.
* **OOM on the driver** from `collect()` on a large DataFrame — the classic mistake of pulling a distributed dataset into one JVM.
* **Cascading recomputation** when an executor holding shuffle output dies: every downstream task that needed that block fails, and Spark recomputes the whole upstream stage. On spot instances this can dominate runtime. (External shuffle service / shuffle-service-on-object-storage exists precisely for this.)
* **The small-file problem** on write — `repartition` before writing, or you produce 10,000 tiny Parquet files (Ch 4 §9.3).
* **Eager-vs-lazy confusion** for people arriving from Pandas (§3.7) — a chain of transformations that "runs instantly" and then takes an hour on the first action.
* **`spark.sql.shuffle.partitions = 200`** (the default) being wildly wrong for both tiny and huge jobs.
***
#### 5.2 Apache Airflow (workflow orchestration) [#52-apache-airflow-workflow-orchestration]
**Problem it solves.** Schedule and manage the **dependency graph between jobs** — which the per-job schedulers (YARN, Spark's own) explicitly do not do (§2.3).
**Why wasn't cron enough?** Cron has no notion of dependencies, no backfill, no retry semantics, no visibility into which of 100 jobs failed, and no way to express "run when all three upstream jobs succeed."
**How it works internally.** A **DAG** of tasks defined in Python. The **scheduler** parses DAG files, creates **DagRuns** per schedule interval, and marks tasks runnable when upstream dependencies are met; **executors** (Celery/Kubernetes) run them; state lives in a metadata database. **Operators** encapsulate integrations — §4.1's "built-in source, sink, and query operators for MySQL, PostgreSQL, Snowflake, Spark, Flink, and dozens of others."
**Monitoring.** **Task duration trend per task** (a slowly growing job is a future incident); **SLA misses**; scheduler loop latency and **DAG parse time** (heavy top-level code in DAG files silently throttles the whole scheduler); queued-vs-running task counts; **retry counts by task** — distinguishing transient from systematic failure, exactly the distinction §4.1 says makes debugging easy.
**Scaling.** More workers for task throughput; but the **scheduler and metadata DB are the real ceiling.** Keep DAG files light.
**What actually breaks.**
* **Top-level code in DAG files** — an API call or heavy import at module scope runs on *every parse*, every few seconds, for every DAG.
* **Tasks that aren't idempotent**, meeting Airflow's retry behaviour → duplicated side effects. This is why §1's "batch jobs avoid side effects" matters operationally.
* **Backfills** running hundreds of DagRuns concurrently and melting a source database.
* **Timezone and `execution_date` semantics** — the single most common source of "the job ran but processed the wrong day's data."
* **Using Airflow as the compute engine** (heavy work inside a PythonOperator) rather than as an orchestrator.
* **A silently succeeding job that produced no data** — success ≠ correctness; this is why §1's third benefit (monitoring jobs comparing to the previous run) exists.
***
#### 5.3 HDFS vs S3 as the batch storage layer [#53-hdfs-vs-s3-as-the-batch-storage-layer]
**Problem each solves.** Store petabyte datasets durably and read them at aggregate bandwidths a single machine can't reach.
**Why HDFS first?** **Data locality** (§2.2) — run the task on the node holding the block, so **you ship the code to the data, not the data to the code.** In 2006, when datacenter networks were slow relative to disk, this was decisive.
**Why S3 now?** §2.2: **"modern datacenter networks are very fast, so this is often acceptable" — and decoupling lets CPU and storage scale independently.** Plus no NameNode to operate, no rebalancing, and vastly better durability economics.
**How they differ operationally — the list that causes real bugs:**
| | HDFS | S3 |
| ---------- | ----------------------------------------------- | --------------------------------------------------------- |
| Rename | **Atomic** | **Copy + delete, nonatomic, O(n) for a "directory"** |
| Listing | Directory listing | **Recursive prefix scan; paginated; eventually complete** |
| Append | Supported | **Generally not** |
| Metadata | **NameNode** (SPOF, memory-bound on file count) | Managed, effectively unbounded |
| Locality | **Yes** | No |
| Cost model | Cluster you run | **Per-request fees + storage; batching matters** |
**Monitoring.** HDFS: **NameNode heap and total file/block count** (the small-file problem is a *NameNode memory* problem here, not just a query-planning one), under-replicated blocks, DataNode volume failures. S3: **request rate and 503 `SlowDown`** per prefix, first-byte latency, incomplete multipart uploads, cost per prefix.
**What actually breaks.**
* **The commit protocol.** Spark/Hadoop output committers historically relied on **atomic rename** to publish results. On S3, rename isn't atomic, so the naive committer is both **slow and unsafe** — hence the S3A committers and, better, **table formats (Iceberg/Delta) whose atomic metadata commit replaces rename entirely.** This is *the* classic HDFS→S3 migration failure.
* **Millions of small files** killing the NameNode (HDFS) or query planning and request cost (S3).
* **Eventual-consistency assumptions in old code** — "list then process" silently missing files.
* **Hot prefix throttling** when a job writes every part file under one date prefix.
***
#### 5.4 Kubernetes / YARN as the job orchestrator [#54-kubernetes--yarn-as-the-job-orchestrator]
**Problem it solves.** Decide *where and when* each task runs, enforce isolation, and reclaim resources on failure — §2.3's three components.
**Why not just SSH and run it?** No global view of capacity, no fairness, no isolation, no automatic retry on node loss, no bin-packing.
**How it works internally.** Resource manager holds cluster state in a consensus store (**ZooKeeper for YARN, etcd for Kubernetes** — Ch 10 §4's "outsource the consensus"). Scheduler matches pending tasks to nodes using heuristics (§2.3: **FIFO, DRF, priority queues, capacity/quota, bin-packing** — because the optimal problem is **NP-hard**). Executors enforce limits with **cgroups**.
**Monitoring.** Pending/unschedulable task count (**and the *reason*** — "no node fits the request" is a capacity-planning signal, not a bug); **queue wait time** per priority class; node resource fragmentation; **preemption rate** (§2.3 — expected and healthy on spot capacity, alarming on on-demand); scheduler decision latency.
**What actually breaks.**
* **Gang-scheduling deadlock** (§2.3): two jobs each hold half the cores they need, neither can proceed, neither releases. Requires a gang/coscheduling plugin with all-or-nothing admission.
* **Starvation** of large jobs by a stream of small ones.
* **Preemption thrash** — low-priority tasks killed and restarted repeatedly, so they never finish and all the work is wasted.
* **Resource requests set from guesswork** — too high wastes the cluster, too low gets tasks OOMKilled or CPU-throttled.
* **The centralized resource manager as a bottleneck** — §2.3 warns about it explicitly, and it shows up as scheduler latency growing with cluster size.
***
# 11.10 Terminology introduced here (/docs/ddia/batch-processing/terminology-introduced-here)
# 11.8 Worked examples (/docs/ddia/batch-processing/worked-examples)
**① Working set: hash vs sort.** 10 billion log lines, 2 million distinct URLs, average URL 60 bytes.
* **Hash:** 2e6 × (60 B + 8 B counter + \~50 B overhead) ≈ **236 MB** → fits comfortably. **Record count is irrelevant.**
* Now 10 billion lines with **500 million distinct URLs** (e.g. keyed by session ID): 5e8 × 118 B ≈ **59 GB** → exceeds a laptop, and sorting (spilling to disk, sequential I/O) wins.
**The lesson: choose by cardinality, not volume.**
**② Shuffle volume.** 1 TB input, 200 map tasks, 200 reduce tasks. Each mapper writes **200 files**, so the shuffle produces **40,000 files** and moves \~1 TB across the network. At 10 Gb/s per node over 20 nodes, that's \~1 TB / 25 GB/s ≈ **40 s of pure network time** — before any computation. **This is why advantage ② (operator fusion, avoiding unnecessary shuffles) matters more than raw CPU.**
**③ Skew.** 100 reduce tasks; keys distributed so one key has 30% of the rows. That reducer does **30% of the work alone**, so wall-clock ≈ 0.30 × total, versus 0.01 × total if perfectly balanced — a **30× worse** stage duration than the ideal, no matter how many machines you add. **Adding nodes cannot fix skew; only changing the key can.**
**④ Recomputation cost on spot instances.** A 5-stage job, each stage 20 minutes, intermediate data in memory. An executor is preempted during stage 5 holding shuffle output from stage 4. Spark must recompute stage 4 for the lost partitions. With a 10% preemption rate per 20-minute window across 100 executors, **you will lose executors most runs** — and without an external shuffle service, expected runtime inflates substantially. **This is the concrete reason MapReduce wrote everything to the DFS, and the concrete cost of Spark's choice not to.**
**⑤ The direct-write anti-pattern, quantified.** 200 parallel tasks × 5,000 records/s each = **1,000,000 writes/s** aimed at a production database sized for 20,000 writes/s. **50× over capacity** — the database doesn't just slow the job, it **takes down the user-facing application.** Via Kafka, the same 1M/s is absorbed by the log's sequential writes and the consumer drains at 20,000/s over \~50× longer — **with production traffic unaffected.**
**⑥ Block size and metadata.** 1 PB of data.
* ext4-style 4 KiB blocks → **2.7 × 10¹¹ blocks** of metadata. Impossible.
* HDFS 128 MB blocks → **8.4 million blocks.** At \~150 bytes of NameNode memory per block, ≈ **1.2 GB** — manageable.
* Same 1 PB stored as **1 billion 1 MB files** → 1 billion blocks → **\~150 GB of NameNode heap.** **This is the small-file problem, and it's a metadata problem, not a data problem.**
***
# 10.3 Consensus (/docs/ddia/consistency-consensus/consensus)
**The three things that are easy on one node and hard with fault tolerance:**
1. **Linearizable database** — easy with one leader; **but how do you FAIL OVER while avoiding split brain? How do you ensure a node that believes it's the leader hasn't been voted out while temporarily paused?**
2. **Linearizable ID generator** — just a counter with atomic fetch-and-add; **what if it crashes?**
3. **Atomic CAS** — may be one CPU instruction on a single node; **how do you make it fault-tolerant?**
> **It turns out ALL OF THESE ARE INSTANCES OF THE SAME FUNDAMENTAL PROBLEM: CONSENSUS.**
**The four best-known algorithms:** **Viewstamped Replication, Paxos, Raft, Zab.** *"These algorithms have quite a few similarities, but THEY ARE NOT THE SAME."* All work in a **non-Byzantine system model** — network communication may be arbitrarily delayed or dropped, nodes may crash/restart/disconnect, **but nodes otherwise follow the protocol and do not behave maliciously.**
#### 3.1 The FLP result — what it actually says [#31-the-flp-result--what-it-actually-says]
> **The FLP result proves NO ALGORITHM IS ALWAYS ABLE TO REACH CONSENSUS if there is a risk that a node may crash. Yet here we are, discussing consensus algorithms. What's going on?**
>
> **First, FLP DOESN'T SAY WE CAN NEVER REACH CONSENSUS; it only says WE CAN'T GUARANTEE A CONSENSUS ALGORITHM WILL ALWAYS TERMINATE.**
>
> **Second, FLP is proved assuming a DETERMINISTIC algorithm in the ASYNCHRONOUS system model — meaning THE ALGORITHM CANNOT USE ANY CLOCKS OR TIMEOUTS. If it CAN use timeouts to suspect a node has crashed (even if the suspicion is sometimes wrong), CONSENSUS BECOMES SOLVABLE. Even allowing RANDOM NUMBERS is sufficient.**
#### 3.2 The many faces of consensus — all equivalent [#32-the-many-faces-of-consensus--all-equivalent]
##### (a) Single-value consensus — the four properties [#a-single-value-consensus--the-four-properties]
| Property | Definition | Kind |
| --------------------- | ----------------------------------------------------------------- | ------------ |
| **Uniform agreement** | **No two nodes decide differently** | **SAFETY** |
| **Integrity** | **After a node has decided one value, it cannot change its mind** | **SAFETY** |
| **Validity** | **If a node decides value v, then v was PROPOSED BY A NODE** | **SAFETY** |
| **Termination** | **Every node that does not crash EVENTUALLY DECIDES a value** | **LIVENESS** |
> **Agreement and integrity define the core idea. VALIDITY RULES OUT TRIVIAL SOLUTIONS — an algorithm that always decides `null` would satisfy agreement and integrity, but not validity.**
>
> **If you don't care about fault tolerance, the first three are EASY: hardcode one node as the "DICTATOR." But if that node fails, the system can no longer decide anything. ALL THE DIFFICULTY ARISES FROM THE NEED FOR FAULT TOLERANCE.**
>
> **TERMINATION formalizes fault tolerance: the algorithm must MAKE PROGRESS even if some nodes fail. And it must decide EVEN IF A CRASHED NODE NEVER COMES BACK.** *(Instead of a software crash, imagine an earthquake causes the datacenter to be destroyed by a landslide — **you must assume your node is buried under 30 feet of mud and is never coming back online.**)*
> ### **Any consensus algorithm requires AT LEAST A MAJORITY OF NODES to be functioning correctly in order to assure termination.** [#any-consensus-algorithm-requires-at-least-a-majority-of-nodes-to-be-functioning-correctly-in-order-to-assure-termination]
>
> **However, most consensus algorithms ensure THE SAFETY PROPERTIES ARE ALWAYS MET, EVEN IF A MAJORITY OF NODES FAIL or a severe network problem occurs. Thus a large-scale outage can STOP THE SYSTEM FROM PROCESSING REQUESTS, BUT IT CANNOT CORRUPT THE CONSENSUS SYSTEM BY CAUSING INCONSISTENT DECISIONS.**
##### (b) CAS ⟺ consensus [#b-cas--consensus]
*(A real example: **conditional writes in object stores** — Ch 6 §1.3.)*
##### (c) Shared logs ⟺ consensus [#c-shared-logs--consensus]
**The five properties of a shared log:**
| Property | Definition |
| --------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Eventual append** | **If a node requests a value be added and doesn't crash, it must EVENTUALLY READ that value in a log entry** |
| **Reliable delivery** | **No entries are lost — if one node reads an entry, eventually EVERY non-crashed node reads it** |
| **Append-only** | **After a node reads an entry it is IMMUTABLE, and new entries can be added only AFTER it, not before.** Rereading gives **the same entries in the same order, even after a crash and restart** |
| **Agreement** | **If two nodes both read entry e, then PRIOR TO e THEY MUST HAVE READ EXACTLY THE SAME SEQUENCE of entries in the same order** |
| **Validity** | **If a node reads an entry containing a value, a node PREVIOUSLY REQUESTED that value's addition** |
*(A shared log is implemented with a **TOTAL ORDER BROADCAST** protocol, also known as **atomic broadcast** or **total order multicast**: to add a value you "broadcast" it; when the protocol "delivers" it, it becomes a log entry.)*
> **Single-leader replication WITHOUT failover does not meet the liveness requirement, since it stops delivering messages if the leader crashes. AS USUAL, THE CHALLENGE IS PERFORMING FAILOVER SAFELY AND AUTOMATICALLY.**
##### (d) Fetch-and-add — the one that ALMOST works [#d-fetch-and-add--the-one-that-almost-works]
##### (e) Atomic commitment ⟺ consensus [#e-atomic-commitment--consensus]
> **The important difference: WITH CONSENSUS IT'S OK TO DECIDE ANY VALUE THAT WAS PROPOSED, WHEREAS WITH ATOMIC COMMITMENT THE ALGORITHM MUST ABORT IF ANY PARTICIPANT VOTED TO ABORT.**
**Atomic commitment's five properties:** uniform agreement · integrity · **validity — *"If a node commits, ALL nodes must have previously voted to commit. If ANY node voted to abort, ALL nodes must abort"*** · **nontriviality — *"If all nodes vote to commit, AND NO COMMUNICATION TIMEOUTS OCCUR, then all nodes must commit"*** (this rules out an algorithm that always aborts) · termination.
#### 3.3 Consensus in practice — shared logs win [#33-consensus-in-practice--shared-logs-win]
> **Which formulation is most useful in practice? MOST CONSENSUS SYSTEMS PROVIDE SHARED LOGS. Raft, Viewstamped Replication, and Zab provide them out of the box. Paxos provides single-value consensus, but in practice most systems using Paxos use MULTI-PAXOS, which also provides a shared log.**
**Why a shared log is such a good primitive:**
| Use | How |
| ------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Database replication** | **Every entry is a write; every replica processes the same writes in the same order with DETERMINISTIC logic ⇒ all replicas end up consistent. This is STATE MACHINE REPLICATION**, the principle behind event sourcing (Ch 3) |
| **Serializable transactions** | **Every entry is a DETERMINISTIC transaction executed as a stored procedure; every node executes them in the same order ⇒ SERIALIZABLE** (Ch 8 §4.1) |
| **Single-value consensus / CAS** | **Decide the value that appears FIRST in the log** |
| **Many instances of single-value consensus** (one per theater seat) | **Include the seat number in the entry; decide the FIRST entry containing that seat number** |
| **Atomic fetch-and-add** | **Put the addend in an entry; the current value is the SUM of all entries so far** |
| **Fencing tokens** | **A simple counter on log entries.** In ZooKeeper this is the **`zxid`** |
| **Stream processing** | Ch 12 |
> ⚠️ **Sharded databases with a strong consistency model often maintain A SEPARATE LOG PER SHARD, which improves scalability BUT LIMITS THE CONSISTENCY GUARANTEES (consistent snapshots, foreign-key references) they can offer ACROSS shards.**
#### 3.4 From single-leader replication to consensus — breaking the circularity [#34-from-single-leader-replication-to-consensus--breaking-the-circularity]
> **It seems we need consensus to elect a leader, and we need a leader to solve consensus. HOW DO WE BREAK OUT OF THIS CONUNDRUM?**
>
> ### **Consensus algorithms DON'T REQUIRE THAT THERE IS ONLY ONE LEADER AT ANY ONE TIME. Instead they make a WEAKER guarantee: they define an EPOCH NUMBER and guarantee that WITHIN EACH EPOCH, THE LEADER IS UNIQUE.** [#consensus-algorithms-dont-require-that-there-is-only-one-leader-at-any-one-time-instead-they-make-a-weaker-guarantee-they-define-an-epoch-number-and-guarantee-that-within-each-epoch-the-leader-is-unique]
**The epoch number under four names:**
| Algorithm | Name |
| --------------------------- | ----------------- |
| **Paxos** | **ballot number** |
| **Viewstamped Replication** | **view number** |
| **Raft** | **term number** |
| (generic) | **epoch number** |
**The two-round structure:**
> **These two rounds look SUPERFICIALLY SIMILAR TO 2PC, BUT THEY ARE VERY DIFFERENT PROTOCOLS:**
>
> | | **Consensus** | **2PC** |
> | --------------- | ------------------------------------- | ------------------------------------------ |
> | Who starts it | **ANY node can start an election** | **ONLY the coordinator can request votes** |
> | What's required | **Only a QUORUM of nodes to respond** | **A YES VOTE FROM EVERY PARTICIPANT** |
#### 3.5 Subtleties [#35-subtleties]
**Every new log entry is synchronously replicated to a quorum before it is confirmed to the client — this ensures the entry won't be lost if the current leader fails.**
**How the algorithms differ on honoring the old leader's entries:**
| Algorithm | Approach |
| --------- | ------------------------------------------------------------------------------------------------------------------------------- |
| **Raft** | **Allows a node to become leader ONLY IF ITS LOG IS AT LEAST AS UP TO DATE as those of a majority of its followers** |
| **Paxos** | **Allows ANY node to become leader, but REQUIRES IT TO BRING ITS LOG UP TO DATE with other nodes before appending new entries** |
**⚠️ The consistency-vs-availability choice inside leader election:**
> **It's ESSENTIAL that the new leader is up to date with any confirmed entries before processing writes or linearizable reads. IF A NODE WITH STALE DATA BECAME LEADER, IT MIGHT WRITE NEW VALUES TO LOG ENTRIES THAT WERE ALREADY WRITTEN by the old leader, VIOLATING THE APPEND-ONLY PROPERTY.**
>
> **In some cases you might choose to WEAKEN the consensus properties to recover more quickly, or to be able to recover at all. Kafka offers UNCLEAN LEADER ELECTION, allowing ANY replica to become leader even if not up to date. Also, in databases with asynchronous replication, you cannot guarantee ANY follower is up to date when the leader fails.**
>
> ### **If you drop the requirement, you may improve performance and availability, BUT YOU ARE ON THIN ICE, SINCE THE THEORY OF CONSENSUS NO LONGER APPLIES. While things will work fine as long as there are no faults, the problems in Ch 9 CAN EASILY CAUSE DATA LOSS OR CORRUPTION.** [#if-you-drop-the-requirement-you-may-improve-performance-and-availability-but-you-are-on-thin-ice-since-the-theory-of-consensus-no-longer-applies-while-things-will-work-fine-as-long-as-there-are-no-faults-the-problems-in-ch-9-can-easily-cause-data-loss-or-corruption]
**Linearizable reads need a quorum too:**
> **Turning writes into log entries and replicating them to a quorum ISN'T ALL THAT'S REQUIRED. If you want LINEARIZABLE READS, THEY ALSO HAVE TO GO THROUGH A QUORUM VOTE, similarly to a write, TO CONFIRM THAT THE NODE THAT BELIEVES ITSELF TO BE LEADER REALLY IS STILL UP TO DATE. Linearizable reads in etcd work like this.**
**Reconfiguration:** most algorithms in standard form **assume a FIXED SET OF NODES.** Extensions make **adding/removing nodes possible — especially useful when adding new regions, or MIGRATING from one location to another (first adding new nodes, then removing old ones).**
#### 3.6 Pros and cons [#36-pros-and-cons]
> ### **Consensus is essentially "SINGLE-LEADER REPLICATION DONE RIGHT," with automatic failover on leader failure, ensuring that NO COMMITTED DATA IS LOST and SPLIT BRAIN IS NOT POSSIBLE, even in the face of all the problems in Ch 9.** [#consensus-is-essentially-single-leader-replication-done-right-with-automatic-failover-on-leader-failure-ensuring-that-no-committed-data-is-lost-and-split-brain-is-not-possible-even-in-the-face-of-all-the-problems-in-ch-9]
>
> **ANY SYSTEM THAT PROVIDES AUTOMATIC FAILOVER BUT DOES NOT USE A PROVEN CONSENSUS ALGORITHM IS LIKELY TO BE UNSAFE.** *(Using one is not a guarantee of whole-system correctness — there are still plenty of places where bugs can lurk — but it's a good start.)*
**The five costs:**
1. **Always requires a STRICT MAJORITY** — three nodes to tolerate one failure, five to tolerate two
2. **Every operation requires communication with a quorum, so YOU CAN'T INCREASE THROUGHPUT BY ADDING MORE NODES — IN FACT, EVERY NODE YOU ADD MAKES THE ALGORITHM SLOWER**
3. **If a network partition cuts off some nodes, ONLY THE MAJORITY PORTION CAN MAKE PROGRESS; the other nodes are BLOCKED**
4. **Timeout tuning is hard in environments with highly variable network delays, especially across regions.** *Too large → slow recovery. Too small → **lots of unnecessary leader elections, resulting in TERRIBLE PERFORMANCE as the system SPENDS MORE TIME CHOOSING LEADERS THAN DOING USEFUL WORK.***
5. **Sensitivity to specific network problems.** **Raft has unpleasant edge cases: if the entire network works correctly EXCEPT ONE CONSISTENTLY UNRELIABLE LINK, Raft can get into situations where LEADERSHIP CONTINUALLY BOUNCES BETWEEN TWO NODES, or the current leader is CONTINUALLY FORCED TO RESIGN, so the system EFFECTIVELY NEVER MAKES PROGRESS.** *(Addressed by a **PRE-VOTE PHASE**.)* **Paxos also depends on leaders and can have similar issues; EGALITARIAN PAXOS (EPaxos) uses a LEADERLESS protocol more robust against poorly performing nodes or connections.**
***
# 10.4 Coordination Services (/docs/ddia/consistency-consensus/coordination-services)
**ZooKeeper, etcd, Consul — modeled after Google's Chubby lock service.**
> **Although they look superficially like any other key-value store, THEY ARE NOT DESIGNED FOR HIGH WRITE VOLUMES OR GENERAL-PURPOSE DATA STORAGE. Instead, they are designed TO COORDINATE AMONG NODES OF ANOTHER DISTRIBUTED SYSTEM.** *(Kubernetes relies on etcd; Spark and Flink in HA mode rely on ZooKeeper.)*
>
> **They hold SMALL amounts of data that FIT ENTIRELY IN MEMORY (although they still write to disk for durability), replicated across multiple nodes via a fault-tolerant consensus algorithm.**
**Four features — note which need consensus and which don't:**
| Feature | Needs consensus? | Detail |
| ------------------------ | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Locks and leases** | **YES** | Built on the atomic, fault-tolerant CAS. **If several nodes concurrently try to acquire the same lease, only one succeeds** |
| **Support for fencing** | **YES** | **Monotonically increasing ID per log entry** — `zxid`/`cversion` in ZooKeeper, revision number in etcd |
| **Failure detection** | no | **Clients maintain a LONG-LIVED SESSION and exchange HEARTBEATS. Even if the connection is temporarily interrupted or a server fails, ANY LEASES HELD REMAIN ACTIVE. But if there's no heartbeat for longer than the lease timeout, the service ASSUMES THE CLIENT IS DEAD AND RELEASES THE LEASE** (ZooKeeper's **ephemeral nodes**) |
| **Change notifications** | no | **A client can request notification whenever certain keys change** — finding out when another client joins (by the value it writes) or fails (its ephemeral nodes disappear). **Saves the client from frequently polling** |
**Four use cases:**
**① Configuration management.** Store timeouts, thread pool sizes as key-value pairs; processes load on startup and subscribe to changes.
> **This DOESN'T NEED the consensus aspect, but it's CONVENIENT if you're already running the service. Alternatively, a process could periodically POLL a file or URL, avoiding a specialized service.**
**② Allocating work to nodes.** Choosing a leader/primary among several instances (necessary for single-leader databases, **also appropriate for job schedulers and similar stateful systems**), and **deciding which SHARD to assign to which node** — rebalancing as nodes join, taking over as nodes fail.
> **These can be achieved by judicious use of atomic operations, ephemeral nodes, and notifications. IT'S NOT EASY, despite libraries like Apache Curator — BUT IT IS STILL MUCH BETTER THAN ATTEMPTING TO IMPLEMENT THE CONSENSUS ALGORITHMS FROM SCRATCH, WHICH WOULD BE VERY PRONE TO BUGS.**
**The key architectural advantage:**
> ### **A dedicated coordination service can run on a FIXED SET OF NODES (usually three or five), REGARDLESS OF HOW MANY NODES ARE IN THE SYSTEM THAT RELIES ON IT. In a storage system with THOUSANDS of shards, running a consensus algorithm over thousands of nodes would be TERRIBLY INEFFICIENT; it's much better to "OUTSOURCE" THE CONSENSUS to a small number of nodes.** [#a-dedicated-coordination-service-can-run-on-a-fixed-set-of-nodes-usually-three-or-five-regardless-of-how-many-nodes-are-in-the-system-that-relies-on-it-in-a-storage-system-with-thousands-of-shards-running-a-consensus-algorithm-over-thousands-of-nodes-would-be-terribly-inefficient-its-much-better-to-outsource-the-consensus-to-a-small-number-of-nodes]
**The data-rate constraint:**
> **The data is quite SLOW-CHANGING — "the node running on IP 10.1.1.23 is the leader for shard 7" — changing on a timescale of MINUTES OR HOURS. Coordination services are NOT INTENDED FOR DATA THAT MAY CHANGE THOUSANDS OF TIMES PER SECOND.** For that, use a conventional database, or **Apache BookKeeper** to replicate fast-changing internal state.
**③ Service discovery.**
> **Convenient — failure detection and change notification make it easy to track instances as they come and go. And if you're already using it for leases and leader election, it makes sense to use it for discovery too.**
>
> **HOWEVER, USING CONSENSUS FOR SERVICE DISCOVERY IS OFTEN OVERKILL. This use case GENERALLY DOESN'T REQUIRE LINEARIZABILITY, and it's MORE IMPORTANT THAT IT IS HIGHLY AVAILABLE AND FAST, since without it everything would grind to a halt. It's therefore usually preferable to CACHE service discovery information** — clients that can't connect bypass the cache, retry with the latest value, and update it; caches may also refresh on a **TTL**. *(DNS-based discovery uses multiple layers of caching for exactly this reason.)*
**ZooKeeper OBSERVERS** support this: **replicas that receive the log and maintain a copy of the data but DO NOT PARTICIPATE IN THE VOTING PROCESS. Reads from an observer are NOT LINEARIZABLE as they might be stale, BUT THEY REMAIN AVAILABLE EVEN IF THE NETWORK IS INTERRUPTED, and they increase read throughput.**
***
# 10.7 Decision cheat sheet (/docs/ddia/consistency-consensus/decision-cheat-sheet)
**Do I need linearizability?**
**Linearizability or serializability — which do I actually want?**
**Serializability** if the problem is *multi-object transactional correctness* (write skew, invariants across rows). **Linearizability** if the problem is *recency of a single object* (is this the latest leader? is this username taken?). **Both (strict serializability)** if you need transactions *and* recency — and know it costs coordination.
**Which ID scheme?**
| Need | Use |
| ------------------------------------------------------ | ------------------------------------------------------------------------------------ |
| Compact, ordered, single-node | **Autoincrement** |
| Distributed, no coordination, don't care about order | **UUIDv4** |
| Distributed, roughly time-ordered, good index locality | **UUIDv7 / Snowflake / ULID** |
| Causally-consistent ordering across nodes | **Lamport clock or HLC** |
| Detect concurrency | **Vector clock** (pay the space) |
| **Linearizable** ordering | **Timestamp oracle** (single node, batched) **or Spanner-style clock + commit wait** |
**How many consensus nodes?**
**3** tolerates 1 failure; **5** tolerates 2. Beyond that you're paying latency for tolerance you don't need. **Never even numbers** — 4 nodes tolerate the same 1 failure as 3, with more coordination.
**Should I use consensus for service discovery?**
**Usually no** — it doesn't need linearizability, and **availability and speed matter more.** Cache aggressively, use TTLs, use observers/read replicas. Use consensus for the *leases and leadership*; use caching for the *lookups*.
**Rule of thumb for automatic failover:**
> **Any system that provides automatic failover but does not use a proven consensus algorithm is likely to be unsafe.** If you're writing your own failover logic, you are writing a consensus algorithm — badly.
***
# 10.11 Forward links (/docs/ddia/consistency-consensus/forward-links)
| Concept here | Where it's developed |
| -------------------------------------------------------- | ------------------------------------ |
| Shared logs as the backbone of stream processing | **Ch 12** — Stream Processing |
| State machine replication and event sourcing | **Ch 3**, **Ch 12** |
| Avoiding linearizability without sacrificing correctness | **Ch 13** — Aiming for Correctness |
| Loosely enforced constraints and compensation | **Ch 13** — Timeliness and Integrity |
| Why timeouts and pauses make this all hard | **Ch 9** |
| Fencing tokens generated by consensus | **Ch 9** §5.2 |
| Two-phase commit and its coordinator problem | **Ch 8** §5 |
| Shard assignment via coordination services | **Ch 7** §4 |
| Replication models this chapter classifies | **Ch 6** |
# 10.2 ID Generators and Logical Clocks (/docs/ddia/consistency-consensus/id-generators-logical-clocks)
**Why a single-node autoincrementing counter is nice:** compact (**64 bits, or 32 if you're sure you'll never exceed 4 billion records — but that is risky**), **and the ORDER of the IDs tells you the order in which records were created.**
> **This single-node ID generator is another example of A LINEARIZABLE SYSTEM. Each request is an atomic FETCH-AND-ADD; linearizability ensures that if Aaliyah's post completes before Bryce's begins, Bryce's ID must be greater.** *(Concurrent posts may be ordered either way, as long as the IDs are unique.)*
**Three problems:**
1. **Not fault-tolerant — a single point of failure**
2. **Slow for records created in another region — potentially a round trip TO THE OTHER SIDE OF THE PLANET just to get an ID**
3. **Could become a bottleneck at high write throughput**
#### 2.1 The alternatives, and what each loses [#21-the-alternatives-and-what-each-loses]
| Scheme | How | What it loses |
| ------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Sharded ID assignment** | One node generates only even numbers, another only odd; generally **reserve some bits for a shard number** | **Ordering.** IDs 16 and 17 don't tell you which message was sent first — **one node might have been ahead of the other** |
| **Preallocated blocks** | Node A claims 1–1,000, node B claims 1,001–2,000; each hands out from its block and requests a new block when running low | **Ordering.** **One message may get an ID in 1,001–2,000 and a LATER message an ID in 1–1,000** |
| **Random UUIDs (v4)** | **Generated locally on any node WITHOUT COMMUNICATION** | **128 bits**, and **the order is RANDOM — comparing two IDs tells you NOTHING about which is newer** |
| **Wall-clock timestamp made unique** | Timestamp in the **most significant bits**, remaining bits filled with a **shard number + per-shard sequence, or a long random value.** Used by **UUIDv7, X's Snowflake, ULIDs, Hazelcast Flake IDs, MongoDB ObjectIDs** | **At best APPROXIMATE ordering.** An earlier write from a slightly fast clock and a later write from a slightly slow clock get inverted; **with clock jumps, EVEN A SINGLE NODE'S TIMESTAMPS MIGHT BE ORDERED INCORRECTLY.** **Unlikely to be linearizable** |
#### 2.2 Logical clocks [#22-logical-clocks]
> **A LOGICAL CLOCK is an ALGORITHM THAT COUNTS THE EVENTS that have occurred. A timestamp from a logical clock DOESN'T TELL YOU WHAT TIME IT IS, but you can compare two timestamps to tell which is earlier and which is later.**
**Three requirements:**
1. **Timestamps are compact (a few bytes) and unique**
2. **Any two can be compared to determine which is earlier — they are TOTALLY ORDERED**
3. **The order is CONSISTENT WITH CAUSALITY: if A happened before B, then A's timestamp \< B's timestamp**
> **A single-node ID generator meets these. The distributed ID generators above DO NOT meet the causal ordering requirement.**
##### Lamport clocks (1978, Leslie Lamport — one of the most-cited papers in distributed systems) [#lamport-clocks-1978-leslie-lamport--one-of-the-most-cited-papers-in-distributed-systems]
> ⚠️ **Although Lamport clocks provide a TOTAL ORDERING, THEY DO NOT PROVIDE LINEARIZABILITY — they are not a way of ensuring a value is up to date. They are MERELY a way of assigning IDs such that if A happened before B, A's ID is less than B's.**
**Two limitations:**
1. **No direct relation to physical time** — you can't find all messages posted on a particular date; **you'd need to store the physical time separately**
2. **If two nodes NEVER COMMUNICATE, one node's increments are never reflected in the other's counter.** So **events generated around the same time on different nodes could have WILDLY DIFFERENT counter values**
##### Hybrid logical clocks (HLC) [#hybrid-logical-clocks-hlc]
> **Combines the advantages of physical time-of-day clocks with the ordering guarantees of Lamport clocks.**
>
> * **Like a PHYSICAL clock: it counts seconds or microseconds**
> * **Like a LAMPORT clock: when one node sees a greater timestamp from another, it MOVES ITS OWN LOCAL VALUE FORWARD to match. So if one node's clock runs fast, THE OTHERS WILL SIMILARLY MOVE THEIR CLOCKS FORWARD when they communicate**
> * **Every generated timestamp is ALSO INCREMENTED, ensuring the clock moves forward MONOTONICALLY even if the underlying physical clock JUMPS BACKWARD (e.g. from NTP adjustments)**
>
> **⇒ You can treat an HLC timestamp ALMOST LIKE a conventional time-of-day timestamp, with the added property that ITS ORDERING IS CONSISTENT WITH HAPPENS-BEFORE. It doesn't depend on special hardware and requires only ROUGHLY SYNCHRONIZED CLOCKS.** *(Used by CockroachDB.)*
**Where these fit with MVCC:** **Lamport clocks and HLCs are a GOOD WAY of generating the transaction IDs that snapshot isolation needs (Ch 8), because they ensure THE SNAPSHOT IS CONSISTENT WITH CAUSALITY.**
##### vs vector clocks [#vs-vector-clocks]
| | **Lamport / HLC** | **Vector clock** |
| ------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------- |
| **Concurrent timestamps** | **Ordered ARBITRARILY. You generally CAN'T TELL whether two timestamps were generated concurrently or one happened before the other** | **CAN detect concurrency: if A has a higher counter for one node and B a higher counter for another, A and B MUST BE CONCURRENT** |
| **Size** | compact | **Much larger — potentially ONE INTEGER FOR EVERY NODE in the system** |
#### 2.3 Why logical clocks aren't enough — the privacy leak [#23-why-logical-clocks-arent-enough--the-privacy-leak]
> **Linearizability requires that if request A COMPLETED before request B BEGAN, then B must have the higher ID — EVEN IF A AND B NEVER COMMUNICATED WITH EACH OTHER. Lamport clocks can ensure only that a node generates timestamps GREATER THAN ANY OTHER TIMESTAMP THAT NODE HAS SEEN; no such guarantees can be made about timestamps IT HASN'T SEEN.**
**Possible fixes, and why they're bad:** *the photos DB could read the account status before writing — "but it's EASY TO FORGET SUCH A CHECK." The app could track the user's latest write timestamp — "but if the user uses a LAPTOP AND A PHONE, THAT'S NOT SO EASY."* **The simplest solution is a LINEARIZABLE ID GENERATOR.**
#### 2.4 Implementing a linearizable ID generator [#24-implementing-a-linearizable-id-generator]
**Option A — a single node doing three things:**
1. **Atomically increment a counter and return its value**
2. **Persist the counter** (so it doesn't generate duplicates after a crash)
3. **Replicate it for fault tolerance** (single-leader replication)
*(Used in practice: **TiDB/TiKV calls it a TIMESTAMP ORACLE, inspired by Google's Percolator.**)*
**The batching optimization:**
> **Avoid a disk write and replication on every request: write a record describing A BATCH of IDs; once persisted and replicated, hand out those IDs in sequence. Before running out, persist the record for the next batch. SOME IDs WILL BE SKIPPED if the node crashes or you fail over — BUT YOU WON'T ISSUE ANY DUPLICATE OR OUT-OF-ORDER IDs.**
**The limits:**
* **You can't easily SHARD it** — multiple shards independently handing out IDs breaks linearizable order
* **You can't easily distribute it across REGIONS** — in a geo-distributed database, **all ID requests must go to a node in a single region**
* **On the upside, the job is very simple, so a single node can handle a large request throughput**
**Option B — Spanner's approach:** rely on a physical clock returning **a RANGE of timestamps** and **wait for the uncertainty interval to elapse** (Ch 9 §3.6).
> **This guarantees linearizable ID assignment WITHOUT ANY COMMUNICATION; even requests in different regions are ordered correctly, WITHOUT WAITING FOR CROSS-REGION REQUESTS. The downside is that you need HARDWARE AND SOFTWARE SUPPORT for tightly synchronized clocks and computing the uncertainty interval.**
#### 2.5 Why even a linearizable ID generator isn't enough for locks [#25-why-even-a-linearizable-id-generator-isnt-enough-for-locks]
> **You could use a logical clock to assign timestamps to lock requests and pick the LOWEST as the winner. If the clock is linearizable, you know future requests will generate greater timestamps.**
>
> ### **But part of the problem is STILL UNSOLVED: HOW DOES A NODE KNOW WHETHER ITS OWN TIMESTAMP IS THE LOWEST? To be sure, IT NEEDS TO HEAR FROM EVERY OTHER NODE that might have generated a timestamp. If one of them has failed or is unreachable, THE SYSTEM WOULD GRIND TO A HALT. This is not the kind of fault-tolerant system we need.** [#but-part-of-the-problem-is-still-unsolved-how-does-a-node-know-whether-its-own-timestamp-is-the-lowest-to-be-sure-it-needs-to-hear-from-every-other-node-that-might-have-generated-a-timestamp-if-one-of-them-has-failed-or-is-unreachable-the-system-would-grind-to-a-halt-this-is-not-the-kind-of-fault-tolerant-system-we-need]
>
> **To implement locks, leases, and similar constructs in a fault-tolerant way, WE NEED SOMETHING STRONGER. WE NEED CONSENSUS.**
***
# 10. Consistency and Consensus (/docs/ddia/consistency-consensus)
> "An ancient adage warns, 'Never go to sea with two chronometers; take one or three.'" — Frederick P. Brooks Jr.
**The two competing philosophies for dealing with replica inconsistency:**
| | **Eventual consistency** | **Strong consistency** |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| **Stance** | **The fact that a system is replicated is MADE VISIBLE to the application; you as the developer are expected to deal with the inconsistencies and conflicts** | **Applications should NOT have to worry about internal details of replication; the system should BEHAVE AS IF IT WERE A SINGLE NODE** |
| **Used by** | Multi-leader, leaderless replication | Single-leader, consensus-based systems |
| **Advantage** | Tolerates faults that break strongly consistent systems; **inevitable if users make changes offline** | **Simpler for you, the application developer** |
| **Disadvantage** | **Can be difficult for applications to deal with** | **A performance cost, and SOME KINDS OF FAULT THAT AN EVENTUALLY CONSISTENT SYSTEM CAN TOLERATE CAUSE OUTAGES** |
> **If your replicas are located in datacenters with fast, reliable communication, STRONG CONSISTENCY IS OFTEN APPROPRIATE because its cost is acceptable.**
**The chapter's three movements:**
> **These topics are NOTORIOUS for being hard to implement correctly. It's very easy to build systems that behave fine when there are no faults but COMPLETELY FALL APART when faced with an unlucky combination of faults or message orderings that their designers hadn't considered.**
***
# 10.1 Linearizability (/docs/ddia/consistency-consensus/linearizability)
**Other names for the same thing:** **atomic consistency, strong consistency, immediate consistency, external consistency.**
> **The basic idea: make a system APPEAR AS IF THERE IS ONLY ONE COPY OF THE DATA, and all operations on it are ATOMIC.**
>
> ### **In a linearizable system, AS SOON AS ONE CLIENT SUCCESSFULLY COMPLETES A WRITE, ALL CLIENTS READING FROM THE DATABASE MUST BE ABLE TO SEE THE VALUE JUST WRITTEN.** [#in-a-linearizable-system-as-soon-as-one-client-successfully-completes-a-write-all-clients-reading-from-the-database-must-be-able-to-see-the-value-just-written]
>
> **Linearizability is a RECENCY GUARANTEE:** the value read is the most recent, up-to-date value, not from a stale cache or replica.
#### 1.1 The sports website — why it's a violation [#11-the-sports-website--why-its-a-violation]
#### 1.2 Building up the definition [#12-building-up-the-definition]
**Vocabulary:** a **register** = one key in a KV store, one row, one document. Operations:
* `Read(x) ⇒ v`
* `Write(x, v) ⇒ r` (r = OK or Error)
* `CAS(x, v_old, v_new) ⇒ r` — atomically set x to v\_new **only if** it currently equals v\_old
**Each bar is a request. The start is when the client SENT it; the end is when the client RECEIVED the response. The client doesn't know exactly when the database processed it — only that it happened somewhere in between.**
**Step 1 — what concurrent operations may return:**
**Step 2 — the extra constraint that makes it linearizable:**
> **If reads concurrent with a write could freely return either value, readers could see the value FLIP BACK AND FORTH several times while a write is going on. THAT IS NOT WHAT WE EXPECT OF A SYSTEM THAT EMULATES "A SINGLE COPY OF THE DATA."**
>
> **We imagine there must be SOME POINT IN TIME (between the start and end of the write) at which the value ATOMICALLY FLIPS from 0 to 1. Thus, IF ONE CLIENT'S READ RETURNS THE NEW VALUE, ALL SUBSEQUENT READS MUST ALSO RETURN THE NEW VALUE, even if the write operation has not yet completed.**
**Step 3 — the full picture with CAS:**
> **Each operation is marked with a vertical line at the time we think it took effect. Those markers are joined in a sequential order, and the result MUST BE A VALID SEQUENCE OF READS AND WRITES FOR A REGISTER (every read returns the value set by the most recent write).**
>
> ### **The requirement of linearizability is that THE LINES JOINING UP THE OPERATION MARKERS ALWAYS MOVE FORWARD IN TIME, NEVER BACKWARD.** [#the-requirement-of-linearizability-is-that-the-lines-joining-up-the-operation-markers-always-move-forward-in-time-never-backward]
**Four details from the worked example worth internalizing:**
1. **B sent a read first, then D sent Write(x,0), then A sent Write(x,1) — and B's read returned 1.** ✔ OK: the database processed D's write, then A's write, then B's read. **Not the order they were sent, but an acceptable order, because the three requests are CONCURRENT.** *(Perhaps B's read was delayed in the network.)*
2. **B's read returned 1 BEFORE A received its "write succeeded" response.** ✔ OK — **the OK response to A was slightly delayed in the network.**
3. **This model assumes NO transaction isolation; another client may change a value at any time.** C reads 1 then reads 2 because B changed it in between. **An atomic CAS can be used to check the value hasn't been concurrently changed.**
4. **The final read by B is NOT linearizable.** It's concurrent with C's CAS (2→4). **In the absence of other requests, returning 2 would be fine. However, A had ALREADY READ 4 before B's read started, so B IS NOT ALLOWED TO READ AN OLDER VALUE THAN A.**
> **It is possible (though COMPUTATIONALLY EXPENSIVE) to TEST whether a system's behavior is linearizable, by recording the timings of all requests and responses and checking whether they can be arranged into a valid sequential order.** *(This is what Jepsen's Knossos checker does — Ch 9.)*
**Linearizability is the STRONGEST consistency model in common use.** It includes read-after-write consistency, monotonic reads, and consistent prefix reads (Ch 6) **and more.**
#### 1.3 Linearizability vs Serializability — the confusion that must be cleared [#13-linearizability-vs-serializability--the-confusion-that-must-be-cleared]
| System | What it provides |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Single-node databases** | **typically linearizable** |
| **CockroachDB** | **serializability and SOME recency guarantees, but NOT strict serializability** — because that would require expensive coordination between transactions |
| **Spanner, FoundationDB** | **strict serializability** |
> **The consistency model and isolation level can be chosen LARGELY INDEPENDENTLY from each other.**
#### 1.4 Where linearizability is genuinely required [#14-where-linearizability-is-genuinely-required]
**① Locking and leader election.**
> **A system using single-leader replication must ensure there is indeed only ONE leader, not several (split brain). No matter how the lease mechanism is implemented, IT MUST BE LINEARIZABLE. It shouldn't be possible for two nodes to acquire the lease at the same time.**
*(**ZooKeeper** and **etcd** implement this. Note: **strictly speaking, ZooKeeper provides linearizable WRITES, but READS may be stale, since there is no guarantee they are served from the current leader. etcd since v3 provides linearizable reads by default.** Libraries like **Apache Curator** provide higher-level recipes — and **many subtle details are involved, e.g. the fencing issue of Ch 9.**)*
*Granular case: **Oracle RAC uses a lock per disk page, with multiple nodes sharing disk storage. Since these linearizable locks are ON THE CRITICAL PATH of transaction execution, RAC deployments usually have A DEDICATED CLUSTER INTERCONNECT NETWORK.***
**② Constraints and uniqueness guarantees.**
> **A username or email must uniquely identify one user; a file storage service can't have two files with the same path. To enforce this AS THE DATA IS WRITTEN, YOU NEED LINEARIZABILITY.**
>
> **It's similar to a lock: registering a username is like ACQUIRING A LOCK ON IT — very similar to an atomic CAS, setting the username to the user's ID provided it is not already taken.**
**Same for:** a bank balance never going negative, not selling more items than are in stock, two people not booking the same seat. **All require A SINGLE UP-TO-DATE VALUE THAT ALL NODES AGREE ON.**
**But the practical escape hatch:** *"In real applications it is sometimes acceptable to treat such constraints LOOSELY — if a flight is overbooked, you can move customers to a different flight and OFFER COMPENSATION. In such cases linearizability may not be needed."* (Ch 13.)
> **A HARD uniqueness constraint requires linearizability. Other kinds of constraints — foreign-key or attribute constraints — CAN be implemented without it.**
**③ Cross-channel timing dependencies — the subtle one.**
> **Notice: if Aaliyah hadn't exclaimed the score, Bryce WOULDN'T HAVE KNOWN his result was stale. The violation was noticed ONLY BECAUSE THERE WAS AN ADDITIONAL COMMUNICATION CHANNEL** (Aaliyah's voice to Bryce's ears).
*(The same race occurs with **push notifications**: the notification arrives quickly, but **the subsequent data fetch goes to a lagging replica and doesn't see the data the notification was about.**)*
> **Linearizability is not the only way to avoid this, but it's the simplest to understand. IF YOU CONTROL THE ADDITIONAL COMMUNICATION CHANNEL** (as with the message queue, but not with Aaliyah and Bryce) **you can use alternatives similar to "reading your own writes" (Ch 6), AT THE COST OF ADDITIONAL COMPLEXITY.**
#### 1.5 Which replication methods are linearizable? [#15-which-replication-methods-are-linearizable]
| Method | Verdict | Detail |
| ------------------------ | ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Single-leader** | **POTENTIALLY** | Linearizable **as long as all reads and writes go to the leader — AND YOU KNOW FOR SURE WHO THE LEADER IS.** A node may think it's the leader when it isn't, and **a delusional leader that keeps serving requests is likely to violate linearizability.** **With asynchronous replication, failover may result in COMMITTED WRITES BEING LOST, violating both durability and linearizability.** *(Sharding doesn't affect it — it's a single-object guarantee.)* |
| **Consensus algorithms** | **LIKELY** | **Essentially single-leader replication with AUTOMATIC LEADER ELECTION AND FAILOVER, carefully designed to prevent split brain.** ZooKeeper uses **Zab**, etcd uses **Raft**. ⚠️ **But using consensus does not GUARANTEE all operations are linearizable: if it allows reads on a node WITHOUT CHECKING IT IS STILL THE LEADER, results may be stale** |
| **Multi-leader** | **NOT** | **Concurrently processes writes on multiple nodes and asynchronously replicates**, producing conflicting writes |
| **Leaderless** | **PROBABLY NOT** | See below |
**Why quorums are NOT automatically linearizable — the counterexample:**
**How to fix it — and what it costs:**
* **The reader must perform READ REPAIR SYNCHRONOUSLY before returning results**
* **Before writing, the writer must READ THE LATEST STATE OF A QUORUM to fetch the greatest prior timestamp and ensure the new write has a greater one**
> **Riak does NOT perform synchronous read repair because of the performance penalty. Cassandra DOES wait for read repair on quorum reads — but IT LOSES LINEARIZABILITY BECAUSE OF ITS USE OF TIME-OF-DAY CLOCKS for timestamps** (Ch 9 §3.4).
>
> **What's more, ONLY LINEARIZABLE READS AND WRITES can be implemented this way; A LINEARIZABLE CAS CANNOT, BECAUSE IT REQUIRES A CONSENSUS ALGORITHM.**
>
> ### **It is safest to assume that a leaderless Dynamo-style system does NOT provide linearizability, even with quorum reads and writes.** [#it-is-safest-to-assume-that-a-leaderless-dynamo-style-system-does-not-provide-linearizability-even-with-quorum-reads-and-writes]
#### 1.6 The cost of linearizability [#16-the-cost-of-linearizability]
**The CAP theorem, stated honestly:**
| Choice | Meaning |
| -------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **CP** — consistent under network partitions | **If your application requires linearizability and some replicas are disconnected, THOSE REPLICAS WILL BE TEMPORARILY UNABLE TO PROCESS REQUESTS: they must either wait or return an error — EITHER WAY, THEY BECOME UNAVAILABLE** |
| **AP** — available under network partitions | **If linearizability isn't required, each replica can process requests INDEPENDENTLY even when disconnected. The application REMAINS AVAILABLE, but ITS BEHAVIOR IS NOT LINEARIZABLE** |
> ### **The unhelpful framing: "pick two out of three" is MISLEADING. Network partitions are A KIND OF FAULT — they aren't something you CHOOSE, they WILL HAPPEN whether you like it or not. The only way to guarantee no partitions is to have NO NETWORK — that is, only one replica — but then you don't have high availability either.** [#the-unhelpful-framing-pick-two-out-of-three-is-misleading-network-partitions-are-a-kind-of-fault--they-arent-something-you-choose-they-will-happen-whether-you-like-it-or-not-the-only-way-to-guarantee-no-partitions-is-to-have-no-network--that-is-only-one-replica--but-then-you-dont-have-high-availability-either]
>
> **A better phrasing: EITHER CONSISTENT OR AVAILABLE WHEN PARTITIONED. A more reliable network makes this choice less often, but at some point THE CHOICE IS INEVITABLE.**
**The book's verdict on CAP — read this carefully:**
> **CAP as formally defined is OF VERY NARROW SCOPE. It considers only ONE consistency model (linearizability) and ONE kind of fault (network partitions, which per Google data cause LESS THAN 8% OF INCIDENTS). It says nothing about network delays, dead nodes, or other trade-offs.**
>
> **The formalization of AVAILABILITY does not match the usual meaning of the term. MANY HIGHLY AVAILABLE (fault-tolerant) SYSTEMS DO NOT MEET CAP'S IDIOSYNCRATIC DEFINITION OF AVAILABILITY. Moreover, some system designers choose (WITH GOOD REASON) to provide NEITHER linearizability NOR CAP's form of availability — SO THOSE SYSTEMS ARE NEITHER CP NOR AP.**
>
> **CAP DESERVES CREDIT for a culture shift — it helped trigger the NoSQL movement — but IT HAS LITTLE PRACTICAL VALUE FOR DESIGNING SYSTEMS TODAY, and has been SUPERSEDED BY MORE PRECISE RESULTS. There is a lot of misunderstanding and confusion around CAP, and IT DOES NOT HELP US UNDERSTAND SYSTEMS BETTER, SO IT'S BEST NOT TO DWELL ON IT.**
*(**PACELC** generalizes it: during a **P**artition choose **A** or **C**; **E**lse, when there's no partition, choose **L**atency or **C**onsistency. **But it inherits several of CAP's problems, such as the counterintuitive definitions.**)*
#### 1.7 The REAL reason linearizability is rare: latency, not fault tolerance [#17-the-real-reason-linearizability-is-rare-latency-not-fault-tolerance]
> **Surprisingly few systems are linearizable in practice. EVEN RAM ON A MODERN MULTI-CORE CPU IS NOT LINEARIZABLE.** If a thread on one core writes to a memory address and a thread on another core reads it shortly after, **it is not guaranteed to read the value written** — unless a **memory barrier / fence** is used.
>
> **Why? Every CPU core has its OWN CACHE AND STORE BUFFER. Reads are served from the cache; changes are asynchronously written to main memory. There are now MULTIPLE COPIES OF THE DATA, asynchronously updated ⇒ LINEARIZABILITY IS LOST.**
>
> ### **It makes NO SENSE to use the CAP theorem to justify the multi-core memory consistency model. Within one computer we assume reliable communication, and we don't expect one CPU core to keep operating if disconnected from the rest. THE REASON FOR DROPPING LINEARIZABILITY IS PERFORMANCE, NOT FAULT TOLERANCE.** [#it-makes-no-sense-to-use-the-cap-theorem-to-justify-the-multi-core-memory-consistency-model-within-one-computer-we-assume-reliable-communication-and-we-dont-expect-one-cpu-core-to-keep-operating-if-disconnected-from-the-rest-the-reason-for-dropping-linearizability-is-performance-not-fault-tolerance]
>
> **The same is true of many distributed databases: they drop linearizable guarantees PRIMARILY TO INCREASE PERFORMANCE, not so much for fault tolerance. LINEARIZABLE SYSTEMS TEND TO BE HIGHER LATENCY — ALL THE TIME, NOT ONLY DURING A NETWORK FAULT.**
**And there is no clever way out:**
> **Attiya and Welch prove that IF YOU WANT LINEARIZABILITY, THE RESPONSE TIME OF READ AND WRITE REQUESTS IS AT LEAST PROPORTIONAL TO THE UNCERTAINTY OF DELAYS IN THE NETWORK.** In a network with highly variable delays — most computer networks — **the response time of linearizable reads and writes is INEVITABLY GOING TO BE HIGH. A FASTER ALGORITHM FOR LINEARIZABILITY DOES NOT EXIST.**
***
# 10.6 Production failure catalog for this chapter (/docs/ddia/consistency-consensus/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ------------------------------------------------------- | ------------------------------------------------------------------------ |
| Two users see different "current" values seconds apart | **Not linearizable** — stale replica read |
| A transcoder processes an old version of a file | **Cross-channel race**: message queue faster than storage replication |
| Push notification arrives before the data it references | Same — two channels, no recency guarantee |
| Two nodes both act as leader | **Split brain** — leader election without consensus |
| A username was registered twice | Uniqueness enforced without **linearizable CAS** |
| Committed writes lost after failover | **Asynchronous replication** + promotion; or **unclean leader election** |
| Quorum reads still return stale values | **w + r > n does NOT imply linearizability** |
| Cassandra "strong consistency" isn't | **Time-of-day-clock LWW** breaks linearizability |
| Reads from a consensus cluster are stale | **Read served without a quorum check** that the leader is still current |
| ZooKeeper read returned old data | **ZooKeeper reads are not linearizable** without `sync` |
| Photo visible despite a prior privacy change | **Non-linearizable ID generator** + MVCC snapshot |
| Cluster spends all its time electing leaders | **Election timeout** too small vs real jitter/GC |
| Leadership bounces between two nodes forever | **Raft edge case** on one bad link (fixed by pre-vote) |
| Adding nodes made the cluster slower | Consensus needs a quorum per operation — **more nodes = slower** |
| Minority partition frozen | **Consensus requires a strict majority** — by design |
| etcd went read-only, Kubernetes stopped | **Storage quota exceeded** |
| Duplicate primary keys across the fleet | **Colliding machine IDs** in a Snowflake-style generator |
| ID generator refuses to issue IDs | **Clock stepped backward** |
| Every write got slower after a GPS fault | **TrueTime ε widened** → longer commit wait |
| Coordination service melted under load | **Used as a database** for fast-changing data |
***
# 10.9 Self-test (/docs/ddia/consistency-consensus/self-test)
State the two philosophies for handling replica inconsistency and one situation where each is the only viable choice.
Define linearizability in one sentence. Why is it called a *recency* guarantee?
In the sports example, why does the violation depend on Aaliyah *speaking*? What is that channel called in general?
Given a write concurrent with three reads, which reads *must* return the old value, which *must* return the new, and which may return either?
State the extra constraint beyond "concurrent reads may return either." What would go wrong without it?
Distinguish linearizability from serializability on three axes. What is the combination called?
Give three categories of situation that genuinely require linearizability, and one type of constraint that does *not*.
Draw the video-transcoder race. Why does linearizable storage fix it, and what is the alternative if you control the queue?
For each replication method, say whether it can be linearizable and what breaks it.
Reproduce the quorum counterexample. What two changes would make Dynamo-style quorums linearizable, and what can *never* be made linearizable that way?
State the CAP trade-off properly. Why is "pick two of three" misleading?
Give three specific criticisms the book makes of CAP as a design tool.
Why is RAM on a multi-core CPU not linearizable? What does that tell you about *why* systems drop linearizability?
State the Attiya–Welch result and its practical consequence.
Give four distributed ID schemes and say precisely what ordering property each loses.
State the three requirements of a logical clock. Which one do sharded/UUID/timestamp schemes fail?
Give the two Lamport clock update rules and the comparison rule. Trace the chat example.
Give two limitations of Lamport clocks, and explain how HLCs fix each.
Why can't you tell from two Lamport timestamps whether the events were concurrent? What data structure can, and what does it cost?
Walk through the privacy-leak example. Which exact property of linearizability is violated, and why can't an HLC provide it?
Describe the batched timestamp oracle. What does batching sacrifice, and what does it preserve?
Why can't you shard a linearizable ID generator?
Why is a linearizable ID generator still insufficient for fault-tolerant locking?
What does the FLP result actually prove? Name the two ways real systems escape it.
State the four properties of single-value consensus. Which is the liveness property, and what is the minimum-node requirement it imposes?
Show both directions of the CAS ⟺ consensus equivalence.
State the five properties of a shared log. Sketch both directions of shared-log ⟺ consensus.
Why does fetch-and-add fail to solve consensus for three nodes but succeed for two? What is the term for this?
What is the one crucial difference between consensus and atomic commitment?
List six things you can build on top of a shared log.
How do consensus algorithms escape "you need a leader to elect a leader"? Name the epoch number in three algorithms.
Explain the quorum-overlap requirement between the two voting rounds. What does it prevent?
Give two ways consensus voting differs from 2PC.
What is unclean leader election, what does it buy, and what does it cost?
Why must linearizable *reads* also go through a quorum?
List five costs of consensus. Which one means you cannot scale throughput by adding nodes?
Which coordination-service features require consensus and which don't? Why is it still convenient to get the latter from the same service?
What is the architectural argument for a *fixed* 3–5 node coordination service in a system with thousands of shards?
Why is using consensus for service discovery usually overkill? What are ZooKeeper observers for?
you're building a multi-region ticketing platform. Requirements: (a) a seat is sold at most once globally, (b) a user's "my tickets" page must show a purchase immediately after buying, (c) browsing seat availability must stay fast and available even during a region partition, (d) event IDs must be sortable by creation time. For each requirement, state which consistency guarantee you need, which mechanism provides it, what it costs in latency, and what degrades during a partition. Identify at least one requirement you would deliberately weaken and justify it.
# 10.5 Technology deep dives (/docs/ddia/consistency-consensus/technology-deep-dives)
***
#### 5.1 Raft (etcd, CockroachDB, TiKV, Consul, RabbitMQ quorum queues) [#51-raft-etcd-cockroachdb-tikv-consul-rabbitmq-quorum-queues]
**Problem it solves.** Replicate a log across a set of nodes with automatic, *safe* leader failover — so that no acknowledged write is ever lost and split brain is impossible.
**Why wasn't single-leader replication enough?** It has no safe automatic failover: promoting a lagging follower loses committed writes; two nodes can both believe they're leader (Ch 6 §1.4).
**Why wasn't Paxos enough?** Raft was explicitly designed for *understandability* — Paxos's Multi-Paxos extension is what people actually deploy, and it's notoriously underspecified in the literature. Raft prescribes leader election, log replication, and membership change concretely.
**How it works internally.** Three roles (follower / candidate / leader). **Terms** are the epoch numbers: monotonically increasing, at most one leader per term. Election timeout randomized (typically 150–300 ms) to avoid split votes. **Election restriction:** a candidate only wins if its log is at least as up-to-date as the voter's — this is Raft's version of §3.5's "new leader must be up to date." A leader appends to its log, replicates via `AppendEntries` (which doubles as the heartbeat), and **commits an entry once a majority have it** — *and only entries from its own current term directly*, which is the subtle rule that prevents a committed entry from being overwritten. `PreVote` was added to fix the §3.6 leadership-bouncing pathology.
**Deployment.** 3 or 5 members (odd, so a majority exists), **spread across AZs but rarely across regions** — every write costs one quorum round trip, so cross-region raises write latency to the inter-region RTT. Learners/non-voting members for adding capacity or a new region without changing the quorum.
**Monitoring.**
* **Leader elections per hour** — should be \~0. Nonzero and rising means your election timeout is fighting your GC pauses or network jitter (§3.6 cost #4)
* **Commit latency p99** and **`fsync` latency on the WAL** — the latter bounds everything
* **Follower lag / `matchIndex` gap** per member — a lagging follower is a member that can't become leader safely
* **Proposal failure rate**; **quorum-loss events**
* etcd specifically: **DB size vs quota** (a full etcd goes read-only — a classic Kubernetes outage), compaction and defrag status, **watch count**
**Backup.** Periodic **snapshots** + the log. **Losing a majority of members loses the cluster**, so snapshots must live off-cluster.
**What actually breaks.**
* **Election storms.** Timeouts too aggressive relative to real network/GC behavior; the cluster spends more time electing than working — verbatim §3.6.
* **`fsync` latency spikes** on a slow disk stalling every write cluster-wide, because commit requires a durable quorum.
* **etcd exceeding its storage quota** and flipping to read-only, taking Kubernetes with it.
* **Cross-region deployment** making every write pay an inter-continental RTT.
* **Losing quorum** (2 of 3 members) — the cluster is *safe* but *unavailable*, and recovery requires a deliberate, dangerous force-new-cluster operation.
* **Reading from a follower and assuming linearizability** — §1.5's warning: consensus doesn't make *all* operations linearizable.
***
#### 5.2 ZooKeeper (Zab) and the coordination-service pattern [#52-zookeeper-zab-and-the-coordination-service-pattern]
**Problem it solves.** Give *other* distributed systems the primitives they need — leases, fencing tokens, failure detection, change notification — without each of them implementing consensus.
**Why not embed consensus in every system?** §4: with thousands of shards, running consensus over thousands of nodes is terribly inefficient. **Outsource it to 3–5 nodes.**
**How it works internally.** **Zab** (a Paxos-family protocol with a strong ordering guarantee). A hierarchical namespace of **znodes**; **ephemeral** znodes tied to a session (the failure detector); **sequential** znodes producing monotonic counters (the fencing tokens); **watches** as one-shot change notifications. `zxid` is a 64-bit value: **high 32 bits = epoch, low 32 bits = counter within the epoch** — so it is simultaneously a fencing token and a leadership generation marker.
**Deployment.** Ensemble of 3 or 5; **observers** for read scaling and remote regions (§4). Session timeout tuned above worst-case GC pause. Apache Curator for recipes (leader election, distributed locks, barriers) — **because the raw API is easy to misuse.**
**Monitoring.** Outstanding requests; **session expirations per hour** (each is a potential zombie — Ch 9 §5.2); watch count (unbounded watches are a known way to melt an ensemble); znode count and data size; **`fsync` time**; leader election count; follower sync latency.
**What actually breaks.**
* **Using ZooKeeper reads as linearizable.** §1.4's footnote: **writes are linearizable, reads may be stale.** You must issue a `sync` before a read if you need recency.
* **Watch semantics misunderstood** — watches are *one-shot* and *may miss intermediate states*; code that assumes it sees every change is wrong.
* **Session timeout shorter than a GC pause** → the ephemeral node vanishes → the lease is reassigned → zombie.
* **Storing too much or too fast-changing data** — §4's constraint. This is not a database.
* **Watch explosion** from a client registering a watch per key across a large keyspace.
* **A herd effect** when a leader's ephemeral node disappears and every candidate wakes at once (Curator's recipes exist to avoid exactly this).
***
#### 5.3 Spanner / TrueTime — linearizability without a coordination round trip [#53-spanner--truetime--linearizability-without-a-coordination-round-trip]
**Problem it solves.** Globally distributed strict serializability where a single-node timestamp oracle (§2.4) would force every transaction in every region through one place.
**Why not a timestamp oracle?** §2.4's limits: can't shard it, can't distribute it across regions — **in a geo-distributed database, all ID requests would go to a node in a single region.**
**Why not HLCs?** §2.3: they give causal ordering, not linearizability — the privacy-leak example is exactly the failure.
**How it works internally.** **TrueTime** returns `[earliest, latest]` from GPS + atomic clocks per datacenter, keeping ε ≈ 7 ms (Ch 9 §3.6). A read/write transaction picks a commit timestamp and then **waits out ε before releasing locks** ("commit wait"), guaranteeing that any later transaction gets a strictly greater timestamp **without any communication.** Reads at a timestamp are lock-free and served from any sufficiently up-to-date replica.
**Monitoring.** **TrueTime ε** (if it grows, commit-wait grows and every write slows — this is a *hardware health* metric with a direct latency consequence); commit-wait duration; per-Paxos-group leader locality; transaction retry/abort rate; **participant count per transaction** (2PC over Paxos groups: more groups = more expensive).
**What actually breaks.** ε widening after a GPS or time-master failure, silently taxing every write; cross-region transactions costing consensus round trips plus commit wait; and the general trap of assuming a globally distributed transaction is as cheap as a local one.
***
#### 5.4 Distributed ID generation (Snowflake, UUIDv7, ULID, timestamp oracles) [#54-distributed-id-generation-snowflake-uuidv7-ulid-timestamp-oracles]
**Problem it solves.** Unique primary keys, generated without a bottleneck, ideally sortable by creation time.
**Why not autoincrement?** §2: SPOF, cross-region latency, throughput bottleneck.
**Why not UUIDv4?** §2.1: 128 bits and **random order — comparing two tells you nothing about which is newer**, which also destroys B-tree insert locality (Ch 4 §9.2: random inserts into a clustered index cause page splits and poor fill factor).
**How Snowflake-style works internally.** `[41 bits ms timestamp][10 bits machine ID][12 bits sequence]` = 64 bits, k-sortable, \~4,096 IDs/ms/node. **UUIDv7** is the standardized version of the same idea: 48-bit Unix ms timestamp in the high bits, then random. **ULID** is the same shape with Crockford base32 text encoding.
**Monitoring.** Sequence exhaustion (>4,096/ms on one node → the generator must wait or error); **clock rollback events** — the failure that forces a Snowflake node to *refuse to issue IDs*; machine-ID collisions (the nastiest failure, and easy to cause with autoscaling that reuses IDs).
**What actually breaks.**
* **Clock moving backward** — Snowflake's only safe response is to **stop issuing IDs** until the clock catches up, which is a hard outage caused by NTP.
* **Duplicate machine IDs** after an autoscaler recycles an instance ordinal → **duplicate primary keys**, discovered much later.
* **Assuming timestamp-prefixed IDs are causally ordered.** They are approximately time-ordered and **not linearizable** (§2.1) — which is precisely the privacy-leak trap.
* **Hot shard from monotonic keys** (Ch 7 §3.1): the very sortability that helps B-trees sends all writes to one shard.
***
# 10.10 Terminology introduced here (/docs/ddia/consistency-consensus/terminology-introduced-here)
# 10.8 Worked examples (/docs/ddia/consistency-consensus/worked-examples)
**① Is this history linearizable?**
**No.** C read 1; D's read *begins after C's completes*, so D must return 1 or newer. Returning 0 moves the linearization point **backward in time** — the one thing forbidden. (B returning 0 is fine: it overlaps the write.)
**② The quorum counterexample, formalized.** n=3, w=3, r=2. A's read set \{r1,r2} ∩ writer's set = nonempty ✔. B's read set \{r2,r3} ∩ writer's set = nonempty ✔. **Both satisfy the intersection property, yet B (later) reads older than A.** The intersection guarantee only says *some replica you read has seen the latest completed write* — **it says nothing about which value you return when replicas disagree, nor about ordering between two reads.**
**③ Lamport vs linearizable.** Node P and node Q never communicate. P performs op₁ at wall-clock 10:00:00, gets Lamport ts (1, P). Q performs op₂ at 10:00:05 — genuinely later — gets (1, Q). Comparing: **(1,P) \< (1,Q)** by node-ID tiebreak, so the order *happens* to be right. Now swap the node names: **the order is wrong**, and nothing detects it. **Lamport ordering is only meaningful along communication paths.**
**④ Cost of a quorum read.** 3 nodes, intra-AZ RTT 0.5 ms, inter-AZ 1.5 ms. A linearizable read = leader must confirm leadership with a quorum ⇒ **≥ 1 inter-AZ RTT ≈ 1.5 ms**, versus \~0.05 ms for a local stale read — **30×**. Across regions (RTT 70 ms), the same read is **1,400× slower.** This is Attiya–Welch made concrete: **response time proportional to network delay uncertainty.**
**⑤ Fetch-and-add's consensus number.** 3 proposers, counter starts at 0. P reads 0, Q reads 1, R reads 2. **Q and R know they lost but not who won.** If P crashes before announcing, Q and R can neither decide P's value (they don't know it) nor decide their own (P might return). **Termination fails ⇒ consensus number \< 3.** With 2 proposers, exchanging values *first* makes the loser able to infer the winner's value — **consensus number exactly 2.**
**⑥ Why quorums must overlap across the two votes.** Leader L₁ (epoch 5) is partitioned. L₂ elected in epoch 6 with quorum \{A,B,C}. L₁ tries to append with quorum \{C,D,E}. **The intersection is \{C}, and C has seen epoch 6 ⇒ C refuses ⇒ L₁'s append fails.** If the quorums could be disjoint (\{A,B,C} and \{D,E,F} on 6 nodes), **both leaders could append conflicting entries — split brain.** This single overlap requirement is what makes epochs safe.
**⑦ Election timeout budgeting.** Worst observed GC pause 800 ms; network p99.9 RTT 20 ms. An election timeout of 500 ms → **every long GC on the leader triggers an election** → §3.6's election storm. Set it above the worst pause (e.g. 1,500 ms), accept slower failover, *or* fix the pauses (Ch 9 §4.3). **The timeout is a statement about your worst-case pause, not about your network.**
***
# 3.4 DataFrames, Matrices, and Arrays (/docs/ddia/data-models-query-languages/dataframes-matrices-arrays)
Models you'll meet in **analytical or scientific** contexts that **rarely feature in OLTP**.
**DataFrames** are supported by **R, Pandas (Python), Apache Spark, ArcticDB, Dask**. Popular for **preparing data to train ML models**, and widely used for data exploration, statistical analysis, and visualization.
**Superficially like a relational table or spreadsheet.** Supports relational-like bulk operators: apply a function to all rows, filter on a condition, group by columns and aggregate others, and **join** — *which on DataFrames is typically called **merge***.
**The key difference:** instead of a declarative query language, **a DataFrame is manipulated through a series of commands that modify its structure and content.** This matches how data scientists work: **incrementally "wrangling" data into a form that answers their question**, usually on **a private copy of the dataset, often on their local machine**, with the end result possibly shared.
**DataFrame APIs go far beyond relational databases**, and the model is often used in very un-relational ways. **A common use: transform data from a relational-like representation into a matrix or multidimensional array — the form many ML algorithms expect.**
**Getting non-numeric data into a matrix (a matrix can contain only numbers):**
* **Dates** → **scaled to floating-point numbers within a suitable range**
* **Categorical columns with a small fixed set of values** (movie genre) → **one-hot encoding**: create a column per possible value ("comedy", "drama", "horror"), put **1** in the column matching the row's genre and **0** in the others. **This generalizes easily to movies that fit several genres** (multiple 1s).
Once numeric, the data is amenable to **linear algebra operations, which form the basis of many ML algorithms** — e.g. a movie recommender. **DataFrames are flexible enough to let data gradually evolve from a relational form into a matrix representation, while giving the data scientist control over the most suitable representation.**
**Array databases** (e.g. **TileDB**) specialize in **large multidimensional arrays of numbers**, used for **scientific datasets: geospatial raster data on a regularly spaced grid, medical imaging, astronomical telescope observations.** DataFrames are also used in **finance for time-series data** (asset prices and trades over time). Because of their popularity, **DataFrames have been added to batch frameworks like Spark and Flink** (Ch 11).
***
# 3.7 Decision cheat sheet (/docs/ddia/data-models-query-languages/decision-cheat-sheet)
**Relational or document?**
Document if the data is a **tree of one-to-many** relationships typically loaded whole, the items are genuinely one-to-**few**, and you rarely need to reference nested items directly. Relational if you have many-to-one/many-to-many, need to address items by ID, or need joins.
**In practice, use a hybrid** — Postgres with `jsonb` columns is the default correct answer for most applications, and the convergence trend says so explicitly.
**Normalize or denormalize this specific field?**
Ask two questions: **How fast does it change?** (fast → normalize + hydrate) and **What dominates cost — reads or writes, and are they dominated by outliers?** Don't answer per-table; answer per-field. The X timeline denormalizes the *join result* and normalizes the *contents*.
**Schema-on-write or schema-on-read?**
Schema-on-write when records are expected to have the same structure — a schema then documents and enforces it. Schema-on-read when data is **heterogeneous** (too many object types to table each one) or the structure is **controlled by an external system that may change at any time**.
**When do I actually need a graph database?**
When queries traverse **a variable number of hops not known in advance**, over data with **arbitrary many-to-many connectivity**. If your traversals are always exactly 2 hops, a relational join is fine and cheaper. Beware supernodes and don't plan on sharding.
**Should I use event sourcing?**
Yes if: intent matters, auditability is required, you need multiple divergent read models, or reversibility is valuable. No if: the domain is CRUD, you have no need for history, or you can't commit to keeping event-processing **deterministic** and **replayable** forever. Half-hearted event sourcing is worse than none.
**Star, snowflake, or OBT?**
Star by default — simpler for analysts. Snowflake when dimension normalization actually pays. OBT when storage is cheap and you need the join gone.
***
# 3.3 Event Sourcing and CQRS (/docs/ddia/data-models-query-languages/event-sourcing-cqrs)
**The setup:** in every model so far, **data is queried in the same form it is written.** But in complex applications it can be hard to find a single representation satisfying all the ways data needs to be queried and presented.
**The move:** write data in one form, then **derive representations optimized for different types of reads.** We saw this with systems-of-record vs derived data, and ETL. Now push it further:
> **If we're going to derive one representation from another anyway, we can choose representations optimized for writing and reading, respectively. How would you model data if you wanted to optimize it for ONLY writing, with efficient queries of no concern?**
**Answer: an event log.** Encode each write as a **self-contained string (perhaps JSON) including a timestamp**, and **append** it. **Events are immutable — never changed or deleted, only appended to (later events may supersede earlier ones).** An event can contain arbitrary properties.
#### 3.1 The conference example [#31-the-conference-example]
A conference management system is a genuinely complex domain:
* Individual attendees register and pay by card
* **Companies order seats in bulk, pay by invoice, and later assign seats to individuals**
* Seats reserved for speakers, sponsors, volunteers
* Reservations may be canceled
* **The organizer might change the capacity by moving to a different room**
> **With all this going on, simply calculating the number of available seats becomes a challenging query.**
**Definitions:**
* **Event sourcing** — using events as the **source of truth**, expressing **every state change as an event**
* **CQRS (command query responsibility segregation)** — maintaining **separate read-optimized representations derived from the write-optimized representation**
Both terms originated in the **DDD community**, though similar ideas are old — e.g. **state machine replication**.
**The command→event lifecycle:**
1. A request from a user is a **command** — it must first be **validated**
2. Once executed and **determined to be valid** (e.g. there were enough seats), **it becomes a fact**, and the corresponding **event is appended to the log**
3. **Therefore the event log contains only valid events, and a consumer building a materialized view is NOT ALLOWED TO REJECT AN EVENT**
> **Name your events in the PAST TENSE** ("the seats were booked") — an event records that something *has happened*. **Even if the user later cancels, the fact remains true that they formerly held a booking**; the cancellation is a **separate event added later.**
**Event sourcing vs star-schema fact table** — similar (both are collections of past events), but:
| | Fact table | Event log |
| ----- | ----------------------------------------- | ---------------------------------------------------------------------------------------------- |
| Shape | **All rows have the same set of columns** | **Many event types, each with different properties** |
| Order | **Unordered collection** | **Order is important** — a booking made then canceled must not be processed in the wrong order |
#### 3.2 Advantages [#32-advantages]
1. **Events communicate intent.** "The booking was canceled" is far easier to understand than *"the `active` column on row 4001 of `bookings` was set to false, three rows were deleted from `seat_assignments`, and a refund row was inserted into `payments`."* **Those row modifications may still happen when a view processes the event — but driven by an event, the reason for the updates becomes much clearer.**
2. **Reproducibility.** A key principle: views are derived from the log **in a reproducible way.** You should **always be able to delete the materialized views and recompute them by processing the same events in the same order with the same code.** If the view-maintenance code had a bug, **delete the view and recompute with the fixed code.** Finding the bug is easier too, because **you can rerun the view-maintenance code as often as you like and inspect its behavior.**
3. **Multiple views, each optimized for particular queries.** Stored in the same database as the events or a different one; **any data model**; **denormalized for fast reads**. You can even **keep a view only in memory and never persist it**, as long as recomputing from the log on restart is acceptable.
4. **Easy evolution.** Present existing information in a new way → build a new view from the existing log. Support new features → **add new event types or new properties to existing types (older events remain unmodified).** **Chain new behaviors off existing events** — e.g. when an attendee cancels, offer their seat to the next person on the waiting list.
5. **Reduced irreversibility.** If an event was written in error, **write a subsequent deletion event to reverse it; downstream views incorporate it automatically and correct the data.** In a database where you update and delete directly, **a committed transaction is often difficult to reverse.** This ties straight back to Ch 2's *irreversibility is the main obstacle to evolvability.*
6. **Audit log** of what has occurred — valuable in **regulated industries requiring auditability.**
7. **Higher write throughput than databases, because of sequential access patterns.** A temporary burst is absorbed by the log, and **downstream view maintainers catch up at their own pace without being overwhelmed** — natural backpressure.
#### 3.3 Downsides (all three are real and commonly underestimated) [#33-downsides-all-three-are-real-and-commonly-underestimated]
1. **External information breaks determinism.** An event contains a price in one currency; a view needs it converted. **Fetching the exchange rate from an external source at processing time is wrong — you'd get a different result if you recomputed the view on another date.** To keep processing **deterministic** you must either **include the exchange rate in the event itself**, or **have a way of querying the historical rate at the event's timestamp that always returns the same result for the same timestamp.**
2. **Immutability vs GDPR.** Users may **request deletion of their data**. If the log is per-user you can delete that user's whole log — **but that doesn't work if the log contains events relating to multiple users.** Options: **store personal data outside the event**, or **encrypt it with a key you can later delete (crypto-shredding)** — **but both make it harder to recompute derived state when needed.**
3. **Reprocessing with externally visible side effects.** **You probably don't want to resend confirmation emails every time you rebuild a materialized view.**
#### 3.4 Implementation [#34-implementation]
Implementable on any database. Purpose-built: **EventStoreDB, MartenDB (on PostgreSQL), Axon Framework.** You can also use **Apache Kafka** to store the event log with **stream processors** keeping views up to date (Ch 12).
> **The only important requirement: the event storage system must guarantee that ALL materialized views process the events in EXACTLY the same order as they appear in the log.** As Ch 10 shows, **this is not always easy to achieve in a distributed system.**
***
# 3.11 Forward links (/docs/ddia/data-models-query-languages/forward-links)
| Concept here | Where it's developed |
| ----------------------------------------------------------------- | ------------------------------------- |
| How documents/rows/graphs become bytes | **Ch 4** — Storage and Retrieval |
| Secondary indexes into document fields | **Ch 4** |
| Column-oriented storage for star schemas | **Ch 4** |
| Full-text search and vector search | **Ch 4** — "Full-Text Search" |
| Schemas, schema evolution, why JSON is problematic as an encoding | **Ch 5** — Encoding and Evolution |
| Atomicity across multiple documents | **Ch 8** — Transactions |
| Keeping denormalized copies consistent via streams | **Ch 12** — Stream Processing |
| Guaranteeing all views see events in the same order | **Ch 10** — Consistency and Consensus |
| DataFrames in batch frameworks | **Ch 11** — Batch Processing |
| State machine replication / shared logs | **Ch 10** |
**Models the chapter deliberately leaves out:** **sequence similarity search** for genome data (specialized software like GenBank); **double-entry accounting ledgers** (TigerBeetle; and distributed ledgers in cryptocurrencies/blockchains, which build value transfer into the data model); and **full-text search**, a large specialist subject touched on in Ch 4.
# 3.2 Graph-Like Data Models (/docs/ddia/data-models-query-languages/graph-like-data-models)
**The decision rule:** mostly one-to-many (tree-structured) with few other relationships → **document**. **Many-to-many relationships very common, connections increasingly complex → graph.**
A graph has **vertices** (nodes, entities) and **edges** (relationships, arcs).
| Graph type | Vertices | Edges |
| ----------------- | --------- | ------------------- |
| Social graph | people | who knows whom |
| Web graph | web pages | HTML links |
| Road/rail network | junctions | roads/railway lines |
Well-known algorithms operate on them: **shortest path** for map navigation; **PageRank** on the web graph for page popularity and search ranking.
**Two representations:**
**Graphs are not limited to homogeneous data** — an equally powerful use is **storing completely different types of objects in a single database**:
* **Facebook maintains a single graph** with many vertex and edge types: vertices are people, locations, events, check-ins, and comments; edges say who is friends with whom, which check-in happened at which location, who commented on which post, who attended which event.
* **Search engines use knowledge graphs** to record facts about entities common in queries — organizations, people, places — obtained by crawling and analyzing website text. Some sites (**Wikidata**) publish graph data in structured form directly.
**Running example** (Figure 3-6): Lucy from Idaho and Alain from Saint-Lô, France; married, living in London. Each person and location is a vertex; relationships are edges.
#### 2.1 Property graphs [#21-property-graphs]
Implemented by **Neo4j, Memgraph, KùzuDB**, and others.
**Each vertex has:** a unique identifier · a **label** (string) describing the type of object · a set of **outgoing** edges · a set of **incoming** edges · a collection of **properties** (key-value pairs).
**Each edge has:** a unique identifier · the **tail vertex** (where it starts) · the **head vertex** (where it ends) · a **label** describing the kind of relationship · a collection of **properties**.
You can think of a graph store as **two relational tables**:
```sql
CREATE TABLE vertices (
vertex_id integer PRIMARY KEY,
label text,
properties jsonb
);
CREATE TABLE edges (
edge_id integer PRIMARY KEY,
tail_vertex integer REFERENCES vertices (vertex_id),
head_vertex integer REFERENCES vertices (vertex_id),
label text,
properties jsonb
);
CREATE INDEX edges_tails ON edges (tail_vertex);
CREATE INDEX edges_heads ON edges (head_vertex);
```
**Three important aspects:**
1. **Any vertex can have an edge connecting it with any other vertex.** There is **no schema restricting which kinds of things can be associated.**
2. **Given any vertex, you can efficiently find both its incoming and outgoing edges**, and thus traverse the graph **both forward and backward** — that's exactly why there are indexes on *both* `tail_vertex` and `head_vertex`.
3. **Different labels for different kinds of vertices and relationships let you store several kinds of information in a single graph while keeping a clean data model.**
> The `edges` table is the **many-to-many join table, generalized to allow many types of relationship in the same table.** There may also be indexes on labels and properties.
**A real limitation:** an edge associates only **two** vertices, whereas a relational join table can represent **three-way or higher-degree relationships** via multiple FK references on one row. Workarounds: **create an extra vertex per join-table row** with edges to/from it, or use a **hypergraph**.
**Why graphs are good for evolvability** — the example is subtle and worth keeping:
* **Different regional structures in different countries** (France has *départements* and *régions*; the US has counties and states)
* **Quirks of history** such as a country within a country
* **Varying granularity of data** — Lucy's current residence is a *city*, but her birthplace is specified only at the level of a *state*
All three are difficult in a traditional relational schema, and trivial in a graph. Then: add food allergies (a vertex per allergen, an edge person→allergen), link allergens to foods containing them, and **query what's safe for each person to eat.** **As you add features, a graph easily extends to accommodate changes in the application's data structures.**
#### 2.2 Cypher [#22-cypher]
Query language for property graphs, originally from Neo4j, now the open standard **openCypher**. Supported by Neo4j, Memgraph, KùzuDB, Amazon Neptune, Apache AGE (storage in PostgreSQL). *(Named after the character in The Matrix; unrelated to cryptographic ciphers.)*
**Insert:**
```cypher
CREATE
(namerica :Location {name:'North America', type:'continent'}),
(usa :Location {name:'United States', type:'country' }),
(idaho :Location {name:'Idaho', type:'state' }),
(lucy :Person {name:'Lucy' }),
(idaho) -[:WITHIN ]-> (usa) -[:WITHIN]-> (namerica),
(lucy) -[:BORN_IN]-> (idaho)
```
Symbolic names (`usa`, `idaho`) are **not stored** — they exist only within the query to wire up edges.
**Query — people who emigrated from the US to Europe:**
```cypher
MATCH
(person) -[:BORN_IN]-> () -[:WITHIN*0..]-> (:Location {name:'United States'}),
(person) -[:LIVES_IN]-> () -[:WITHIN*0..]-> (:Location {name:'Europe'})
RETURN person.name
```
Read as: find any vertex `person` where (1) it has an outgoing `BORN_IN` edge to a vertex from which you can follow **a chain of outgoing `WITHIN` edges** until reaching a `Location` named *United States*; and (2) the same vertex has an outgoing `LIVES_IN` edge from which a chain of `WITHIN` edges reaches a `Location` named *Europe*.
**`*0..` means "follow this edge zero or more times" — like the `*` operator in a regular expression.** This is the essential capability, and the thing SQL struggles with.
**Two possible execution strategies** (the optimizer chooses — this is declarative):
* **Forward:** scan all people, examine each person's birthplace and residence, filter.
* **Backward:** if there's an index on `name`, efficiently find the US and Europe vertices, follow **incoming** `WITHIN` edges to enumerate all locations inside each, then look for people via **incoming** `BORN_IN`/`LIVES_IN` edges at those locations.
#### 2.3 The same query in SQL — why the model matters [#23-the-same-query-in-sql--why-the-model-matters]
> **Every edge you traverse in a graph query is effectively a join with the `edges` table. In a relational database you usually know in advance which joins you need. In a graph query, you may traverse a VARIABLE number of edges — the number of joins is not fixed in advance.**
A person's `LIVES_IN` may point to a street, city, district, region, or state; a city is `WITHIN` a region, a region `WITHIN` a state, a state `WITHIN` a country. The target may be **directly adjacent or several levels away.**
SQL's tool for this is the **recursive common table expression** (`WITH RECURSIVE`):
```sql
WITH RECURSIVE
-- in_usa: vertex IDs of all locations within the United States
in_usa(vertex_id) AS (
SELECT vertex_id FROM vertices
WHERE label = 'Location' AND properties->>'name' = 'United States'
UNION
SELECT edges.tail_vertex FROM edges
JOIN in_usa ON edges.head_vertex = in_usa.vertex_id
WHERE edges.label = 'within'
),
-- in_europe: same, starting from Europe
in_europe(vertex_id) AS (
SELECT vertex_id FROM vertices
WHERE label = 'location' AND properties->>'name' = 'Europe'
UNION
SELECT edges.tail_vertex FROM edges
JOIN in_europe ON edges.head_vertex = in_europe.vertex_id
WHERE edges.label = 'within'
),
-- born_in_usa: people born somewhere within the US
born_in_usa(vertex_id) AS (
SELECT edges.tail_vertex FROM edges
JOIN in_usa ON edges.head_vertex = in_usa.vertex_id
WHERE edges.label = 'born_in'
),
-- lives_in_europe: people living somewhere within Europe
lives_in_europe(vertex_id) AS (
SELECT edges.tail_vertex FROM edges
JOIN in_europe ON edges.head_vertex = in_europe.vertex_id
WHERE edges.label = 'lives_in'
)
SELECT vertices.properties->>'name'
FROM vertices
JOIN born_in_usa ON vertices.vertex_id = born_in_usa.vertex_id
JOIN lives_in_europe ON vertices.vertex_id = lives_in_europe.vertex_id;
```
> **A 4-line Cypher query requires 31 lines in SQL. That's how much difference the right choice of data model and query language makes.**
And that's just the beginning — there are further details around **handling cycles** and **choosing breadth-first vs depth-first traversal**.
Other options: Oracle's **hierarchical** SQL extension, TigerGraph's **GSQL**, **PGQL**. The **ISO GQL standard (2024), based on Cypher**, is published but not yet widely adopted — hopefully leading to greater uniformity.
#### 2.4 Triple stores and RDF [#24-triple-stores-and-rdf]
**Mostly equivalent to the property graph model, using different words for the same ideas.** Worth knowing because the tools and languages are valuable additions to your toolbox.
**All information is stored as three-part statements: `(subject, predicate, object)`.** In `(Jim, likes, bananas)`: Jim = subject, likes = predicate (verb), bananas = object.
*Precision note: real systems store extra metadata. **AWS Neptune uses quads** (adds a graph ID); **Datomic uses 5-tuples** (adds a transaction ID and a Boolean indicating deletion). They keep the subject-predicate-object core, so the book still calls them triple stores.*
**The subject is a vertex. The object is one of two things:**
| Object is… | Meaning | Example |
| -------------------------------------- | ------------------------------------------------------------------------ | -------------------------------------------------------------------- |
| **A primitive value** (string, number) | predicate + object = **key + value of a property** on the subject vertex | `(lucy, birthYear, 1989)` ≡ vertex `lucy` with `{"birthYear": 1989}` |
| **Another vertex** | predicate = **edge label**; subject = tail vertex; object = head vertex | `(lucy, marriedTo, alain)` |
**Turtle** (a subset of Notation3) is the readable encoding:
```turtle
@prefix : .
_:lucy a :Person; :name "Lucy"; :bornIn _:idaho.
_:idaho a :Location; :name "Idaho"; :type "state"; :within _:usa.
_:usa a :Location; :name "United States"; :type "country"; :within _:namerica.
_:namerica a :Location; :name "North America"; :type "continent".
```
`_:someName` names a vertex; **the name means nothing outside the file** — it exists only so we know which triples refer to the same vertex. Semicolons let you say multiple things about one subject.
**The Semantic Web.** Triple stores were motivated by the early-2000s effort to publish data in standardized machine-readable form for internet-wide exchange. **The Semantic Web as originally envisioned did not succeed**, but its legacy lives on in: **JSON-LD**, biomedical **ontologies**, **Facebook's Open Graph protocol** (used for link unfurling), knowledge graphs like **Wikidata**, and **Schema.org** vocabularies. **Even with no interest in the Semantic Web, triples can be a good internal data model for applications.**
**RDF** is the underlying data model (Turtle is one encoding; RDF/XML is another, more verbose one; Apache Jena converts between them).
**RDF's quirk:** because it's designed for **internet-wide data exchange**, subject/predicate/object are often **URIs** — `` rather than plain `WITHIN`. **The reasoning: you should be able to combine your data with someone else's, and if they attach a different meaning to `within`, you won't get a conflict, because their predicate is actually ``.** The URL **need not resolve to anything** — it's just a namespace. Declare the prefix once at the top and forget it.
#### 2.5 SPARQL [#25-sparql]
Query language for RDF triple stores. *(Recursive acronym: SPARQL Protocol and RDF Query Language; pronounced "sparkle.")* **It predates Cypher — and Cypher's pattern matching is borrowed from SPARQL**, which is why they look similar.
```sparql
PREFIX :
SELECT ?personName WHERE {
?person :name ?personName.
?person :bornIn / :within* / :name "United States".
?person :livesIn / :within* / :name "Europe".
}
```
Equivalences (SPARQL variables start with `?`):
> **Because RDF doesn't distinguish between properties and edges — it just uses predicates for both — you can use the SAME syntax for matching properties and for traversing edges.** That's a genuine elegance advantage over property graphs.
Supported by Amazon Neptune, AllegroGraph, Blazegraph, OpenLink Virtuoso, Apache Jena.
#### 2.6 Datalog [#26-datalog]
**Much older than SPARQL or Cypher** — from 1980s academic research. **Less well known among software engineers and not widely supported in mainstream databases, but it ought to be better known: very expressive, especially powerful for complex queries.** Used by **Datomic, LogicBlox, CozoDB, and LinkedIn's LIquid.** It's based on a **relational** data model, not a graph — but **recursive queries on graphs are a particular strength.**
**Contents are `facts`, each corresponding to a row in a relational table.** `location(2, "United States", "country")` means the `location` table has a row with those column values.
```prolog
location(1, "North America", "continent").
location(2, "United States", "country").
location(3, "Idaho", "state").
within(2, 1). /* US is in North America */
within(3, 2). /* Idaho is in the US */
person(100, "Lucy").
born_in(100, 3). /* Lucy was born in Idaho */
```
Edges (`within`, `born_in`, `lives_in`) are **two-column join tables**.
```prolog
within_recursive(LocID, PlaceName) :- location(LocID, PlaceName, _). /* Rule 1 */
within_recursive(LocID, PlaceName) :- within(LocID, ViaID), /* Rule 2 */
within_recursive(ViaID, PlaceName).
migrated(PName, BornIn, LivingIn) :- person(PersonID, PName), /* Rule 3 */
born_in(PersonID, BornID),
within_recursive(BornID, BornIn),
lives_in(PersonID, LivingID),
within_recursive(LivingID, LivingIn).
us_to_europe(Person) :- migrated(Person, "United States", "Europe"). /* Rule 4 */
```
> **Cypher and SPARQL jump in right away with SELECT; Datalog takes a small step at a time.**
**How it works:** rules **derive new virtual tables** from underlying facts. These derived tables are **like virtual SQL views** — not stored, but queryable like stored tables. The name and columns come from the part **before `:-`**; the content from the pattern-matching **after `:-`**. **A rule applies if the system can find a match for all patterns on the right-hand side; when it applies, it's as though the left-hand side was added to the database** (variables replaced by matched values).
Trace of the recursion:
By repeated application of rules 1 and 2, `within_recursive` yields **all locations contained in any other location.** Rule 3 finds people with a birthplace and residence; rule 4 pins those to *United States* and *Europe*.
> **Datalog requires a different kind of thinking: complex queries are built up rule by rule, with one rule referring to others — like breaking code into functions that call each other. And just as functions can be recursive, Datalog rules can invoke themselves (rule 2), which is what enables graph traversals.**
#### 2.7 GraphQL — deliberately the *least* powerful [#27-graphql--deliberately-the-least-powerful]
**By design much more restrictive than the others.** Intended for **OLTP queries**; its purpose is to let **client software on a user's device** (mobile app, JS frontend) **request a JSON document with a particular structure containing exactly the fields needed to render its UI.**
**The benefit:** developers can **rapidly change queries in client code without changing server-side APIs.**
**The costs:**
* Organizations adopting GraphQL **often need tooling to convert queries into requests to internal services**, which commonly use REST or gRPC (Ch 5)
* **Authorization, rate limiting, and performance** are additional concerns
> **The language is intentionally limited BECAUSE GraphQL queries come from untrusted sources.** It does not allow anything expensive to execute, since otherwise users could (perhaps unintentionally) cause a **denial-of-service** by running lots of expensive queries. Specifically:
>
> * **No recursive queries** (unlike Cypher, SPARQL, SQL, Datalog)
> * **No arbitrary search conditions** — you can't ask "find people born in the US now living in Europe" unless the service owners **explicitly choose to offer that search functionality**
Example — a Slack/Discord-style chat app:
```graphql
query ChatApp {
channels {
name
recentMessages(latest: 50) {
timestamp
content
sender { fullName imageUrl }
replyTo { content sender { fullName } }
}
}
}
```
**The response mirrors the query structure exactly — those attributes, no more and no less.**
```json
{ "data": { "channels": [ { "name": "#general", "recentMessages": [
{ "timestamp": 1693143014, "content": "Hey! How are y'all doing?",
"sender": {"fullName": "Aaliyah", "imageUrl": "https://..."}, "replyTo": null },
{ "timestamp": 1693143024, "content": "Great! And you?",
"sender": {"fullName": "Caleb", "imageUrl": "https://..."},
"replyTo": { "content": "Hey! How are y'all doing?",
"sender": {"fullName": "Aaliyah"} } } ] } ] } }
```
> **The advantage: the server does not need to know which attributes the client requires to render its UI — the client simply requests what it needs.** If the UI changes to show the replyTo sender's profile picture, the client adds `imageUrl` to the query **with no server-side changes.**
**Two deliberate duplication choices, both justified the same way:**
* The sender's name and image are **embedded in each message**, so if one user sends multiple messages, the info repeats. **In principle this could be reduced, but GraphQL accepts a larger response to make it simpler to render the UI.**
* `replyTo` **duplicates the replied-to content and sender name.** Returning just an ID would force **an additional client→server request if that ID isn't among the 50 messages returned.** **Duplicating makes it much simpler to work with the data.**
**Server-side:** the database can store data **more normalized** and perform the joins to answer the query (store a message with the sender's user ID and the replied-to message ID, then resolve). **But only joins explicitly declared in the GraphQL schema can be requested by the client** — that's the DoS guardrail.
> **Despite the name and the JSON-shaped response, GraphQL can be implemented on top of ANY type of database — relational, document, or graph.**
***
# 3. Data Models and Query Languages (/docs/ddia/data-models-query-languages)
> "The limits of my language mean the limits of my world." — Ludwig Wittgenstein
**Why data models matter more than they look:** they have a profound effect not only on **how the software is written**, but on **how we think about the problem we are solving.**
***
# 3.0 The layer stack (/docs/ddia/data-models-query-languages/layer-stack)
Most applications are built by layering one data model on another. For each layer the key question is: **how is it represented in terms of the next-lower layer?**
**Each layer hides the complexity of the layers below by providing a clean data model.** That's what lets database vendors' engineers and application developers work together effectively without knowing each other's internals.
#### Declarative query languages — the key terminology note [#declarative-query-languages--the-key-terminology-note]
SQL, Cypher, SPARQL, and Datalog are **declarative**: you specify **the pattern of the data you want** — what conditions results must meet, how they should be transformed (sorted, grouped, aggregated) — **but not how to achieve it.** The **query optimizer** decides which indexes and join algorithms to use, and in which order.
With imperative languages (Python, Java) you write the algorithm: which operations, in which order.
**Why declarative wins:**
1. More concise and easier to write than an explicit algorithm.
2. **More importantly, it hides implementation details of the query engine**, so the database can introduce performance improvements **without any changes to your queries.**
3. It enables **automatic parallelism** — the DB can execute the query across multiple CPU cores and machines without you implementing that. In a handcoded algorithm, implementing parallel execution yourself is a lot of work.
***
# 3.6 Production failure catalog for this chapter (/docs/ddia/data-models-query-languages/production-failure-catalog-chapter)
| Symptom | Underlying modeling decision |
| ----------------------------------------------------------- | ------------------------------------------------------------------------------- |
| Page makes 200 DB queries to render 50 rows | **N+1**, from an ORM or a GraphQL resolver without batching |
| Rename a company → 400,000 stale copies of the old logo URL | Denormalized human-meaningful data with no update process |
| Two sources of the same relationship disagree | Many-to-many stored on **both** sides |
| Timeline shows a stale like count / old avatar | Fast-changing data was denormalized into a materialized view |
| One document is 14 MB and every read pulls all of it | Embedded a genuine one-to-**many** where the model only supports one-to-**few** |
| Migration locks the table and takes the site down | Schema-on-write `UPDATE` rewriting every row |
| Read code full of `if (!user.first_name)` forever | Schema-on-read with no migration and no shape observability |
| Recursive SQL query never terminates | Cycles in the graph; no visited-set, no depth bound |
| One graph query melts the cluster | **Supernode** with millions of edges |
| Public API DoS'd by one clever query | GraphQL without depth/complexity limits |
| Rebuilt projection produces different numbers | **Non-deterministic** event processing (external lookup, `now()`) |
| Rebuild sends 200,000 confirmation emails | Side effects inside a replayable projection |
| Cannot honour a GDPR erasure request | Immutable multi-user event log; no crypto-shredding designed in |
| Analytics dashboard is wrong after a source rename | Schema drift with schema-on-read semantics in the pipeline |
| Model works in training, garbage in production | One-hot encoder not persisted — training/serving skew |
***
# 3.1 Relational vs Document Models (/docs/ddia/data-models-query-languages/relational-vs-document-models)
#### 1.1 History — and why it matters that the challengers all failed [#11-history--and-why-it-matters-that-the-challengers-all-failed]
The relational model: proposed by **Edgar Codd in 1970**. Data organized into **relations (tables)**, each an **unordered collection of tuples (rows)**.
It was **originally a theoretical proposal, and many people doubted it could be implemented efficiently.** By the mid-1980s RDBMSs and SQL had won for regularly structured data.
> **Each competitor generated a lot of hype in its time, but none lasted.** Instead, SQL grew to incorporate other types of data.
**NoSQL** was never a single technology — it was a loose set of ideas around **new data models, schema flexibility, scalability, and open source licensing.** **NewSQL** aimed at NoSQL scalability + relational data model and transactional guarantees. Both were very influential in design; **as the principles became widely adopted, use of the terms faded.**
**The lasting effect of NoSQL: the document model** (usually JSON), popularized by **MongoDB and Couchbase** — although **most relational databases have now added JSON support** too.
#### 1.2 The object-relational impedance mismatch [#12-the-object-relational-impedance-mismatch]
If data is in relational tables and application code is object-oriented, **an awkward translation layer is needed** between objects and tables/rows/columns.
*(The term is borrowed from electronics: every circuit has an impedance on its inputs and outputs; power transfer is maximized when output and input impedances match. A mismatch causes signal reflections and other troubles.)*
**ORM frameworks** (ActiveRecord, Hibernate) reduce boilerplate but are widely criticized:
| ORM problem | Detail |
| --------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Leaky abstraction | ORMs are complex and **can't completely hide the differences**, so developers still think about both representations |
| OLTP-only | Data engineers exposing data for analytics work with the **underlying relational representation**, so the relational schema design still matters |
| Limited reach | Many ORMs work **only with relational OLTP databases** — poor support for search engines, graph databases, NoSQL |
| Generated schemas | Auto-generated relational schemas may be **awkward for direct users and inefficient on the DB**; customizing generation is complex and negates the benefit |
| **N+1 query problem** | Display N comments each with an author name → ORM issues **one query per comment** to look up the author = **N+1 queries**, instead of one join. Avoiding it requires explicitly telling the ORM to eager-fetch |
**ORM advantages** (the book is even-handed):
* For data well-suited to relational, **some translation is inevitable** — ORMs reduce the boilerplate for simple and repetitive cases (complex queries can still be handwritten)
* Some help with **caching query results**, reducing DB load
* Some help with **schema migrations** and other administrative activities
#### 1.3 The document model for one-to-many [#13-the-document-model-for-one-to-many]
The LinkedIn-résumé example. `first_name`/`last_name` appear once per user → columns. But **positions, education, and contact info are one-to-many** → separate tables with foreign keys to `users`… or one JSON document:
```json
{
"user_id": 251, "first_name": "Barack", "last_name": "Obama",
"headline": "Former President of the United States of America",
"region_id": "us:91",
"positions": [
{"job_title": "President", "organization": "United States of America"},
{"job_title": "US Senator (D-IL)","organization": "United States Senate"}
],
"education": [
{"school_name": "Harvard University", "start": 1988, "end": 1991},
{"school_name": "Columbia University", "start": 1981, "end": 1983}
],
"contact_info": {"website": "https://barackobama.com", "x": "https://x.com/barackobama"}
}
```
**Advantages claimed for the document form:**
* Reduced impedance mismatch with application code
* Lack of schema (discussed below)
* **Better locality**: fetching the relational profile needs **multiple queries or a messy multiway join**; in JSON all the relevant information is **in one place, making the query both faster and simpler**
* The **tree structure is made explicit**
**The important caveat:** a one-to-many relationship is sometimes called **one-to-few** — a résumé has a small number of positions. **If you have a genuinely large number of related items** — e.g. thousands of comments on a celebrity's post — **embedding them all in the same document may be too unwieldy, so the relational approach is preferable.**
#### 1.4 Normalization, denormalization, and joins [#14-normalization-denormalization-and-joins]
Why is `region_id` an ID rather than the string `"Washington, DC, United States"`? If the UI has a free-text field, a string is fine. But a **standardized list with a drop-down** gives:
* **Consistent style and spelling** across profiles
* **Disambiguation** — "Washington" alone: the DC or the state?
* **Ease of updating** — the name is stored once, so a city rename propagates everywhere
* **Localization** — the standardized list can be translated, so the region displays in the viewer's language
* **Better search** — the region list can encode that Washington is on the US East Coast, which the raw string doesn't reveal
**Definition:** storing an **ID** = more **normalized** (human-meaningful information stored in exactly one place; everything referring to it uses an ID that has meaning only within the database). Storing the **text** = **denormalized** (duplicating human-meaningful information in every record).
> **The core argument for IDs: because an ID has no meaning to humans, it never needs to change.** The ID stays the same even when the information it identifies changes. **Anything meaningful to humans may need to change** — and if it's duplicated, all redundant copies must be updated, requiring more code, more writes, more disk space, and **risking inconsistency** when some copies are updated and others aren't.
**Downside of normalization:** every display of a record containing an ID needs an extra lookup — a **join**.
```sql
SELECT users.*, regions.region_name
FROM users JOIN regions ON users.region_id = regions.id
WHERE users.id = 251;
```
**Document databases can store both normalized and denormalized data**, but they're associated with denormalization because (a) JSON makes it easy to add denormalized fields, and (b) **weak join support in many document databases makes normalization inconvenient.** Some don't support joins at all → you join in application code (fetch a document with an ID, then a second query to resolve it). MongoDB offers `$lookup` in an aggregation pipeline:
```js
db.users.aggregate([
{ $match: { _id: 251 } },
{ $lookup: { from: "regions", localField: "region_id",
foreignField: "_id", as: "region" } }
])
```
**The trade-off, stated crisply:**
| | Write cost | Read cost |
| ---------------- | ----------------------------------------------------- | --------------------------- |
| **Normalized** | **Faster** (one copy) | **Slower** (requires joins) |
| **Denormalized** | **More expensive** (more copies to update, more disk) | **Faster** (fewer joins) |
> **View denormalization as a form of derived data** — you need a *process* for updating the redundant copies.
And beyond the cost of the updates: **what about consistency if a process crashes halfway through?** Databases with **atomic transactions** make consistency easier, **but not all databases offer atomicity across multiple documents.** Consistency can also be maintained via **stream processing** (Ch 12).
**Where each fits:**
* **Normalization → better for OLTP**, where both reads and updates must be fast
* **Denormalization → often better for analytics**, where updates are bulk and read-only query performance dominates
* **Small-to-moderate scale → normalized is often best**: no multi-copy consistency worries, and join cost is acceptable
* **Very large scale → the cost of joins can become problematic**
#### 1.5 The X/Twitter timeline: a masterclass in *partial* denormalization [#15-the-xtwitter-timeline-a-masterclass-in-partial-denormalization]
Recall Ch 2's materialized timeline. It's the cache of the result of the expensive `posts ⋈ follows` join, and fan-out is how the denormalized copy is kept consistent.
**But the crucial detail: X's materialized timeline does NOT store the post text.** Each entry stores only:
* the **post ID**
* the **ID of the user who posted it**
* a little extra info to identify reposts and replies
i.e., it's the precomputed result of:
```sql
SELECT posts.id, posts.sender_id FROM posts
JOIN follows ON posts.sender_id = follows.followee_id
WHERE follows.follower_id = current_user
ORDER BY posts.timestamp DESC LIMIT 1000
```
**So reading a timeline still performs two joins**, in application code:
1. Look up post IDs → fetch actual post content **plus statistics like like/reply counts**
2. Look up sender IDs → fetch username, profile picture, other details
This is called **hydrating the IDs**.
**Why store only IDs?** Because **the data they refer to is fast-changing**:
* Like and reply counts may change **multiple times per second** on a popular post
* Users regularly change their username or profile photo
* The timeline must show the **latest** counts and picture when viewed
* **Denormalizing them would also significantly increase storage cost**
> **This example shows that having to perform joins when reading data is NOT, as sometimes claimed, an impediment to creating high-performance, scalable services.** Hydration scales easily because **it parallelizes well, and the cost doesn't depend on how many accounts you follow or how many followers you have.**
**The generalizable rule:** the most scalable approach **denormalizes some things and leaves others normalized.** Decide per-field by asking:
1. **How often does this information change?** (fast-changing → keep normalized, hydrate on read)
2. **What is the cost of reads vs writes** — *which may be dominated by outliers* (users with many follows/followers)
> **Normalization and denormalization are not inherently good or bad — they are trade-offs in read/write performance and implementation effort.**
#### 1.6 Many-to-one and many-to-many [#16-many-to-one-and-many-to-many]
| Relationship | Example in the résumé | Shape |
| ---------------------------- | ---------------------------------------------------------------------------- | ---------------------------- |
| **One-to-many / one-to-few** | one résumé → several positions; each position belongs to one résumé | tree |
| **Many-to-one** | many people live in the same region; each person lives in one region | FK reference |
| **Many-to-many** | a person worked at several organizations; an organization has many employees | **associative / join table** |
In the relational model a many-to-many is an **associative table (join table)**: each `position` row associates one user ID with one organization ID.
> **Many-to-one and many-to-many relationships do not easily fit within one self-contained JSON document; they lend themselves to a normalized representation.**
In a document model you'd reference by ID:
```json
{ "user_id": 251, "first_name": "Barack", "last_name": "Obama",
"positions": [
{"start": 2009, "end": 2017, "job_title": "President", "org_id": 513},
{"start": 2005, "end": 2008, "job_title": "US Senator (D-IL)", "org_id": 514}
] }
```
**Querying many-to-many "in both directions"** — all organizations a person worked for, *and* all people who worked at an organization. Two options:
| Option | How | Cost |
| ---------------------------------- | --------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Store IDs on both sides** | résumé lists org IDs; org document lists résumé IDs | **Denormalized — the relationship is stored twice and the two could become inconsistent** |
| **Store once + secondary indexes** | keep the relationship in one place; index both directions | Normalized. In relational: index both `user_id` and `org_id` on `positions`. In document: index the `org_id` field *inside* the `positions` array |
**Many document databases and relational databases with JSON support can create indexes on values inside a document** — this is what makes the normalized document approach viable.
#### 1.7 Star and snowflake schemas (analytics) [#17-star-and-snowflake-schemas-analytics]
Data warehouses are usually relational, with conventions optimized for business analysts: **star schema, snowflake schema, dimensional modeling, and one big table (OBT)**. ETL translates operational data into the chosen schema.
**Fact table** — each row is an **event that occurred at a particular time** (a customer's purchase of a product; or, for web analytics, a page view or click).
> **Facts are usually captured as individual events, because this allows maximum flexibility of analysis later. This means the fact table can become extremely large — a big enterprise may have many petabytes of transaction history, mostly fact tables.**
Fact table columns are either:
* **Attributes** — e.g. the price at which the product was sold and the cost from the supplier (so profit margin can be computed)
* **Foreign keys to dimension tables** — the **who, what, where, when, how, and why** of the event
**Even date and time get a dimension table**, because it lets you encode extra facts about dates (public holidays), enabling queries that differentiate holiday from non-holiday sales.
The name **star schema** comes from the visual: fact table in the middle, dimension tables around it like rays.
**Snowflake schema** = dimensions further broken into **subdimensions** (separate `brand` and `category` tables referenced by FK from `dim_product` rather than stored as strings). **More normalized than star schemas — but star schemas are often preferred because they're simpler for analysts to work with.**
**Tables are wide:** fact tables frequently have **over a hundred columns, sometimes several hundred**. Dimension tables are wide too — `dim_store` might include which services each store offers, whether it has an in-store bakery, square footage, opening date, last remodel date, and distance to the nearest highway.
**Relationship shape:** star/snowflake schemas consist **mostly of many-to-one relationships**. Other types could exist in principle but **are often denormalized to simplify queries** — e.g. a multi-item transaction is **not represented explicitly**; the fact table just has a separate row per product purchased, and those rows happen to share the same customer ID, store ID, and timestamp.
**One big table (OBT)** takes denormalization further: **drop the dimension tables entirely and fold their information into denormalized columns on the fact table** — essentially precomputing the fact↔dimension joins. **More storage, sometimes faster queries.**
> **Why aggressive denormalization is safe here:** the data is a **log of historical data that is not going to change** (except to correct errors). **The consistency and write-overhead problems of OLTP denormalization are not as pressing in analytics.**
#### 1.8 When to use which model [#18-when-to-use-which-model]
**Arguments for document:** schema flexibility · better performance due to locality · closer to the application's object model.
**Arguments for relational:** better support for joins, many-to-one, and many-to-many.
> **If your data has a document-like structure — a tree of one-to-many relationships where the entire tree is typically loaded at once — a document model is probably a good idea.** The relational technique of **shredding** (splitting a document-like structure across multiple tables) **can lead to cumbersome schemas and unnecessarily complicated application code.**
**Document model limitations:**
* **You cannot refer directly to a nested item.** You must say "the second item in the list of positions for user 251." If you need to reference nested items, **relational works better — any item is directly addressable by its ID.**
* **Conversely, ordered/reorderable lists favor documents.** A to-do list or issue tracker where users drag-and-drop to reorder: in a document, items (or their IDs) sit in a JSON array that *is* the order. **Relational databases have no standard way of representing reorderable lists** — the workarounds are sorting by an integer column (**requiring renumbering when inserting into the middle**), maintaining a **linked list of IDs**, or **fractional indexing**.
#### 1.9 Schema-on-read vs schema-on-write [#19-schema-on-read-vs-schema-on-write]
Most document databases (and JSON support in relational databases) **do not enforce any schema**. (XML support in relational databases usually *does* offer optional schema validation.) No schema = arbitrary keys and values can be added, and readers have **no guarantees about what fields exist.**
> **"Schemaless" is misleading** — the code reading the data **usually assumes some structure**. There *is* an **implicit schema**; it's just **not enforced by the database.**
| Term | Meaning | Programming-language analogy |
| ------------------- | --------------------------------------------------------------------- | --------------------------------------- |
| **Schema-on-read** | Structure is implicit, interpreted only when data is read | **Dynamic (runtime) type checking** |
| **Schema-on-write** | Schema is explicit; the DB ensures all data conforms **when written** | **Static (compile-time) type checking** |
**Just as the static/dynamic typing debate has no clear winner, neither does this one.**
**The difference is most visible when changing the data format.** Say you store a user's full name in one field and now want first and last name separately.
**Document (schema-on-read):** just start writing new documents with the new fields, and handle old ones in application code:
```js
if (user && user.name && !user.first_name) {
// Documents written before Dec 8, 2023 don't have first_name
user.first_name = user.name.split(" ")[0];
}
```
*Downside: **every part of your application that reads from the database must now handle old formats**, possibly written long ago.*
**Relational (schema-on-write):** migrate.
```sql
ALTER TABLE users ADD COLUMN first_name text DEFAULT NULL;
UPDATE users SET first_name = split_part(name, ' ', 1); -- PostgreSQL
UPDATE users SET first_name = substring_index(name, ' ', 1); -- MySQL
```
**Adding a column with a default is fast and unproblematic even on large tables.** But **the `UPDATE` is likely slow on a large table since every row is rewritten**, and other schema operations (e.g. changing a column's datatype) typically require **copying the entire table**.
Tools exist for background, no-downtime schema change, but **migrations on large databases remain operationally challenging.** And the clever escape hatch: **add the column with a `NULL` default (fast) and fill it in at read time — exactly as you would with a document database.**
**Schema-on-read is advantageous when the data is heterogeneous:**
* Many types of objects, and **it isn't practicable to put each type in its own table**
* **The structure is determined by external systems you don't control and that may change at any time**
**But when all records are expected to have the same structure, schemas are a useful mechanism for documenting and enforcing that structure.**
#### 1.10 Data locality for reads and writes [#110-data-locality-for-reads-and-writes]
A document is usually stored as **a single continuous string** — JSON, XML, or a binary variant like MongoDB's **BSON**. If the app often needs the **entire** document (to render a page), this **storage locality has a performance advantage**; splitting across tables requires **multiple index lookups**, which may mean more disk seeks and more time.
> **The locality advantage applies ONLY if you need large parts of the document at the same time.** The database typically **loads the entire document**, which is wasteful if you need only a small part of a large one. **And on updates, the entire document usually needs to be rewritten.**
>
> **Therefore: keep documents fairly small and avoid frequent small updates.**
**Locality is not exclusive to the document model:**
| System | Mechanism |
| --------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- |
| **Google Spanner** | Schema can declare that a table's rows be **interleaved (nested) within a parent table** — same locality property, relational model |
| **Oracle** | **Multi-table index cluster tables** |
| **Bigtable / HBase / Accumulo** (wide-column) | **Column families**, serving a similar locality purpose |
#### 1.11 Query languages for documents [#111-query-languages-for-documents]
Relational → SQL. Document databases vary widely: **key-value access by primary key only** / **secondary indexes into document values** / **rich query languages**.
* **XML**: XQuery and XPath — complex queries, **including joins across multiple documents**, results formatted as XML
* **JSON**: **JSON Pointer** and **JSONPath** are the XPath equivalents
* **MongoDB aggregation pipeline** — a query language for collections of JSON documents
Aggregation example — sharks sighted per month:
```sql
-- PostgreSQL
SELECT date_trunc('month', observation_timestamp) AS observation_month,
sum(num_animals) AS total_animals
FROM observations
WHERE family = 'Sharks'
GROUP BY observation_month;
```
```js
// MongoDB aggregation pipeline — same query
db.observations.aggregate([
{ $match: { family: "Sharks" } },
{ $group: {
_id: { year: { $year: "$observationTimestamp" },
month: { $month: "$observationTimestamp" } },
totalAnimals: { $sum: "$numAnimals" }
} }
]);
```
> The aggregation pipeline is **similar in expressiveness to a subset of SQL**, with a JSON-based syntax rather than SQL's English-sentence style. **The difference is perhaps a matter of taste.**
#### 1.12 Convergence [#112-convergence]
> **Document and relational databases started as very different approaches and have grown more similar over time.**
* **Relational** added JSON types and query operators, and **indexing of properties inside documents**
* **Document** (MongoDB, Couchbase, RethinkDB) added **joins, secondary indexes, and declarative query languages**
**This is good news, because the two models work best when combined in the same database.** Many document databases need relational-style references; many relational databases have sections where schema flexibility helps. **Relational–document hybrids are a powerful combination.**
**Historical footnote worth knowing:** Codd's original 1970 relational model allowed something like JSON — he called them **nonsimple domains**: a value in a row need not be a primitive; it can be a **nested relation**, giving arbitrarily nested trees. **That's comparable to the JSON/XML support added to SQL over 30 years later.**
***
# 3.10 Self-test (/docs/ddia/data-models-query-languages/self-test)
Give the four layers of the data-model stack and say who owns each.
What are the two advantages of declarative query languages beyond conciseness? Which one is more important, and why?
Why did every challenger to the relational model fail while SQL absorbed their ideas?
What is the impedance mismatch? Give three criticisms of ORMs and two genuine advantages.
Explain the N+1 problem and how you avoid it.
Why does the document model suit the résumé example? Name the specific limit expressed by "one-to-few."
Give five reasons to store a region ID rather than the region name.
State the normalization trade-off in one sentence for reads and one for writes. Why is denormalization "a form of derived data"?
What exactly does X's materialized timeline store, and what does it deliberately *not* store? Give both reasons.
What is hydration, and why does the book say join-on-read is *not* an obstacle to scalability?
Which relationship types don't fit in a single document? Give the two ways to query a many-to-many bidirectionally and the risk of each.
Draw a star schema. What is in a fact table row, and what do the dimensions represent?
Why is a fact table's multi-item transaction not represented explicitly?
What is OBT, and why is aggressive denormalization safe in analytics but not OLTP?
Contrast schema-on-read with schema-on-write using the static/dynamic typing analogy. Show the migration for splitting a name field both ways.
Why is "schemaless" misleading? When is schema-on-read genuinely better?
What is the locality advantage of documents, and what are its two limits? Name three non-document systems that provide locality.
Why does a graph query mean a variable number of joins? Why is that hard in SQL?
Describe the property graph model. Why are there indexes on *both* tail and head vertex?
Give three things about the Lucy/Alain example that are hard in a relational schema.
What does `*0..` mean, and what is its regex analogy? Give the two execution strategies an optimizer might choose.
Map triple-store concepts onto property-graph concepts. What are the two possible kinds of object?
Why does RDF use URIs as predicates?
What is the one syntactic advantage SPARQL gets from RDF not distinguishing properties from edges?
How does Datalog differ in style from Cypher and SPARQL? Trace the `within_recursive` rules for Idaho.
Why is GraphQL deliberately restricted? Name two things it forbids and the reason.
Why does GraphQL accept duplication in its responses? Give both cases.
Define event sourcing and CQRS. What is a command, when does it become an event, and what may a projection never do?
Why must events be named in the past tense?
Give two differences between an event log and a star-schema fact table.
List four advantages of event sourcing and all three downsides. Which downside is a determinism problem?
What is the one hard requirement on the event storage system, and why is it difficult in a distributed system?
What is a DataFrame, what is `merge`, and why do data scientists prefer it to SQL?
Explain one-hot encoding and why a sparse matrix doesn't fit a relational database.
you're modeling a project-management product: workspaces contain projects, projects contain tasks, tasks have ordered subtasks users drag to reorder, tasks have many assignees, and users want full-text search plus "show me everything blocking this task" (transitive dependencies of unknown depth). Choose a model (or models) for each part, justify each against a specific trade-off from this chapter, and identify the one query that would push you toward a graph database — and the one that would make you regret it.
# 3.5 Technology deep dives (/docs/ddia/data-models-query-languages/technology-deep-dives)
***
#### 5.1 The relational model / SQL databases (PostgreSQL, MySQL) [#51-the-relational-model--sql-databases-postgresql-mysql]
**Problem it solves.** Store data with regular structure so that **arbitrary queries can be answered efficiently without knowing the queries in advance**, with the database — not the application — responsible for choosing indexes, join order, and parallelism.
**Why wasn't the alternative enough?** The **hierarchical and network models** (1970s–80s) required the application to know the physical access paths; changing a query meant changing the storage layout. **Codd's insight was to separate the logical model from the physical access path** — that separation is why declarative queries and query optimizers exist at all, and it's why relational survived four waves of challengers.
**How it works internally.** Tables → rows stored in heap files or a clustered index (Ch 4); B-tree indexes for lookups; a **query planner** that estimates cardinalities from statistics and picks among nested-loop / hash / merge joins; MVCC for isolation (Ch 8); a write-ahead log for durability and replication (Ch 6).
**Deployment.** Primary + replicas; connection pooling in front (**PgBouncer** — Postgres's per-connection process model makes this near-mandatory above a few hundred clients); migrations as versioned, reviewed, forward-only scripts.
**Monitoring.** Slow query log and `pg_stat_statements` (total time by normalized query — the single highest-value view); **replication lag**; **connection count vs `max_connections`**; cache hit ratio; **transaction ID wraparound / autovacuum progress** (Postgres-specific and genuinely dangerous); table and index bloat; lock waits and deadlock rate.
**Scaling.** Read replicas for read scaling; connection pooling; partitioning by range/hash for very large tables; then sharding at the application layer (Ch 7). Vertical scaling gets you much further than folklore suggests.
**Backup.** Base backup + **WAL archiving** for point-in-time recovery (PITR). **`pg_dump` is not a backup strategy for a large production database** — it can't do PITR and restore time is measured in hours. **Test restores on a schedule**; an untested backup is a hypothesis.
**What actually breaks in production.**
* **Long-running transactions block vacuum**, bloat accumulates, and eventually the table's performance collapses — or the database refuses writes to prevent XID wraparound.
* **`ALTER TABLE` taking an ACCESS EXCLUSIVE lock** behind a queue of waiting queries: the migration itself is fast, but every query behind the lock piles up and the site goes down for the duration. Use `lock_timeout` and `CONCURRENTLY` variants.
* **Connection storms** — an app tier scaling out past `max_connections` during a traffic spike.
* **The planner flipping to a bad plan** after statistics drift, turning a 5 ms query into a 5-minute sequential scan.
* **Unbounded `IN (...)` lists** generated by ORMs.
* **Replica promoted with data loss** because replication was asynchronous (Ch 6).
***
#### 5.2 Document databases (MongoDB, Couchbase) and JSON columns in relational DBs [#52-document-databases-mongodb-couchbase-and-json-columns-in-relational-dbs]
**Problem it solves.** Store self-contained tree-shaped records with good locality, without a rigid up-front schema, in a shape close to the application's objects.
**Why wasn't relational enough?** Shredding a document-like structure across tables gives **cumbersome schemas and unnecessarily complicated application code**, plus a multiway join or N queries to reassemble one logical object.
**Why isn't it always the answer?** Many-to-one and many-to-many relationships don't fit a single document; you cannot address nested items directly; and joins are weak or absent.
**How it works internally.** A document is stored as **one contiguous encoded string** (**BSON** in MongoDB). B-tree indexes can be built on **paths inside the document**, including fields inside arrays (multikey indexes). Updates typically **rewrite the whole document**; if the new version doesn't fit in place, it's relocated, which invalidates and updates every index entry pointing at it.
**Deployment.** Replica sets (primary + secondaries with automatic election); sharded clusters add config servers and routers (`mongos`).
**Monitoring.** Document size distribution (the p99 matters far more than the mean); **working set vs RAM** — once indexes stop fitting in memory, performance falls off a cliff; index hit rate and `COLLSCAN` counts; replication oplog window (how far a secondary can fall behind before needing a full resync); lock/ticket saturation.
**Scaling.** Shard on a key with high cardinality and even distribution; **avoid monotonically increasing shard keys** (timestamps, ObjectIds), which send all writes to one shard.
**Backup.** Snapshots plus oplog replay for PITR. A `mongodump` of a sharded cluster is **not** consistent across shards unless you coordinate it.
**What actually breaks.**
* **Unbounded array growth** — the classic "embed comments in the post document" design working fine until a post gets 40,000 comments, at which point every read pulls megabytes and every write rewrites them. **This is exactly the "one-to-few, not one-to-many" warning.**
* **Implicit schema drift.** Five years of writes, six generations of shape, and read code full of `if (!user.first_name)` branches that nobody dares delete — because nobody knows whether any document still lacks the field.
* **Application-side joins with no transaction** — the two documents fetched are from different points in time.
* **Dual-write inconsistency** when the relationship is stored on both sides for bidirectional querying.
* **Monotonic shard key** → one hot shard doing all the writes while the rest idle.
***
#### 5.3 Graph databases (Neo4j, Memgraph, KùzuDB, Amazon Neptune) [#53-graph-databases-neo4j-memgraph-kùzudb-amazon-neptune]
**Problem it solves.** Queries that traverse a **variable, not-known-in-advance number of hops** across highly connected heterogeneous data.
**Why wasn't relational enough?** Every traversed edge is a join with the edges table, and **you don't know how many joins you need until you run the query.** SQL can express this with `WITH RECURSIVE` — at roughly 8× the code, with cycle handling and traversal-order control left to you.
**Why wasn't a document database enough?** Documents model trees. Graphs model arbitrary many-to-many connectivity, in both directions.
**How it works internally.** Two logical tables (vertices, edges) with indexes on **both** `tail_vertex` and `head_vertex` so traversal works forward and backward. Native graph engines go further with **index-free adjacency**: a vertex record stores direct pointers to its edge records, so a hop is a pointer dereference rather than an index lookup — this is what makes deep traversals O(hops) rather than O(hops × log n). Query execution then becomes a **pattern-matching problem**, and the optimizer chooses which end of the pattern to start from (the "forward vs backward" choice in §2.2).
**Deployment.** Usually a single primary with read replicas — **graph databases are notoriously hard to shard**, because a good partition of a highly connected graph doesn't exist (min-cut on a social graph is bad by definition). Neptune and Neo4j clusters replicate rather than partition for this reason.
**Monitoring.** Query traversal depth and expanded-node counts (the real cost driver, not row counts); page-cache hit rate; heap pressure during large traversals; supernode degree distribution.
**Scaling.** Vertically first, then read replicas. If you truly need to partition, expect cross-partition traversals to dominate cost.
**Backup.** Full + incremental snapshots; the graph is usually a system of record, so treat it accordingly.
**What actually breaks.**
* **Supernodes.** One vertex with millions of edges (a celebrity, a "USA" location node, a shared category) makes every traversal through it explode. Mitigations: edge-type partitioning, degree-aware query planning, or modeling the supernode away.
* **Unbounded variable-length patterns** — `*0..` with no upper bound over a cyclic graph, producing a query that never finishes. Always bound the depth in production.
* **Cycles** causing infinite traversal when the engine doesn't deduplicate visited vertices.
* **Sharding attempts** that turn every query into a distributed join.
* **Schemaless flexibility becoming schemaless chaos** — since *any vertex can connect to any vertex*, nothing stops a bad writer from creating relationships the query code never anticipated.
***
#### 5.4 GraphQL servers [#54-graphql-servers]
**Problem it solves.** Let clients specify exactly the fields their UI needs, so UI changes don't require server changes, and so mobile clients don't over-fetch or make N round trips.
**Why wasn't REST enough?** Fixed response shapes cause **over-fetching** (endpoints return more than a screen needs) and **under-fetching** (a screen needs 4 endpoints, so 4 round trips, or a bespoke endpoint per screen that the backend team must ship).
**How it works internally.** A schema defines types and the allowed traversals. Execution walks the query tree, calling a **resolver** per field. The naive implementation is **catastrophically N+1**: resolving `sender` for 50 messages calls the user resolver 50 times. The fix is **DataLoader-style batching and per-request caching** — collect the IDs requested within a tick, issue one batched query, distribute results.
**Deployment.** A gateway in front of REST/gRPC internal services (which is where most of the operational cost lands — **organizations adopting GraphQL often need tooling to convert queries into requests to internal services**). Persisted/allow-listed queries in production.
**Monitoring.** Per-**field** resolver latency and error rate (not per-endpoint — there is only one endpoint); query depth and complexity score distribution; **resolver call counts per request** (the N+1 detector); cache hit rate in the batch loaders.
**Scaling.** **Persisted queries** (clients send a hash, server has the allow-listed document) both bounds the query space and shrinks requests. Complexity/depth limits. Automatic Persisted Queries + CDN caching for public read traffic.
**What actually breaks.**
* **N+1 resolvers** — the default failure mode, and the reason DataLoader exists.
* **Denial of service by query complexity** — deeply nested or wide queries. The book flags this as *the* reason the language is deliberately limited; in practice you still need depth limits, complexity scoring, and timeouts, because the schema alone doesn't bound cost.
* **Authorization at the wrong layer.** Because any field can be reached by many paths, endpoint-level authz doesn't work — **authorization must be enforced per field/resolver**, and this is where GraphQL security bugs live.
* **Rate limiting is hard** — one POST to `/graphql` may be 10 or 10,000 units of work, so request-count limits are meaningless; you need **cost-based** limits.
* **Caching is hard** — one URL, POST bodies, no HTTP cache semantics for free.
* **Schema deprecation** — you cannot see who uses a field without field-level usage telemetry.
***
#### 5.5 Event sourcing platforms (EventStoreDB, MartenDB, Kafka + stream processors) [#55-event-sourcing-platforms-eventstoredb-martendb-kafka--stream-processors]
**Problem it solves.** Complex business domains where **no single representation serves all reads**, where **intent matters**, where **auditability is required**, and where you want **reversibility** rather than destructive updates.
**Why wasn't a normal database enough?** A committed `UPDATE`/`DELETE` destroys the prior state and the *reason* for the change. You cannot re-derive a new view of history you no longer have, and you cannot easily reverse a committed transaction.
**Why wasn't plain CDC enough?** CDC gives you row-level diffs (Ch 12) — the *what*, not the *why*. `active=false` doesn't tell you whether it was a cancellation, a fraud block, or a data fix.
**How it works internally.**
* Events appended to a **per-aggregate stream** with a **monotonically increasing sequence number**.
* **Optimistic concurrency on append**: "append these events expecting the stream to be at version N" — this is how command validation stays correct under concurrency, and it's the mechanism most homegrown implementations forget.
* Rebuilding state = **fold** over the stream. Because folds get slow for long streams, real systems add **snapshots** (checkpoint state at version N, replay only from there).
* **Projections** consume the log in order and write read models. Each projection tracks its own **checkpoint/offset**, which is what makes "delete the view and rebuild" possible.
**Deployment.** Event store (or Kafka topic with `cleanup.policy=compact` off — event sourcing needs **retention forever**, not compaction, unless you're keeping only latest-per-key). Projection workers as separate deployables so they can be rebuilt independently. A schema registry for event types (Ch 5), because **old events are never rewritten and must stay readable forever.**
**Monitoring.**
* **Projection lag** per view (events behind head) — the number that determines whether users see stale data
* Rebuild duration per projection — this is your recovery-time budget, and it grows with history forever
* Append conflict rate (optimistic concurrency retries) — rising rate means an aggregate is a contention hotspot
* Stream length distribution — a stream growing without bound signals a modeling error
* **Dead-letter count for events a projection failed on** — and remember: *a projection is not allowed to reject an event*, so any dead letter is a bug, not a business case
**Scaling.** Partition by aggregate ID (order matters *within* an aggregate, rarely globally). Snapshot long streams. Run projections independently and in parallel — they don't need to keep pace with each other.
**Backup.** The log **is** the backup: views are disposable and rebuildable. So backup discipline concentrates entirely on the event store — and the thing to test is **rebuild time**, because that's your real RTO.
**What actually breaks.**
* **Non-deterministic projections** — the currency-conversion trap, but also `now()`, random IDs, and calls to external services inside a projection. Rebuild produces different numbers than the original run and nobody can explain why.
* **Rebuild time growing past the maintenance window.** Fine at 10M events, a weekend outage at 10B.
* **Side effects on replay** — resending confirmation emails during a rebuild. Requires a strict separation between projections (pure) and reactors/process managers (effectful, with their own idempotency).
* **GDPR erasure vs immutability** — crypto-shredding works, but **destroying the key makes those events unreadable forever, so any future rebuild silently loses them.**
* **Event schema evolution.** You will read 2019 events in 2027. Versioning and upcasting are mandatory, not optional (Ch 5).
* **Modeling aggregates too large** — one "Conference" stream with a million events means every command replays a million events, or depends entirely on snapshots being healthy.
* **Ordering assumptions across streams.** Order is guaranteed *within* a stream. Cross-aggregate ordering in a distributed system is exactly the hard problem of Ch 10.
***
#### 5.6 DataFrame / array systems (Pandas, Spark, TileDB, NumPy) [#56-dataframe--array-systems-pandas-spark-tiledb-numpy]
**Problem it solves.** Bridge relational data and the **numeric matrices ML algorithms require**, with an interactive, incremental "wrangling" workflow.
**Why wasn't SQL enough?** Feature engineering needs custom code; pivoting into a **thousands-of-columns sparse matrix** doesn't fit relational storage; and the workflow is exploratory and imperative rather than declarative.
**How it works internally.** Columnar in-memory representation (Pandas → NumPy arrays; increasingly **Apache Arrow** as the shared memory format, which is what lets Pandas/Polars/DuckDB/Spark exchange data with zero copies). Operations are vectorized over whole columns. **Spark DataFrames are lazy** — the chain of transformations builds a logical plan that Catalyst optimizes before any execution, which is why Spark can do predicate pushdown that Pandas cannot.
**Deployment.** Local single-machine (Pandas, Polars, DuckDB) is right far more often than people assume — see Ch 1's *more nodes are not always faster*. Distributed (Spark, Dask) when the data genuinely exceeds one machine.
**Monitoring.** Peak memory vs available (the dominant failure mode); shuffle volume and skew in Spark; task duration skew across partitions; spill-to-disk volume.
**Scaling.** Prefer a bigger machine before a cluster. Then partition, and above all **avoid wide shuffles** — a `groupBy` on a skewed key is the standard Spark cliff.
**What actually breaks.**
* **OOM on a `pivot`** that materializes a dense matrix from sparse data — the exact transformation §4 describes, done without a sparse representation.
* **Silent dtype coercion** — a column becomes `object` because one row had a string, and everything downstream slows by 100× or computes wrong.
* **Chained-assignment / `SettingWithCopyWarning`** in Pandas, where an update lands on a copy and is silently discarded.
* **Data skew in Spark** — 199 tasks finish in 10 s and one runs for 4 hours, because one key holds 60% of the rows.
* **Training/serving skew** — the one-hot encoding built at training time has different category ordering than at serving time, so the model gets garbage features. **The encoder must be persisted with the model, not recomputed.**
* **Notebook irreproducibility** — cells executed out of order, so the DataFrame in memory doesn't correspond to any sequence of code anyone can rerun.
***
# 3.8 Terminology introduced here (/docs/ddia/data-models-query-languages/terminology-introduced-here)
# 3.9 Worked examples (/docs/ddia/data-models-query-languages/worked-examples)
**① The N+1 problem, costed.** A page shows 50 comments, each needing its author's name.
* **N+1:** 1 query for comments + 50 author lookups = **51 round trips.** At 0.5 ms each = **25.5 ms** of pure network time, before any work.
* **One join:** **1 round trip = 0.5 ms.**
* Now put the database in another AZ (RTT 1.5 ms): **76.5 ms vs 1.5 ms — 51×.**
**The bug is invisible in development against localhost and catastrophic in production.** That's why the chapter says these are "subtle bugs difficult to find by testing."
**② Normalize or denormalize? Run the numbers per field.** A social feed, 5,000 reads/s, 100 writes/s.
* **Post text** (changes \~never): denormalizing costs 100 writes/s × fan-out; saves 5,000 lookups/s. **Denormalize.**
* **Like count** (changes \~10×/s per hot post): denormalizing into 1M timelines = **10M writes/s.** Absurd. **Normalize, hydrate on read.**
* **Username/avatar** (changes rarely, but *appears* in millions of timelines): a single username change would rewrite millions of rows. **Normalize, hydrate on read.**
**Exactly the X design — and note the decision differs per field within the same record.**
**③ Why hydration scales and the join didn't.** Hydrating 1,000 post IDs + 1,000 sender IDs:
* **Parallelizable:** 2 batched multi-gets, not 2,000 round trips
* **Cost independent of graph shape:** a user following 10 people and one following 10,000 both hydrate \~1,000 timeline entries
Contrast the original `posts ⋈ follows` join, whose cost is **linear in followee count** — 200 lookups for a normal user, 10,000 for a power user. **The join's cost was unbounded; hydration's is bounded by page size.**
**④ Document locality, measured.** A 200 KB user document; a page needs 2 KB of it.
* **Read:** the database loads **all 200 KB** — **100× the needed bytes.**
* **Write:** updating one field rewrites **all 200 KB.** At 100 updates/s that's **20 MB/s of write amplification** for 200 B of actual change.
**"Keep documents fairly small and avoid frequent small updates"** is this arithmetic.
**⑤ Recursive traversal: Cypher vs SQL, and why depth matters.** Finding all locations within the US.
* **Cypher:** `-[:WITHIN*0..]->` — 4 lines total.
* **SQL:** a recursive CTE, \~31 lines, and **you must handle cycles and choose traversal order yourself.**
* **Fixed 2-hop?** A plain SQL join is fine and cheaper. **The variable-depth requirement is what buys the graph database its complexity.**
**⑥ Star schema I/O.** `fact_sales`: 1B rows × 100 columns × 8 bytes = **800 GB**. Query touches 3 columns.
* **Row store:** reads \~800 GB.
* **Column store:** 1e9 × 3 × 8 = **24 GB**; with 4:1 compression, **\~6 GB**.
* **OBT (dimensions folded in):** the join disappears, but the fact table grows — say 140 columns. **Query still reads only its 3 columns**, so OBT costs storage, not scan time. That's the trade.
**⑦ Event sourcing rebuild time.** 500M events; a projection processes 50,000 events/s.
Rebuild = 500e6 / 50e3 = **10,000 s ≈ 2.8 hours.** In three years at the same rate: **\~8.5 hours.** **Your rebuild time is your recovery objective, and it grows monotonically forever.** This is why snapshots exist, and why "we can always replay" needs a number attached.
***
# 2.1 Case study: social network home timelines (/docs/ddia/defining-nonfunctional-requirements/case-study-social-network)
The numbers (X/Twitter-shaped, simplified):
| Quantity | Value |
| ------------------------------------- | -------------------------------- |
| Posts per day | 500 million |
| Posts per second (average) | **5,800** |
| Posts per second (spike) | **150,000** |
| Average follows / followers per user | **200 / 200** |
| Range of followers | a handful → 100M+ (Barack Obama) |
| Simultaneously online users (assumed) | 10 million |
| Timeline freshness target | **5 seconds** |
#### 1.1 The naive design: query on read [#11-the-naive-design-query-on-read]
Relational schema: `users`, `posts`, `follows`.
```sql
SELECT posts.*, users.* FROM posts
JOIN follows ON posts.sender_id = follows.followee_id
JOIN users ON posts.sender_id = users.id
WHERE follows.follower_id = current_user
ORDER BY posts.timestamp DESC
LIMIT 1000
```
Execution: use `follows` to find everyone `current_user` follows → look up recent posts by those users → sort by timestamp → take the most recent 1,000.
**Why it doesn't work — do the arithmetic:**
**400 million lookups per second — and that's the *average* case.** Some users follow tens of thousands of accounts, and for them this query is very expensive and hard to make fast.
Two separate problems identified:
* **Polling** wastes work re-asking a question whose answer usually hasn't changed
* **Query-on-read** re-does the same expensive merge for every request
#### 1.2 The fix: push + materialize [#12-the-fix-push--materialize]
1. **Push instead of poll** — the server actively pushes new posts to followers who are online.
2. **Precompute the query result** — store, per user, a data structure containing their home timeline. On every post, look up all the poster's followers and **insert that post into each follower's timeline — like delivering a message to a mailbox.**
On login, hand over the precomputed timeline. For notifications, the client just **subscribes to the stream of posts being added to their timeline.**
**Fan-out** = the factor by which one initial request multiplies into downstream requests.
**The new arithmetic:**
```txt
5,800 posts/s × 200 followers = ~1,160,000 timeline writes/second
```
**\~1.16 million writes/s vs 400 million lookups/s — a \~350× saving.**
And during a spike, timeline deliveries **don't have to be immediate** — enqueue them, accept that posts temporarily take a bit longer to appear. **Timelines stay fast to load throughout the spike, because reads are served from cache.**
This is **materialization**; the timeline cache is a **materialized view**. The trade: *speeds up reads, costs more work on writes.*
#### 1.3 The two extreme cases (this is the real lesson) [#13-the-two-extreme-cases-this-is-the-real-lesson]
| Extreme | Problem | Solution |
| --------------------------------------------------- | --------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| **User follows a huge number of prolific accounts** | Very high write rate to their materialized timeline | That user isn't reading all of it anyway → **it's OK to drop some of their timeline writes** and show only a sample |
| **Celebrity with millions of followers posts** | Must insert into millions of timelines | **Dropping is NOT OK.** Handle celebrity posts separately: **store them separately and merge at read time** with the materialized timeline |
This **hybrid fan-out** (write-path for the many, read-path for the few) is the canonical answer to skew, and it recurs in Ch 7 (hot-spot handling). Even with it, "handling celebrities on a social network can require a lot of infrastructure."
***
# 2.8 Decision cheat sheet (/docs/ddia/defining-nonfunctional-requirements/decision-cheat-sheet)
**Which latency number do I put in the SLO?**
p50 for "typical user experience," **p99 for the SLO**, p999 only if your slowest requests correlate with your most valuable customers (the Amazon case). Skip p9999 — too expensive, too noisy, diminishing returns.
**Where do I measure?**
**Client side.** Server-side timers exclude queueing delay, which is where the variance lives. Instrument both and treat the gap as a signal in itself.
**Read-heavy feed: query on read, or materialize on write?**
Materialize when read volume × query cost ≫ write volume × fan-out. Go **hybrid** the moment the fan-out distribution has a long tail — write-path for the many, read-path merge for the few.
**How do I stop a retry storm?**
Jittered exponential backoff + a **retry budget** (token bucket capping retries as a fraction of traffic) + circuit breakers + server-side load shedding with **bounded** queues. Any one alone is insufficient.
**Vertical or horizontal?**
Vertical while it's cheap and simple — it usually is, for longer than people assume. Horizontal when you need multi-DC fault tolerance, elasticity, or you've hit the price cliff. Remember horizontal costs you explicit sharding plus all of Ch 9.
**How far ahead should I design for scale?**
**One order of magnitude.** Not two. Expect to rethink the architecture at each 10×.
**When is more automation wrong?**
When load is predictable (manual scaling has fewer surprises), and when the automation makes failures harder to troubleshoot than a manual procedure would be.
***
# 2.2 Describing Performance (/docs/ddia/defining-nonfunctional-requirements/describing-performance)
Two metric families:
* **Response time** — elapsed time from the user making a request to receiving the answer. Unit: seconds/ms/µs.
* **Throughput** — requests per second, or data volume per second, being processed. For a given hardware allocation there is a **maximum** throughput. Unit: "somethings per second."
In the case study: *posts/s* and *timeline writes/s* are throughput; *time to load the home timeline* and *time until a post reaches followers* are response times.
#### 2.1 The relationship: queueing [#21-the-relationship-queueing]
**Why:** when a request arrives at a highly loaded system, the CPU is likely already handling an earlier request, so the new one **waits**. As throughput approaches hardware maximum, queueing delays increase **sharply**.
> **Which metric matters to whom:** response time is what *users* care about most. Throughput determines the *computing resources required* (how many servers) and therefore the **cost** of serving the workload. A system is **scalable** if its maximum throughput can be significantly increased by adding computing resources.
#### 2.2 Metastable failure — when an overloaded system won't recover [#22-metastable-failure--when-an-overloaded-system-wont-recover]
The vicious cycle:
**Key property: even when the original load is removed, the system may stay overloaded until it is rebooted or reset.** That's what makes it *metastable* — the failure state is self-sustaining. This causes serious production outages.
**Countermeasures — memorize this table, it's the most operationally reusable content in the chapter:**
| Side | Technique | What it does |
| ------ | ------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------ |
| Client | **Exponential backoff** | increase *and randomize* the time between successive retries (randomization = jitter, to break synchronized retry waves) |
| Client | **Circuit breaker** | temporarily stop sending requests to a service that recently errored or timed out |
| Client | **Token bucket** | rate-limit retries to a fixed budget |
| Server | **Load shedding** | detect approaching overload and *proactively reject* requests |
| Server | **Backpressure** | send responses asking clients to slow down |
| Both | Queueing & load-balancing algorithm choice | can materially change behavior under saturation |
#### 2.3 Latency vs response time — precise definitions [#23-latency-vs-response-time--precise-definitions]
| Term | Meaning |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Response time** | What the **client** sees. Includes **all** delays incurred anywhere in the system. |
| **Service time** | Duration the service is **actively processing** the request. |
| **Queueing delay** | Waiting, at several possible points: for a CPU to become available; for an outbound network buffer when other tasks on the machine are sending a lot of data. |
| **Latency** | Catchall for time when the request is **not being actively processed** — i.e., it is *latent*. |
| **Network latency / delay** | Time the request and response spend traveling through the network. |
**Sources of random per-request delay:** a context switch to a background process; a lost network packet and TCP retransmission; a **garbage collection pause**; a **page fault** forcing a disk read; **mechanical vibrations in the server rack**.
**Head-of-line blocking.** A server processes only a small number of things in parallel (bounded by CPU cores). It takes **only a small number of slow requests** to hold up all subsequent ones. Those later requests may have *fast service times* but the client still sees a slow response.
> **Therefore: queueing delay is not part of service time, and this is exactly why you must measure response times on the CLIENT side.** Server-side timers systematically under-report the thing users actually experience.
#### 2.4 Average, median, percentiles [#24-average-median-percentiles]
Response time is **a distribution, not a number**. Variation in network delay is called **jitter**.
* **Mean (arithmetic average)** — useful for *estimating throughput limits*. Bad for "typical" response time: it doesn't tell you **how many users actually experienced that delay**.
* **Median = p50** — sort fastest→slowest, take the halfway point. p50 = 200 ms means half of requests are faster, half slower. Good metric for "how long users typically wait."
* **p95 / p99 / p999** — how bad the outliers are. p95 = 1.5 s means 5 out of 100 requests take ≥ 1.5 s.
**Tail latencies** (high percentiles) directly affect user experience.
**The Amazon argument (important and non-obvious):** Amazon specifies internal service response times at the **99.9th percentile**, even though it affects only 1 in 1,000 requests — because **the customers with the slowest requests are usually those with the most data on their accounts, i.e., the ones who bought the most, i.e., the most valuable customers.** Conversely Amazon judged optimizing p99.99 (1 in 10,000) **too expensive for insufficient benefit** — very high percentiles are easily affected by random events outside your control and returns diminish.
#### 2.5 What the latency-vs-revenue data actually says [#25-what-the-latency-vs-revenue-data-actually-says]
The book is unusually careful here, and it's worth copying the skepticism:
| Study | Claim | Status |
| ---------------------- | -------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Google 2006 | 400 ms → 900 ms slowdown ≈ 20% drop in traffic and revenue | **Often cited, unreliable** |
| Google 2009 | +400 ms latency → only **0.6%** fewer searches/day | Contradicts the above |
| Bing 2009 | +2 s load time → **4.3%** less ad revenue | — |
| Akamai (recent) | +100 ms → up to **7%** lower ecommerce conversion | **Same study shows very fast page loads also correlate with LOWER conversion** — because the fastest pages are often ones with no useful content (404s). No attempt to separate content effects from load-time effects → **probably not meaningful** |
| Yahoo (following year) | **20–30% more clicks** on fast searches when the fast/slow difference is ≥1.25 s | **Controlled for search-result quality** — the most trustworthy of these |
Lesson: be suspicious of the folklore numbers; the causal claim is real but weaker and better-measured than the slide-deck version.
#### 2.6 Tail latency amplification [#26-tail-latency-amplification]
High percentiles matter **most** in backend services called multiple times per end-user request.
**Even if the calls are parallel, the request waits for the slowest one.** One slow call makes the entire end-user request slow. And **the more backend calls per end-user request, the higher the probability of hitting at least one slow call** — so a *larger proportion* of end-user requests end up slow than the per-service percentile suggests.
Quick math intuition: if each backend call independently has a 1% chance of exceeding its p99, then with 100 backend calls, `1 − 0.99¹⁰⁰ ≈ 63%` of end-user requests hit at least one p99-slow call. **Your service's p99 becomes your user's median.**
#### 2.7 SLOs and SLAs [#27-slos-and-slas]
* **SLO (service level objective)** — a target. Example: *median response time \< 200 ms, p99 \< 1 s, and ≥99.9% of valid requests return non-error responses.*
* **SLA (service level agreement)** — a **contract specifying what happens if the SLO is not met** (e.g., customers entitled to a refund).
The book notes plainly: **defining good availability metrics for SLOs/SLAs is not straightforward in practice.**
#### 2.8 Computing percentiles efficiently [#28-computing-percentiles-efficiently]
Need: a rolling window (e.g. last 10 minutes), recomputed every minute for a dashboard.
* Simplest: keep all response times in the window and sort. Works; often too expensive.
* Approximation libraries at minimal CPU/memory cost: **HdrHistogram, t-digest, OpenHistogram, DDSketch**.
> ⚠️ **Averaging percentiles is mathematically meaningless.** Not "imprecise" — meaningless. You cannot average p99 across machines, or average p99 over time to downsample a graph. **The right way to aggregate response-time data is to add the histograms.** Almost every hand-rolled dashboard gets this wrong.
***
# 2.12 Forward links (/docs/ddia/defining-nonfunctional-requirements/forward-links)
| Concept here | Where it's developed |
| --------------------------------------------------- | ------------------------------------------------ |
| Materialized views | **Ch 4** (data cubes), **Ch 13** (derived state) |
| Exactly-once semantics for fan-out | **Ch 12** — Stream Processing |
| Sharding for shared-nothing | **Ch 7** — Sharding |
| Tolerating node loss | **Ch 6** (Replication), **Ch 10** (Consensus) |
| Rolling upgrades & schema evolution | **Ch 5** — Encoding and Evolution |
| Timeouts, unbounded delays, GC pauses | **Ch 9** — Trouble with Distributed Systems |
| Automatic vs manual rebalancing | **Ch 7** — Sharding |
| Complexity of the distributed setup you're avoiding | **Ch 9** |
# 2. Defining Nonfunctional Requirements (/docs/ddia/defining-nonfunctional-requirements)
> "The Internet was done so well that most people think of it as a natural resource like the Pacific Ocean, rather than something that was man-made." — Alan Kay
**Functional requirements** = what screens, what buttons, what each operation does.
**Nonfunctional requirements** = fast, reliable, secure, legally compliant, maintainable. Usually unwritten because they seem obvious — and *just as important*: **an app that is unbearably slow or unreliable might as well not exist.**
This chapter covers four of them:
1. **Performance** — defining and measuring it
2. **Reliability** — continuing to work correctly even when things go wrong
3. **Scalability** — efficiently adding capacity as load grows
4. **Maintainability** — keeping it workable long-term
***
# 2.5 Maintainability (/docs/ddia/defining-nonfunctional-requirements/maintainability)
Software doesn't wear out or suffer material fatigue. But requirements evolve, the environment changes (dependencies, platform), and bugs need fixing.
> **The majority of the cost of software is not initial development but ongoing maintenance** — fixing bugs, keeping systems operational, investigating failures, adapting to new platforms, modifying for new use cases, repaying technical debt, adding features.
Legacy pain compounds: outdated technologies few engineers understand (mainframes, COBOL); **institutional knowledge of how and why the system was designed lost as people leave**; fixing other people's mistakes. Because systems are intertwined with the human organizations they support, **maintenance is as much a people problem as a technical one.**
> **Every system we create today will one day become a legacy system, if it is valuable enough to survive.**
Three principles:
| Principle | Goal |
| ---------------- | -------------------------------------------------------------------------------------------------------------- |
| **Operability** | Make it easy for the organization to keep the system running smoothly |
| **Simplicity** | Make it easy for new engineers to understand — well-understood, consistent patterns; no unnecessary complexity |
| **Evolvability** | Make it easy to change the system later, for unanticipated use cases |
#### 5.1 Operability [#51-operability]
> "Good operations can often work around the limitations of bad (or incomplete) software, but good software cannot run reliably with bad operations."
**Automation is essential** at thousands of machines — manual maintenance would be unreasonably expensive. **But automation is two-edged:**
* There will always be edge cases (rare failure scenarios) requiring manual intervention, and **the cases that can't be automated tend to be the most complex — so greater automation requires a MORE skilled operations team**
* **An automated system that goes wrong is often harder to troubleshoot than one where an operator does some steps manually**
* Therefore **more automation is not always better for operability.** The sweet spot depends on your application and organization.
**What data systems can do to be operable:**
* Support **monitoring** of key metrics and **observability** tools for runtime behavior
* **Avoid dependency on individual machines** — let machines be taken down for maintenance while the system keeps running
* Good documentation and an **easy-to-understand operational model**: *"If I do X, Y will happen"*
* **Good defaults, but freedom to override them**
* **Self-healing where appropriate, but manual control over system state when needed**
* **Predictable behavior, minimizing surprises**
#### 5.2 Simplicity [#52-simplicity]
Complexity slows everyone down and raises maintenance cost. A project mired in it is a **big ball of mud**. In complex software there is **greater risk of introducing bugs when making a change**, because hidden assumptions, unintended consequences, and unexpected interactions are more easily overlooked.
**Simplicity is subjective — there is no objective standard.** The book's own counterexample: *is a system that hides a complex implementation behind a simple interface simpler than one with a simple implementation that exposes more internal detail?* No settled answer.
**Essential vs accidental complexity** (essential = inherent to the problem domain; accidental = arising only from limitations of our tooling) is a useful frame but **also flawed, because the boundary shifts as tooling evolves.**
**Abstraction is the best tool we have.** A good abstraction hides implementation detail behind a clean façade and can serve a wide range of applications. Beyond reuse efficiency, **quality improvements in the abstracted component benefit every application that uses it.**
Examples of abstraction:
* **High-level languages** hide machine code, CPU registers, and system calls
* **SQL** hides complex on-disk and in-memory data structures, **concurrent requests from other clients**, and **inconsistencies after crashes**
Application-level abstraction methodologies: **design patterns**, **domain-driven design (DDD)**. This book is instead about **general-purpose abstractions you build applications on: database transactions, indexes, and event logs.** You can implement DDD on top of these foundations.
#### 5.3 Evolvability [#53-evolvability]
Requirements are in constant flux: new facts learned, unanticipated use cases, changed business priorities, new feature requests, platform replacement, legal/regulatory change, growth forcing architectural change.
Agile gives an organizational framework; TDD and refactoring are its technical tools. **Evolvability** is the word for agility **at the data-system level** (across several applications/services with different characteristics).
**Evolvability is closely linked to simplicity and abstraction.** Loosely coupled, simple systems are easier to modify than tightly coupled, complex ones.
> **The major factor making change difficult in large systems is IRREVERSIBILITY.** Migrating from one database to another: **if you cannot switch back when the new one has problems, the stakes are much higher.** Irreversible actions must be taken very carefully. **Minimizing irreversibility improves flexibility.**
This is the design principle that most directly translates into daily practice: prefer dual-writes over cutovers, expand-then-contract migrations over rename-in-place, feature flags over branch-and-deploy, and shadow traffic over big-bang switches.
***
# 2.7 Production failure catalog for this chapter (/docs/ddia/defining-nonfunctional-requirements/production-failure-catalog-chapter)
| Symptom | Underlying concept |
| ------------------------------------------------- | --------------------------------------------------------------------- |
| Latency fine at 60% load, catastrophic at 85% | Queueing delays rise sharply near capacity |
| Outage persists after the traffic spike ends | **Metastable failure** / retry storm |
| One user action produces 81 backend calls | Retries stacked at multiple layers, no retry budget |
| p99 looks fine, users complain constantly | Measured server-side; missing queueing + network latency |
| Dashboard p99 disagrees with reality | Averaged percentiles across time or instances |
| Benchmark says p99 = 12 ms, prod says 900 ms | Coordinated omission in the load generator |
| End-user p50 ≈ backend p99 | **Tail latency amplification** across many backend calls |
| A "fast" endpoint is slow only for big customers | Per-user data volume skew — the Amazon p999 argument |
| Everything fails at once, same error | **Correlated software fault** — same bug on every node |
| SSDs all die in the same week, 4 years in | Firmware bug at 32,768 hours (2¹⁵ counter overflow) |
| Wrong results, no crash, no error | Silent data corruption — \~1 in 1,000 machines has a bad CPU core |
| Whole rack unavailable despite redundancy | Correlated component failures; redundancy assumed independence |
| Leading cause of outages overall | **Operator configuration changes**, not hardware |
| Incident recurs after "we told Bob to be careful" | Blame instead of blameless postmortem; sociotechnical cause untouched |
| Migration can't be rolled back | **Irreversibility** — the main obstacle to evolvability |
| Autoscaler oscillates and causes incidents | Automation added where load was predictable enough for manual scaling |
***
# 2.3 Reliability and Fault Tolerance (/docs/ddia/defining-nonfunctional-requirements/reliability-fault-tolerance)
Reliability ≈ **"continuing to work correctly, even when things go wrong."**
"Working correctly" typically expects:
* The application performs the function the user expected
* It tolerates the user making mistakes or using it in unexpected ways
* Performance is good enough for the use case, under expected load and data volume
* It prevents unauthorized access and abuse
#### 3.1 Fault vs failure [#31-fault-vs-failure]
| Term | Definition |
| ----------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| **Fault** | A particular **part** of a system stops working correctly — a hard drive malfunctions, a machine crashes, a dependency has an outage |
| **Failure** | The **system as a whole** stops providing the required service to the user — i.e., it does not meet the SLO |
They're the same thing at different levels. A drive that stops working has "failed" as a drive; in a system with multiple drives, that's merely a *fault*, which the bigger system may tolerate by having a copy elsewhere.
**Fault-tolerant** = continues providing the required service in spite of certain faults. A part whose fault escalates into whole-system failure is a **single point of failure (SPOF)**.
Case-study example: during fan-out, a machine updating materialized timelines crashes. Fault tolerance requires another machine to take over **without missing any posts that should have been delivered, and without duplicating any** — that's **exactly-once semantics** (Ch 12).
**Fault tolerance is always bounded**: "at most 2 drives at once," "at most 1 of 3 nodes." Tolerating *any* number of faults is meaningless — if all nodes crash, nothing can be done. (If Earth is swallowed by a black hole, you'd need hosting in space; good luck with the budget line item.)
#### 3.2 Fault injection and chaos engineering [#32-fault-injection-and-chaos-engineering]
**Counterintuitively, in fault-tolerant systems it makes sense to *increase* the rate of faults deliberately** — e.g., randomly killing processes without warning.
**Why:** many critical bugs are due to **poor error handling**. Deliberately inducing faults ensures the fault-tolerance machinery is **continually exercised and tested**, raising confidence it will work when faults occur naturally. **Chaos engineering** is the discipline built around this.
The exception where prevention beats cure: **security**. If an attacker compromised the system and got sensitive data, that **cannot be undone**. This book mostly deals with curable faults.
#### 3.3 Hardware faults — the actual numbers [#33-hardware-faults--the-actual-numbers]
| Component | Failure characteristics |
| -------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Magnetic HDD | **2%–5% fail per year.** In a 10,000-disk cluster → **on average one disk failure per day.** Getting more reliable, but rates remain significant |
| SSD | **0.5%–1% fail per year.** Small bit errors auto-corrected, but **uncorrectable errors occur \~once per year per drive**, even in nearly-new drives — **a higher error rate than magnetic HDDs** |
| PSUs, RAID controllers, memory modules | Fail too, less often than disks |
| **CPU cores** | **\~1 in 1,000 machines has a core that occasionally computes the WRONG result**, likely from manufacturing defects. Sometimes crashes; **sometimes just returns a wrong answer** |
| RAM | Corruption from cosmic rays or permanent physical defects. **Even with ECC, >1% of machines hit an uncorrectable error per year**, typically crashing the machine and requiring module replacement. Certain **pathological memory access patterns flip bits with high probability** (Rowhammer) |
| Datacenter | Power outage, network misconfiguration; or permanent destruction by fire, flood, earthquake. A **solar storm** could damage power grids and undersea cables. Rare, but catastrophic if you can't tolerate losing a DC |
> The "1 in 1,000 CPUs silently computes wrong results" fact is the scariest one here. It breaks the assumption every program makes. It's why large operators run continuous verification and why Byzantine-ish thinking (Ch 9) isn't purely academic.
**Small system:** these are rare enough to ignore as long as you can replace faulty hardware easily.
**Large system:** hardware faults **happen often enough that they become part of normal system operation.**
#### 3.4 Redundancy — and its limits [#34-redundancy--and-its-limits]
First response: hardware redundancy. RAID across disks in a machine; dual power supplies; hot-swappable CPUs; datacenter batteries and diesel generators. This can keep a machine running for years.
> **Redundancy is most effective when component faults are INDEPENDENT** — when one fault doesn't change the likelihood of another. **Experience shows significant correlations between component failures.** Whole-rack and whole-datacenter unavailability still happens more often than we'd like.
Hence the cloud posture: **focus less on the reliability of individual machines; make services highly available by tolerating faulty nodes at the software level.** Cloud providers expose **availability zones** to tell you which resources are physically co-located — co-located resources are more likely to fail together.
**Operational bonus of machine-level fault tolerance:** a single-server system needs planned downtime to reboot for OS security patches; a multi-node fault-tolerant system is patched by **restarting one node at a time without affecting users** — a **rolling upgrade** (Ch 5).
#### 3.5 Software faults — the correlated, dangerous kind [#35-software-faults--the-correlated-dangerous-kind]
Hardware faults are weakly correlated but **mostly independent**. **Software faults are often HIGHLY correlated, because many nodes run the same software and therefore have the same bugs.** Harder to anticipate, and they **cause many more system failures than uncorrelated hardware faults.**
Real examples given:
* **The 2012 leap second**: a Linux kernel bug caused many Java applications to hang **simultaneously**, taking down several internet services.
* **The 32,768-hour SSD firmware bug**: all SSDs of certain models fail after precisely 32,768 hours (\<4 years) of operation, **rendering data unrecoverable**. (Note the number: 2¹⁵ — a signed 16-bit counter overflowing.)
* **Runaway process** consuming a shared limited resource: CPU, memory, disk space, network bandwidth, or threads. A process consuming too much memory on a large request gets OOM-killed; a client-library bug drives far higher request volume than anticipated.
* **A dependency slows down, becomes unresponsive, or returns corrupted responses.**
* **Emergent behavior from interactions between systems** that doesn't occur when each is tested in isolation.
* **Cascading failures**: one component's problem overloads another, slowing it, which brings down a third.
**The general shape:** these bugs **lie dormant for a long time until an unusual set of circumstances triggers them.** At that point it's revealed that **the software made an assumption about its environment** — usually true, and eventually not.
**No quick solution.** Lots of small things help: carefully thinking about assumptions and interactions; thorough testing; **process isolation**; **allowing processes to crash and restart**; avoiding feedback loops like retry storms; measuring, monitoring, and analyzing behavior **in production**.
#### 3.6 Humans and reliability [#36-humans-and-reliability]
> **One study of large internet services found that CONFIGURATION CHANGES BY OPERATORS were the leading cause of outages, with hardware faults playing a role in only 10%–25% of cases.**
But: **blaming people for mistakes is counterproductive.** "Human error" is **not the cause** of an incident — it's a **symptom of a problem with the sociotechnical system** in which people are doing their best. Complex systems have **emergent behavior**; unexpected interactions between components also cause failures.
**Technical measures that reduce the impact of human mistakes:**
* Thorough testing — handwritten tests **and property testing over lots of random inputs**
* **Rollback mechanisms** for quickly reverting config changes
* **Gradual rollouts** of new code
* Detailed, clear **monitoring**; **observability** tools for diagnosing production issues
* **Well-designed interfaces that encourage the right thing and discourage the wrong thing**
**And the honest organizational point:** all of these cost time and money. Given a choice between more features and more testing, many organizations understandably choose features. So when a preventable mistake occurs, blaming the individual makes no sense — **the problem is the organization's priorities.**
**Blameless postmortems**: after an incident, people share full details **without fear of punishment**, so others can learn to prevent similar problems. The process may reveal a need to change business priorities, invest in neglected areas, change incentives, or escalate a systemic issue to management.
> **Be suspicious of simplistic answers.** "Bob should have been more careful deploying that change" is not productive — but **neither is "we must rewrite the backend in Haskell."** Management should learn how the sociotechnical system actually works from the people who work with it daily, and improve it from that feedback.
#### 3.7 How important is reliability? — the Post Office Horizon scandal [#37-how-important-is-reliability--the-post-office-horizon-scandal]
Reliability isn't only for nuclear plants and air traffic control. A few minutes or hours of outage is tolerable in many applications; **permanent data loss or corruption is catastrophic.** (Consider a parent whose entire photo/video record of their children lives in your app. Would they even know how to restore from a backup?)
**The Horizon case (1999–2019):** hundreds of British Post Office branch managers were convicted of theft or fraud because the accounting software showed shortfalls in their accounts. Many of those shortfalls were **software bugs**. Convictions were eventually overturned — **probably the largest miscarriage of justice in British history.**
The enabling cause is legal, not technical: **English law assumed computers operate correctly, and therefore that computer-produced evidence is reliable, unless evidence exists to the contrary.** Engineers may laugh at the idea of bug-free software, but that's "little solace to the people who were wrongfully imprisoned, declared bankruptcy, or even committed suicide."
Sacrificing reliability to cut development cost is sometimes right (prototyping for an unproven market) — **but be conscious that you're cutting corners, and keep the consequences in mind.**
***
# 2.4 Scalability (/docs/ddia/defining-nonfunctional-requirements/scalability)
**Scalability = a system's ability to cope with increased load.**
#### 4.1 The "you're not Google" caveat, handled properly [#41-the-youre-not-google-caveat-handled-properly]
For a new product with few users, the overriding engineering goal is **keeping the system simple and flexible** so you can adapt as you learn what customers need. Worrying about hypothetical future scale is counterproductive:
* Best case: wasted effort and premature optimization
* **Worst case: it locks you into an inflexible design and makes the application harder to evolve**
#### 4.2 Scalability is not a label [#42-scalability-is-not-a-label]
**It is meaningless to say "X is scalable" or "Y doesn't scale."** The real questions:
1. If the system grows **in a particular way**, what are our options for coping?
2. How can we **add computing resources** to handle the additional load?
3. **Based on current growth projections, when will we hit the limits of our current architecture?**
#### 4.3 Understanding load [#43-understanding-load]
You must first describe **current** load before you can ask "what if it doubles?"
Usually a **throughput** measure: requests/second, GB of new data per day, checkouts per hour. Sometimes **the peak** of a variable quantity (simultaneously online users).
**Other statistical characteristics that change the access pattern and hence the scalability requirement:**
* read:write ratio
* cache hit rate
* **number of data items per user** (followers, in the case study)
* whether the average case matters, or **whether the bottleneck is dominated by a small number of extreme cases**
Then ask the growth question in two symmetrical forms:
* **Fix resources, increase load** → how is performance affected?
* **Fix performance, increase load** → how much must resources increase?
Goal: **stay within SLA while minimizing cost.**
**Linear scalability**: doubling resources handles twice the load at the same performance — considered good. Occasionally you do better than linear (economies of scale, better distribution of peak load). **Much more likely, cost grows faster than linearly** — e.g., with a lot of data, processing a single write may involve more work than with a small amount of data, even for the same request size.
#### 4.4 Shared-memory vs shared-disk vs shared-nothing [#44-shared-memory-vs-shared-disk-vs-shared-nothing]
| Architecture | Also called | Reality |
| ------------------ | ------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Shared-memory** | vertical scaling, scaling up | Individual CPU cores aren't getting much faster, but you can buy more cores/RAM/disk. Parallelism via threads sharing RAM. **Cost grows faster than linearly** — a machine with 2× resources costs significantly more than 2×, and **because of bottlenecks it likely can't actually handle 2× the load** |
| **Shared-disk** | — | Multiple machines with independent CPU/RAM, data on a shared disk array over a fast network (**NAS / SAN**). Traditionally used for on-premises data warehousing. **Contention and locking overhead limit scalability** |
| **Shared-nothing** | horizontal scaling, scaling out | Each node has own CPU, RAM, disk; all coordination in software over a conventional network. Potential for **linear scaling**, best price/performance hardware (especially in cloud), easy resource adjustment, **fault tolerance across datacenters and regions**. Costs: **explicit sharding** (Ch 7) and **all the complexity of distributed systems** (Ch 9) |
**The cloud-native variant.** Some cloud-native databases use **separate services for storage and transaction execution**, with multiple compute nodes sharing one storage service. Superficially like shared-disk — **but it avoids the old scalability problems because the storage service exposes a specialized API designed for the database's specific needs, rather than a filesystem (NAS) or block device (SAN) abstraction.** That API-level specialization is the whole trick.
#### 4.5 Principles for scalability [#45-principles-for-scalability]
> **There is no generic, one-size-fits-all scalable architecture ("magic scaling sauce").**
Illustration: a system for **100,000 requests/s × 1 kB** looks *very* different from one for **3 requests/minute × 2 GB** — even though both move \~100 MB/second.
> **An architecture appropriate for one level of load is unlikely to cope with 10× that load.** On a fast-growing service you will probably **rethink your architecture on every order-of-magnitude load increase.** It is usually **not worth planning more than one order of magnitude ahead.**
**Principle 1 — decompose.** Break the system into smaller components that operate **largely independently**. This underlies microservices, sharding (Ch 7), stream processing (Ch 12), and shared-nothing. The challenge is **knowing where to draw the line** between what belongs together and what belongs apart.
**Principle 2 — don't over-complicate.**
* If a single-machine database will do the job, prefer it to a complicated distributed setup.
* **Autoscaling is cool, but if your load is fairly predictable, a manually scaled system may have fewer operational surprises.**
* A system with 5 services is simpler than one with 50.
* Good architectures usually involve a **pragmatic mixture** of approaches.
***
# 2.11 Self-test (/docs/ddia/defining-nonfunctional-requirements/self-test)
Why is a nonfunctional requirement "just as important" as functionality?
Compute the query-on-read load for the social network from first principles. Which two separate problems does the naive design have?
What is fan-out? Why does materializing timelines trade write cost for read cost, and why is that trade favourable here?
Give the two extreme cases in the timeline design. Why is dropping writes acceptable in one and not the other?
Distinguish response time, service time, queueing delay, and latency. Which does the client experience?
Why must response times be measured on the client side?
What is head-of-line blocking, and why does a small number of slow requests hurt so many fast ones?
Why is the mean a bad "typical" latency? What is it good for?
Reproduce the Amazon p999 argument — why that percentile, and why *not* p9999?
Why is averaging percentiles meaningless, and what is the correct aggregation?
Define tail latency amplification. Why does parallelizing backend calls not help?
Distinguish an SLO from an SLA.
Define a metastable failure. Draw the loop. Why doesn't removing the load fix it?
Name three client-side and two server-side overload protections, and say what each does.
Distinguish a fault from a failure. What is a SPOF?
Why does it make sense to *increase* the rate of faults deliberately? What class of bug does this find?
Give the annual failure rates for HDDs and SSDs. Which fact in that list is the most unsettling, and why?
Why is redundancy less effective than the arithmetic suggests?
Why are software faults more dangerous than hardware faults? Give two real examples from the chapter.
What was the leading cause of outages in the study cited? Why is "human error" the wrong conclusion?
What is a blameless postmortem for, and what two simplistic answers should you distrust?
What was the Post Office Horizon scandal, and what legal assumption enabled it?
Why is "X is scalable" a meaningless statement? What three questions replace it?
Beyond throughput, name three statistical characteristics of load that change your design.
Compare shared-memory, shared-disk, and shared-nothing on cost curve, scaling limit, and what they demand of you.
How does the cloud-native storage/compute split differ from classic shared-disk, and why does that difference matter?
Why is there no "magic scaling sauce"? How far ahead should you plan?
Why is more automation not always better for operability?
Give three ways abstraction reduces complexity, and one reason "simplicity" resists definition.
Why is irreversibility the main obstacle to evolvability? Name three practices that reduce it.
you own a notification service handling 50,000 events/s, fanning out to email, push, and SMS providers. p99 is 400 ms and rising; during provider outages the whole service becomes unavailable for 20 minutes even after the provider recovers. Diagnose the likely cause, propose the specific protections, define the SLO you'd publish, and say exactly which metrics you'd add — including one you'd deliberately count that isn't an error.
# 2.6 Technology deep dives (/docs/ddia/defining-nonfunctional-requirements/technology-deep-dives)
***
#### 6.1 The timeline fan-out service (write-path materialization) [#61-the-timeline-fan-out-service-write-path-materialization]
**Problem it solves.** Serve a per-user merged, time-ordered feed at interactive latency for 10M concurrent users, where the query is a 200-way merge that would otherwise run 2M times per second.
**Why wasn't the alternative enough?**
* *Query on read (the SQL JOIN):* 400M lookups/s. Also unbounded per-user cost — users following 10,000 accounts can't be served.
* *Polling every 5 s:* re-asks a question whose answer usually hasn't changed; multiplies the above by online-user count.
* *A plain cache of the query result:* invalidation is the problem — any of 200 followees posting invalidates it, so the hit rate collapses exactly for active users.
**How it works internally.**
* On each post: look up followers → append a reference (post ID + timestamp, not the post body) to each follower's timeline structure.
* Timeline structure = a **bounded, sorted list per user** (Redis sorted set / a capped list), holding maybe the newest \~800–1,000 entries, with older reads falling back to the slow path.
* Fan-out runs **asynchronously via a queue**, so a spike delays delivery rather than rejecting posts.
* **Hybrid**: celebrity posts are *not* fanned out. They're stored once and merged at read time. The read path becomes `merge(materialized_timeline, recent_posts_of_followed_celebrities)`.
* Delivery to online clients uses a **subscription/push** channel, not polling.
**Deployment.** A queue (Kafka/BullMQ/SQS) between the write API and a pool of fan-out workers; a keyspace-sharded cache cluster holding timelines; a separate WebSocket/push tier holding client subscriptions.
**Monitoring.**
* **Fan-out lag** — time from post accepted → present in the last follower's timeline. This is the SLO ("5 seconds") and must be measured at a **high percentile**, not the mean.
* Queue depth and consumer lag per partition
* Fan-out factor distribution (p50 / p99 / max) — this is how you detect a new celebrity before they hurt you
* Timeline cache hit rate and eviction rate
* Dropped-write counter (deliberate, for the follow-many case) — you *must* count what you deliberately drop, or you can't tell it from a bug
**Scaling.** Shard timelines by user ID. Scale fan-out workers horizontally, partitioning the queue by *poster* ID. The scaling limit is not throughput but **skew**: one celebrity post is one queue message that expands to 100M writes. That single message must be **split into sub-batches** or it becomes a stuck partition.
**Backup.** Timelines are **derived data** — the backup is the ability to rebuild them from `posts` + `follows`. Test that rebuild, because you will need it after a bad deploy corrupts them. The systems of record (`posts`, `follows`, `users`) need real backups.
**What actually breaks in production.**
* **A user crosses the celebrity threshold and nobody notices.** The threshold must be dynamic and monitored, not a constant someone picked in 2019.
* **The fan-out queue backs up during a spike and never drains**, because delivery cost grew with the backlog — a metastable failure.
* **Duplicate or missing posts after a worker crashes** mid-fan-out: exactly-once semantics is genuinely hard (Ch 12). The usual practical answer is idempotent writes keyed on `(timeline_owner, post_id)`.
* **Unfollow/block/delete races** — a post lands in a timeline *after* the user blocked the author, because the fan-out was already in flight. Requires a read-time filter as a backstop; you cannot fix it purely at write time.
* **Cache cluster restart** → a stampede of timeline rebuilds hits the database simultaneously.
* **Deleting a post** requires removing it from N million timelines, or filtering at read time. Almost everyone does read-time filtering and accepts the tombstone cost.
***
#### 6.2 Percentile-tracking / metrics pipelines (HdrHistogram, t-digest, DDSketch, Prometheus) [#62-percentile-tracking--metrics-pipelines-hdrhistogram-t-digest-ddsketch-prometheus]
**Problem it solves.** Continuously compute p50/p95/p99/p999 over a rolling window, cheaply, across many machines.
**Why wasn't the simple approach enough?** Keeping every response time in the window and sorting it every minute is correct but too expensive at high request rates — it's O(n) memory per window and O(n log n) per refresh, per instance.
**Why weren't averages enough?** The mean tells you nothing about how many users experienced a delay, and it's dominated by neither the typical case nor the tail.
**How they work internally.**
* **HdrHistogram** — fixed bucket boundaries with configurable *relative* precision across a huge dynamic range (µs to hours). Recording is O(1) (an array increment). Memory is fixed and known upfront. **Histograms are addable**, which is the whole point.
* **t-digest** — adaptive clustering of the distribution with **much finer resolution at the extremes** (q→0 and q→1) than in the middle, giving very accurate high percentiles in small space. Mergeable.
* **DDSketch** — buckets with **relative-error guarantees** (e.g. every quantile accurate to within 1%), which is the guarantee you actually want for latency. Mergeable.
* **Prometheus histograms** — fixed `le` buckets; `histogram_quantile()` interpolates. Accuracy depends entirely on whether you chose bucket boundaries near your SLO. Prometheus *summaries* compute quantiles client-side and are **not aggregatable across instances** — the classic trap.
**Deployment.** Record in-process; export raw **bucket counts** (never precomputed quantiles) to a collector; aggregate by summing buckets; compute quantiles at query time.
**Monitoring the monitoring.** Metric cardinality (labels × values — the #1 cause of a monitoring system falling over), scrape duration, and samples dropped.
**Scaling.** Cardinality control is the entire scaling story: never label a metric with user ID, request ID, or a raw URL path.
**What actually breaks.**
* **Averaged percentiles.** Someone downsamples a p99 graph from 15 s to 5 min resolution by averaging, or averages p99 across 50 pods. Both are meaningless, and both silently *understate* the tail.
* **Prometheus summaries aggregated across replicas** — same bug, wearing a different hat.
* **Bucket boundaries that don't bracket the SLO** — your SLO is 200 ms and your buckets jump 100 ms → 250 ms → 500 ms, so `histogram_quantile` interpolates linearly across a bucket and reports fiction.
* **Server-side-only measurement**, missing queueing delay and network latency — the system looks healthy while users suffer (see §2.3).
* **Coordinated omission** — a load generator that waits for a response before sending the next request systematically *fails to record* the slow period, because it stopped sending requests during it. HdrHistogram has explicit correction for this; most homegrown benchmarks don't, and they under-report tails by orders of magnitude.
***
#### 6.3 Overload protection (circuit breakers, load shedding, backpressure, token buckets) [#63-overload-protection-circuit-breakers-load-shedding-backpressure-token-buckets]
**Problem it solves.** Prevent a system near capacity from entering the self-sustaining metastable failure state where it can't recover even after load drops.
**Why weren't plain retries enough?** Retries *are the mechanism of the failure*. Naive retry converts a transient slowdown into a permanent outage by multiplying load exactly when the system can least afford it.
**Why wasn't autoscaling enough?** Scaling takes tens of seconds to minutes; a retry storm compounds in seconds. And if the bottleneck is a shared database, adding stateless replicas makes it worse.
**How they work internally.**
* **Exponential backoff with jitter** — retry after `base × 2^attempt`, multiplied by a random factor. Without jitter, all clients that failed at time T retry together at T+1s, T+3s, T+7s: synchronized thundering herds. *Full jitter* (`random(0, base × 2^attempt)`) is the standard recommendation.
* **Circuit breaker** — a state machine per dependency: **CLOSED** (pass through, count failures) → **OPEN** (fail fast immediately, don't even try, for a cooldown period) → **HALF-OPEN** (let a trickle through; success closes it, failure reopens). The value is *failing fast*: it stops the caller's threads from all blocking on a dead dependency.
* **Token bucket** — a bucket refilled at rate R with capacity B. Each request consumes a token; empty bucket = reject. Allows bursts up to B while bounding sustained rate at R. Used both for rate limiting and, importantly, for **budgeting retries** (e.g., retries may consume at most 10% of the token budget — so retries can't dominate under widespread failure).
* **Load shedding** — the server watches a saturation signal (queue depth, or better, **queue wait time**) and rejects requests above a threshold with 503/429. Rejecting cheaply is the point: a fast rejection costs almost nothing, a slow timeout costs a thread and a connection.
* **Backpressure** — propagate "slow down" upstream rather than buffering. In practice: bounded queues (an unbounded queue *is* the bug), TCP flow control, gRPC/HTTP2 flow control, `Retry-After` headers, and reactive-streams style demand signalling.
* **Adaptive concurrency limits** (Netflix's `concurrency-limits`, TCP-Vegas-style) — infer the optimal in-flight request limit from observed latency gradient, instead of hardcoding a thread-pool size.
**Deployment.** Circuit breakers and backoff in the client library (or the service mesh — Envoy/Istio outlier detection). Load shedding at the server edge and *also* inside the service in front of the expensive resource. Rate limits at the API gateway.
**Monitoring.** Rejected-request rate by reason; circuit-breaker state transitions per dependency; retry rate as a *fraction of total requests* (the key early-warning signal); queue wait time (better than queue depth); and, crucially, **goodput** — successful requests per second — not just throughput.
**Scaling.** Limits must be *per-dependency* and *adaptive*. Static thresholds picked once become wrong the moment hardware or traffic mix changes.
**What actually breaks.**
* **Retries at every layer.** Client retries 3×, the SDK retries 3×, the mesh retries 3×, the load balancer retries 3× → one user action becomes 81 backend requests. This is the single most common cause of self-inflicted outages.
* **Circuit breaker with too coarse a granularity** — one shared breaker for a whole service trips because one *endpoint* is degraded, taking out healthy functionality.
* **Unbounded queues** anywhere in the path, which convert "reject fast" into "accept and time out later" — the worst of both worlds, since you did the work *and* the client gave up.
* **Load shedding that sheds the wrong things** — dropping health checks or control-plane traffic, so the orchestrator kills instances that were merely busy, and capacity *drops* under load.
* **No jitter**, producing synchronized retry waves that look like a clean sawtooth on the graph.
* **Timeouts longer than the caller's timeout** — the downstream keeps working on a request no one is waiting for, burning capacity on garbage.
***
#### 6.4 Chaos engineering / fault injection [#64-chaos-engineering--fault-injection]
**Problem it solves.** Fault-tolerance code is the least-exercised code in the system, and **many critical bugs are due to poor error handling**. Untested failover is not failover; it's a hypothesis.
**Why wasn't testing enough?** Unit and integration tests exercise the code you thought of. Production failures come from **emergent behavior between systems that doesn't occur when each is tested in isolation**, and from environmental assumptions that were true until they weren't.
**How it works.** Form a hypothesis about steady-state behavior (a business metric, e.g. orders/minute), inject a fault in a bounded blast radius, and check whether steady state holds. Fault types, roughly in order of increasing realism: process kill → CPU/memory pressure → **latency injection** → packet loss → dependency error injection → **network partition** → zone loss. Latency injection is consistently the highest-value one, because slow is harder than dead.
**Deployment.** Start in staging, then production with a small blast radius, during business hours, with an owner watching and a kill switch. Run experiments on a schedule so they keep testing the *current* system, not the one from six months ago.
**Monitoring.** The steady-state business metric first; then error rates, saturation, and — the actual point of the exercise — **whether the alert fired and whether the runbook worked**.
**What actually breaks.** Running chaos experiments without a kill switch or blast-radius limit; running them at 3 a.m. when nobody can respond; testing only instance termination (the easy case) and never latency or partial failure; and organizations that run chaos experiments but don't fix what they find, which converts the practice into theatre.
***
# 2.9 Terminology introduced here (/docs/ddia/defining-nonfunctional-requirements/terminology-introduced-here)
# 2.10 Worked examples (/docs/ddia/defining-nonfunctional-requirements/worked-examples)
**① The fan-out arithmetic, from scratch.**
Query-on-read: 10M online users ÷ 5 s polling = **2M queries/s**; × 200 followees = **400M lookups/s**.
Fan-out-on-write: 5,800 posts/s × 200 followers = **1.16M writes/s**.
Ratio = 400M / 1.16M ≈ **345×**. And the write path has a further advantage: **it can be queued and delayed during a spike, while the read path cannot** — users are waiting.
**② Where hybrid fan-out becomes mandatory.** A user with 100M followers posts.
Pure write-path: **100M timeline inserts from one event.** At 1M writes/s of total system capacity, that single post consumes **100 seconds of the entire cluster's write budget.** Read-path merge: **1 stored post + a merge at read time.** The threshold isn't a matter of taste — **any follower count where `followers × post_rate` approaches system write capacity forces the read-path treatment.**
**③ Tail latency amplification, precisely.** Each backend call independently exceeds its p99 with probability 0.01.
| Backend calls | P(at least one slow) |
| ------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------- |
| 1 | 1% |
| 10 | 9.6% |
| 50 | 39.5% |
| 100 | **63.4%** |
| 200 | **86.6%** |
| **At 100 calls, your service's p99 has become your user's median.** The only fixes are fewer calls, hedged requests, or a much better per-service p99. | |
**④ Queueing near capacity.** A single-server queue at utilization ρ: average wait ≈ `service_time × ρ/(1−ρ)`.
| ρ | Wait, as a multiple of service time |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------- |
| 0.5 | 1× |
| 0.8 | 4× |
| 0.9 | 9× |
| 0.95 | **19×** |
| 0.99 | **99×** |
| **Going from 80% to 95% utilization — a 19% capacity gain — costs you a 5× increase in wait time.** This is why "we're only at 90% CPU" is not reassuring. | |
**⑤ Retry amplification.** Client retries 3×, SDK retries 3×, service mesh retries 3×, load balancer retries 3×.
Worst case = 3⁴ = **81 backend requests per user action.** During a partial outage where most requests fail, **your own infrastructure multiplies the load 81-fold at exactly the moment it can least afford it.** A retry *budget* (retries ≤ 10% of traffic) caps this at 1.1×, regardless of how many layers exist.
**⑥ Disk failures as normal operation.** 10,000 disks at a 3.5% annual failure rate.
Failures/year = 350 → **\~1 per day.** At 100,000 disks: **\~10 per day.** *"In a large-scale system, hardware faults happen often enough that they become part of normal system operation"* — this is that sentence as a number.
**⑦ Averaging percentiles is meaningless — demonstrated.** Two pods, each 1,000 requests.
Pod A: 990 requests at 10 ms, 10 at 1,000 ms → p99 = 1,000 ms.
Pod B: 1,000 requests at 10 ms → p99 = 10 ms.
**Averaged p99 = 505 ms. True combined p99** (2,000 requests, the 20th slowest): **10 ms**, because only 10 of 2,000 requests exceed it — 0.5%, below the 1% threshold. **The average is off by 50×, and in the dangerous direction of neither being right nor conservative.**
***
# 14.7 The book, in one page (/docs/ddia/doing-right-thing/book-one-page)
| Ch | What it established |
| ------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **1** | **Trade-offs.** Operational vs analytical; cloud vs self-hosted; distributed vs single-node; business needs vs user rights |
| **2** | **Nonfunctional requirements.** Performance (percentiles, tail amplification), reliability (fault vs failure), scalability (no magic sauce), maintainability (operability, simplicity, evolvability) |
| **3** | **Data models.** Relational, document, graph, event sourcing, DataFrames; SQL, Cypher, SPARQL, Datalog, GraphQL |
| **4** | **Storage engines.** LSM-trees and B-trees for OLTP; column-oriented for analytics; inverted, multidimensional, and vector indexes for retrieval |
| **5** | **Encoding and evolution.** Backward/forward compatibility; JSON vs Protobuf vs Avro; dataflow through databases, services, workflows, and events |
| **6** | **Replication.** Single-leader, multi-leader, leaderless; replication lag anomalies; conflict resolution; version vectors |
| **7** | **Sharding.** Key-range vs hash; rebalancing; request routing; local vs global secondary indexes |
| **8** | **Transactions.** ACID's ambiguity; the anomaly table; serial execution, 2PL, SSI; 2PC and why XA fails |
| **9** | **The trouble with distributed systems.** Unreliable networks, unreliable clocks, process pauses; quorums, fencing, Byzantine faults; system models, safety vs liveness |
| **10** | **Consistency and consensus.** Linearizability and its cost; logical clocks; the equivalence of consensus, CAS, shared logs, and atomic commit |
| **11** | **Batch processing.** Unix tools → MapReduce → dataflow engines; the shuffle; workflow orchestration; serving derived data |
| **12** | **Stream processing.** Log-based brokers; CDC and log compaction; event time vs processing time; the three join types; exactly-once |
| **13** | **A philosophy.** Unbundling the database; write path vs read path; integrity without coordination; the end-to-end argument; trust but verify |
| **14** | **Doing the right thing.** Predictive analytics, bias, feedback loops, surveillance, consent, data as a toxic asset — and the responsibility that comes with all of the above |
> ### **"Given the large impact that software and data have on the world, we as engineers must remember that WE CARRY A RESPONSIBILITY TO WORK TOWARD THE KIND OF WORLD THAT WE WANT TO LIVE IN: A WORLD THAT TREATS PEOPLE WITH HUMANITY AND RESPECT. LET'S WORK TOGETHER TOWARD THAT GOAL."** [#given-the-large-impact-that-software-and-data-have-on-the-world-we-as-engineers-must-remember-that-we-carry-a-responsibility-to-work-toward-the-kind-of-world-that-we-want-to-live-in-a-world-that-treats-people-with-humanity-and-respect-lets-work-together-toward-that-goal]
# 14. Doing the Right Thing (/docs/ddia/doing-right-thing)
> "Feeding AI systems on the world's beauty, ugliness, and cruelty, but expecting it to reflect only the beauty is a fantasy." — Vinay Uday Prabhu and Abeba Birhane
**The premise:**
> **"Every system is built for a purpose; EVERY ACTION WE TAKE HAS BOTH INTENDED AND UNINTENDED CONSEQUENCES. The purpose may be as simple as making money, BUT THE CONSEQUENCES MAY BE FAR-REACHING. We, the engineers building these systems, have a RESPONSIBILITY to carefully consider those consequences and to ensure that our decisions do not cause harm."**
>
> **"We talk about data as an abstract thing, BUT REMEMBER THAT MANY DATASETS ARE ABOUT PEOPLE: their behavior, their interests, their identity. WE MUST TREAT SUCH DATA WITH HUMANITY AND RESPECT. USERS ARE HUMANS TOO, AND HUMAN DIGNITY IS PARAMOUNT."**
**On professional ethics as actually practiced:**
> **"There are guidelines to help software engineers navigate these issues, such as the ACM CODE OF ETHICS AND PROFESSIONAL CONDUCT, BUT THEY ARE RARELY DISCUSSED, APPLIED, AND ENFORCED IN PRACTICE. As a result, ENGINEERS AND PRODUCT MANAGERS SOMETIMES TAKE A CAVALIER ATTITUDE TO PRIVACY and potential negative consequences of their products."**
>
> **"A technology is NOT GOOD OR BAD IN ITSELF — what matters is HOW IT IS USED AND HOW IT AFFECTS PEOPLE. This is true of a software system like a search engine IN MUCH THE SAME WAY AS IT IS OF A WEAPON LIKE A GUN. THE ETHICAL RESPONSIBILITY IS OURS TO BEAR; IT IS NOT SUFFICIENT FOR SOFTWARE ENGINEERS TO FOCUS EXCLUSIVELY ON THE TECHNOLOGY AND IGNORE ITS CONSEQUENCES."**
>
> **"In contrast to much of computing, the concepts at the heart of ethics are NOT FIXED OR DETERMINATE in their precise meaning; THEY REQUIRE INTERPRETATION, WHICH MAY BE SUBJECTIVE. What makes something 'good' or 'bad' is not well defined, and SERIOUS DISCOURSE ON THE SUBJECT AMONG COMPUTING PROFESSIONALS IS LACKING."**
>
> ### **"ETHICS IS NOT GOING THROUGH A CHECKLIST TO CONFIRM YOU COMPLY; IT'S A PARTICIPATORY AND ITERATIVE PROCESS OF REFLECTION, IN DIALOG WITH THE PEOPLE INVOLVED, WITH ACCOUNTABILITY FOR THE RESULTS."** [#ethics-is-not-going-through-a-checklist-to-confirm-you-comply-its-a-participatory-and-iterative-process-of-reflection-in-dialog-with-the-people-involved-with-accountability-for-the-results]
***
# 14.4 A practical checklist — with the caveat the chapter itself insists on (/docs/ddia/doing-right-thing/practical-checklist-caveat-chapter)
> ⚠️ **"ETHICS IS NOT GOING THROUGH A CHECKLIST TO CONFIRM YOU COMPLY."** Treat the following as **prompts for the reflective, dialogic, accountable process** the chapter describes — not as a box-ticking exercise. A checklist you complete without argument has told you nothing.
**Before collecting:**
* **What is this data *for*? Is the answer "we might find a use later"?** That is precisely what GDPR's purpose limitation forbids, and precisely what "big data" encourages.
* **Would this still sound acceptable with "surveillance" substituted for "data"?**
* **Does data about one user reveal things about non-users who never agreed to anything?**
* **What is the worst outcome if this dataset leaks, is subpoenaed, is sold in a bankruptcy, or is inherited by a hostile regime?** Consider **all possible future governments**, not today's.
* **Can we not collect it, collect less of it, or collect it in aggregate/anonymized form?**
**Before deciding about people:**
* **Are any inputs proxies for protected traits?** Postal code and IP address predict race. Purchase history predicts pregnancy, illness, sexuality.
* **Is this "how did *you* behave" or "how did people *like you* behave"?** The second is stereotyping.
* **What is the recourse when the model is wrong about an individual?** Not the aggregate — the individual. **Is there an appeal, and does a human review it?**
* **Can you explain a specific decision to the person affected, and to a judge?**
* **Draw the feedback loop.** Does a "no" today make a "no" tomorrow more likely? **Does the system amplify existing differences or combat them?**
* **Who is accountable when it's wrong?** If the answer is "the algorithm," you have no answer.
**Before retaining:**
* **What is the retention period, and is it enforced by a job that actually runs?**
* **Is there a deletion path that reaches every derived copy** — caches, search indexes, warehouses, ML training sets, backups, event logs?
* **Have you tested the deletion path,** the way Ch 13 says you must test restores?
**On consent:**
* **Could a user actually understand what they consented to?** If the derived datasets are ones "users cannot meaningfully understand," consent is a fiction.
* **Can they refuse or withdraw without detriment?** If not, GDPR says it isn't freely given — and more importantly, it isn't consent in any moral sense.
* **Is opting out realistically available, or does the service have network effects that make it "effectively mandatory"?**
***
# 14.1 Predictive Analytics (/docs/ddia/doing-right-thing/predictive-analytics)
#### 1.1 The line that matters [#11-the-line-that-matters]
**Why organizations are structurally biased toward "no":**
> **"Payment networks want to prevent fraudulent transactions, banks want to avoid bad loans, airlines want to avoid hijackings, companies want to avoid hiring ineffective people. FROM THEIR POINT OF VIEW, THE COST OF A MISSED BUSINESS OPPORTUNITY IS LOW, BUT THE COST OF A BAD LOAN OR A PROBLEMATIC EMPLOYEE IS MUCH HIGHER, so it is expected for organizations to want to be cautious. IF IN DOUBT, THEY ARE BETTER OFF SAYING NO."**
**And why that becomes catastrophic in aggregate:**
> ### **"As algorithmic decision making becomes more widespread, someone who has (ACCURATELY OR FALSELY) BEEN LABELED AS RISKY may suffer A LARGE NUMBER OF THOSE 'NO' DECISIONS. Systematically being excluded from JOBS, AIR TRAVEL, INSURANCE COVERAGE, PROPERTY RENTAL, FINANCIAL SERVICES, and other key aspects of society is such a large constraint of an individual's freedom that it has been called 'ALGORITHMIC PRISON.'"** [#as-algorithmic-decision-making-becomes-more-widespread-someone-who-has-accurately-or-falsely-been-labeled-as-risky-may-suffer-a-large-number-of-those-no-decisions-systematically-being-excluded-from-jobs-air-travel-insurance-coverage-property-rental-financial-services-and-other-key-aspects-of-society-is-such-a-large-constraint-of-an-individuals-freedom-that-it-has-been-called-algorithmic-prison]
>
> ### **"In countries that respect human rights, THE CRIMINAL JUSTICE SYSTEM PRESUMES INNOCENCE UNTIL PROVEN GUILTY; ON THE OTHER HAND, AUTOMATED SYSTEMS CAN SYSTEMATICALLY AND ARBITRARILY EXCLUDE A PERSON FROM PARTICIPATING IN SOCIETY WITHOUT ANY PROOF OF GUILT, AND WITH LITTLE CHANCE OF APPEAL."** [#in-countries-that-respect-human-rights-the-criminal-justice-system-presumes-innocence-until-proven-guilty-on-the-other-hand-automated-systems-can-systematically-and-arbitrarily-exclude-a-person-from-participating-in-society-without-any-proof-of-guilt-and-with-little-chance-of-appeal]
#### 1.2 Bias and discrimination [#12-bias-and-discrimination]
**The hope, stated fairly first:**
> **"Decisions made by an algorithm are NOT NECESSARILY ANY BETTER OR ANY WORSE than those made by a human. EVERY PERSON IS LIKELY TO HAVE BIASES, EVEN IF THEY ACTIVELY TRY TO COUNTERACT THEM, and discriminatory practices can become CULTURALLY INSTITUTIONALIZED. THERE IS HOPE that basing decisions on data could be MORE FAIR and give a better chance to people who are often overlooked or disadvantaged."**
**Then the mechanism by which it fails:**
> **"When we develop predictive analytics and AI systems, WE ARE NOT MERELY AUTOMATING A HUMAN'S DECISION by using software to specify the rules; WE ARE LEAVING THE RULES THEMSELVES TO BE INFERRED FROM DATA. However, the patterns learned are OPAQUE: EVEN IF THE DATA INDICATES A CORRELATION, WE MAY NOT KNOW WHY."**
>
> ### **"IF THE INPUT TO AN ALGORITHM CARRIES A SYSTEMATIC BIAS, THE SYSTEM WILL MOST LIKELY LEARN AND AMPLIFY THAT BIAS IN ITS OUTPUT."** [#if-the-input-to-an-algorithm-carries-a-systematic-bias-the-system-will-most-likely-learn-and-amplify-that-bias-in-its-output]
**The proxy problem — why "just don't use protected attributes" doesn't work:**
> ### **"PREDICTIVE ANALYTICS SYSTEMS MERELY EXTRAPOLATE FROM THE PAST; IF THE PAST IS DISCRIMINATORY, THEY CODIFY AND AMPLIFY THAT DISCRIMINATION. IF WE WANT THE FUTURE TO BE BETTER THAN THE PAST, MORAL IMAGINATION IS REQUIRED, AND THAT'S SOMETHING ONLY HUMANS CAN PROVIDE. DATA AND MODELS SHOULD BE OUR TOOLS, NOT OUR MASTERS."** [#predictive-analytics-systems-merely-extrapolate-from-the-past-if-the-past-is-discriminatory-they-codify-and-amplify-that-discrimination-if-we-want-the-future-to-be-better-than-the-past-moral-imagination-is-required-and-thats-something-only-humans-can-provide-data-and-models-should-be-our-tools-not-our-masters]
#### 1.3 Responsibility and accountability [#13-responsibility-and-accountability]
> **"If a human makes a mistake, THEY CAN BE HELD ACCOUNTABLE, and the person affected CAN APPEAL. Algorithms make mistakes too, BUT WHO IS ACCOUNTABLE IF THEY GO WRONG?"**
>
> * **"When a self-driving car causes an accident, WHO IS RESPONSIBLE?"**
> * **"If an automated credit scoring algorithm systematically discriminates against people of a particular race or religion, IS THERE ANY RECOURSE?"**
> * **"If a decision by your ML system comes under judicial review, CAN YOU EXPLAIN TO THE JUDGE HOW THE ALGORITHM MADE ITS DECISION?"**
>
> ### **"PEOPLE SHOULD NOT BE ABLE TO EVADE THEIR RESPONSIBILITY BY BLAMING AN ALGORITHM."** [#people-should-not-be-able-to-evade-their-responsibility-by-blaming-an-algorithm]
**Credit scores vs predictive scoring — the crucial distinction:**
| | **Traditional credit score** | **ML-based predictive scoring** |
| ----------------------- | ---------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Question it answers** | **"HOW DID *YOU* BEHAVE IN THE PAST?"** | **"WHO IS SIMILAR TO YOU, AND HOW DID PEOPLE *LIKE YOU* BEHAVE IN THE PAST?"** |
| **Basis** | **Relevant facts about a person's ACTUAL BORROWING HISTORY** | **A MUCH WIDER RANGE OF INPUTS, MUCH MORE OPAQUE** |
| **Correction** | **Errors in the record CAN BE CORRECTED** *(although agencies "normally do not make this easy")* | **"If a decision is incorrect because of ERRONEOUS DATA, RECOURSE IS ALMOST IMPOSSIBLE"** |
| **Ethical problem** | — | **"Drawing parallels to others' behavior IMPLIES STEREOTYPING PEOPLE — for example, based on where they live (A CLOSE PROXY FOR RACE AND SOCIOECONOMIC CLASS). WHAT ABOUT PEOPLE WHO GET PUT IN THE WRONG BUCKET?"** |
**The statistical fallacy applied to individuals:**
> **"Much data is STATISTICAL in nature, which means that EVEN IF THE PROBABILITY DISTRIBUTION ON THE WHOLE IS CORRECT, INDIVIDUAL CASES MAY WELL BE WRONG. If the average life expectancy in your country is 80 years, THAT DOESN'T MEAN YOU'RE EXPECTED TO DROP DEAD ON YOUR 80TH BIRTHDAY. From the average and the distribution, YOU CAN'T SAY MUCH ABOUT THE AGE TO WHICH ONE PARTICULAR PERSON WILL LIVE."**
>
> ### **"A BLIND BELIEF IN THE SUPREMACY OF DATA FOR MAKING DECISIONS IS NOT ONLY DELUSIONAL BUT ALSO POSITIVELY DANGEROUS."** [#a-blind-belief-in-the-supremacy-of-data-for-making-decisions-is-not-only-delusional-but-also-positively-dangerous]
**The dual-use problem, stated without flinching:**
> **"Analytics can reveal financial and social characteristics of people's lives. ON THE ONE HAND, this power could be used to FOCUS AID AND SUPPORT TO HELP THOSE WHO NEED IT MOST. ON THE OTHER HAND, IT IS SOMETIMES USED BY PREDATORY BUSINESSES SEEKING TO IDENTIFY VULNERABLE PEOPLE AND SELL THEM RISKY PRODUCTS SUCH AS HIGH-COST LOANS OR WORTHLESS COLLEGE DEGREES."**
#### 1.4 Feedback loops [#14-feedback-loops]
**Echo chambers:**
> **"When services become good at predicting the content users want to see, THEY MAY END UP SHOWING PEOPLE ONLY OPINIONS THEY ALREADY AGREE WITH, leading to ECHO CHAMBERS in which STEREOTYPES, MISINFORMATION, AND POLARIZATION CAN BREED. WE ARE ALREADY SEEING THE IMPACT SOCIAL MEDIA ECHO CHAMBERS CAN HAVE ON ELECTION CAMPAIGNS."**
**The self-reinforcing spiral — trace it:**
**A second, non-obvious example:**
> **"Economists found that when GAS STATIONS IN GERMANY introduced ALGORITHMIC PRICES, COMPETITION WAS REDUCED AND PRICES FOR CONSUMERS WENT UP BECAUSE THE ALGORITHMS LEARNED TO COLLUDE."** *(Nobody programmed collusion. It emerged.)*
**The tool for anticipating this — systems thinking:**
> **"We can't always predict when such feedback loops may happen. However, MANY CONSEQUENCES CAN BE PREDICTED BY THINKING ABOUT THE ENTIRE SYSTEM — NOT JUST THE COMPUTERIZED PARTS, BUT ALSO THE PEOPLE INTERACTING WITH IT — an approach known as SYSTEMS THINKING."**
>
> ### **"DOES THE SYSTEM REINFORCE AND AMPLIFY EXISTING DIFFERENCES BETWEEN PEOPLE (e.g., MAKING THE RICH RICHER OR THE POOR POORER), OR DOES IT TRY TO COMBAT INJUSTICE? EVEN WITH THE BEST INTENTIONS, WE MUST BEWARE OF THE POSSIBILITY OF UNINTENDED CONSEQUENCES."** [#does-the-system-reinforce-and-amplify-existing-differences-between-people-eg-making-the-rich-richer-or-the-poor-poorer-or-does-it-try-to-combat-injustice-even-with-the-best-intentions-we-must-beware-of-the-possibility-of-unintended-consequences]
***
# 14.2 Privacy and Tracking (/docs/ddia/doing-right-thing/privacy-tracking)
#### 2.1 The relationship shift [#21-the-relationship-shift]
**The thought experiment:**
> **"Try REPLACING THE WORD *DATA* WITH *SURVEILLANCE*, and observe whether common phrases still sound so good:"**
>
> > *"In our **surveillance**-driven organization we collect real-time **surveillance** streams and store them in our **surveillance** warehouse. Our **surveillance** scientists use advanced analytics and **surveillance** processing in order to derive new insights."*
>
> **"This thought experiment is UNUSUALLY POLEMIC FOR THIS BOOK, *DESIGNING SURVEILLANCE-INTENSIVE APPLICATIONS*, BUT STRONG WORDS ARE NEEDED TO EMPHASIZE THIS POINT."**
#### 2.2 The scale of it [#22-the-scale-of-it]
> ### **"IN OUR ATTEMPTS TO MAKE SOFTWARE 'EAT THE WORLD,' WE HAVE BUILT THE GREATEST MASS SURVEILLANCE INFRASTRUCTURE EVER SEEN."** [#in-our-attempts-to-make-software-eat-the-world-we-have-built-the-greatest-mass-surveillance-infrastructure-ever-seen]
>
> **"We are rapidly approaching a world in which EVERY INHABITED SPACE CONTAINS AT LEAST ONE INTERNET-CONNECTED MICROPHONE — smartphones, smart TVs, voice-controlled assistants, baby monitors, and EVEN CHILDREN'S TOYS that use cloud-based speech recognition. MANY OF THESE DEVICES HAVE A TERRIBLE SECURITY RECORD."**
>
> **"Surveillance of our LOCATION AND MOVEMENTS, our SOCIAL RELATIONSHIPS AND COMMUNICATIONS, our PURCHASES AND PAYMENTS, and our HEALTH DATA has become ALMOST UNAVOIDABLE. A surveillance organization may end up KNOWING MORE ABOUT A PERSON THAN THAT PERSON KNOWS ABOUT THEMSELVES — for example, IDENTIFYING ILLNESSES OR ECONOMIC PROBLEMS BEFORE THAT INDIVIDUAL IS AWARE OF THEM."**
>
> ### **"EVEN THE MOST TOTALITARIAN AND REPRESSIVE REGIMES OF THE PAST COULD ONLY DREAM OF PUTTING A MICROPHONE IN EVERY ROOM AND FORCING EVERY PERSON TO CONSTANTLY CARRY A DEVICE CAPABLE OF TRACKING THEIR LOCATION. YET THE BENEFITS WE GET FROM DIGITAL TECHNOLOGY ARE SO GREAT THAT WE NOW VOLUNTARILY ACCEPT THIS STATE OF TOTAL SURVEILLANCE. THE DIFFERENCE IS JUST THAT THE DATA IS BEING COLLECTED BY CORPORATIONS TO PROVIDE US WITH SERVICES, RATHER THAN GOVERNMENT AGENCIES SEEKING CONTROL."** [#even-the-most-totalitarian-and-repressive-regimes-of-the-past-could-only-dream-of-putting-a-microphone-in-every-room-and-forcing-every-person-to-constantly-carry-a-device-capable-of-tracking-their-location-yet-the-benefits-we-get-from-digital-technology-are-so-great-that-we-now-voluntarily-accept-this-state-of-total-surveillance-the-difference-is-just-that-the-data-is-being-collected-by-corporations-to-provide-us-with-services-rather-than-government-agencies-seeking-control]
**Why we accept it — and why that reasoning is a privilege:**
> **"Perhaps you feel YOU HAVE NOTHING TO HIDE — in other words, YOU ARE TOTALLY IN LINE WITH EXISTING POWER STRUCTURES, YOU ARE NOT A MARGINALIZED MINORITY, AND YOU NEEDN'T FEAR PERSECUTION. NOT EVERYONE IS SO FORTUNATE."**
>
> **"Or perhaps it's because THE PURPOSE SEEMS BENIGN — not overt coercion, MERELY BETTER RECOMMENDATIONS AND MORE PERSONALIZED MARKETING. HOWEVER, COMBINED WITH THE DISCUSSION OF PREDICTIVE ANALYTICS, THAT DISTINCTION SEEMS LESS CLEAR."**
**It's already not benign:**
* **"Behavioral data on CAR DRIVING, tracked by cars WITHOUT DRIVERS' CONSENT, affecting their INSURANCE PREMIUMS"**
* **"Health insurance coverage that depends on people WEARING A FITNESS TRACKING DEVICE"**
* **"The MOVEMENT SENSOR IN A SMARTWATCH can be used to work out WHAT YOU ARE TYPING (E.G., PASSWORDS) WITH FAIRLY GOOD ACCURACY. Sensor accuracy and algorithms for analysis are ONLY GOING TO GET BETTER"**
#### 2.3 Consent and freedom of choice — four reasons "they agreed" fails [#23-consent-and-freedom-of-choice--four-reasons-they-agreed-fails]
**What GDPR actually requires of consent:**
> **Consent must be "FREELY GIVEN, SPECIFIC, INFORMED, AND UNAMBIGUOUS," and the user must be able to "REFUSE OR WITHDRAW CONSENT WITHOUT DETRIMENT" — otherwise it is not considered freely given. Any request must be written "IN AN INTELLIGIBLE AND EASILY ACCESSIBLE FORM, USING CLEAR AND PLAIN LANGUAGE," and "SILENCE, PRE-TICKED BOXES OR INACTIVITY \[DO NOT] CONSTITUTE CONSENT."**
>
> *(Consent is not the only lawful basis — others include **complying with other legislation**, **protecting somebody's life**, and **legitimate interest**, which permits e.g. **fraud prevention** — "which fraudsters would presumably not consent to." **Nevertheless, consent is the most frequently used basis in internet services.**)*
#### 2.4 What privacy actually means [#24-what-privacy-actually-means]
> **"Sometimes people claim that 'PRIVACY IS DEAD' on the grounds that some users are willing to post all sorts of things about their lives to social media. HOWEVER, THIS CLAIM IS FALSE AND RESTS ON A MISUNDERSTANDING OF THE WORD PRIVACY."**
>
> ### **"HAVING PRIVACY DOES NOT MEAN KEEPING EVERYTHING SECRET; IT MEANS HAVING THE FREEDOM TO CHOOSE WHAT TO REVEAL TO WHOM, WHAT TO MAKE PUBLIC, AND WHAT TO KEEP SECRET. THE RIGHT TO PRIVACY IS A *DECISION RIGHT*: IT ENABLES EACH PERSON TO DECIDE WHERE THEY WANT TO BE ON THE SPECTRUM BETWEEN SECRECY AND TRANSPARENCY IN EACH SITUATION. IT IS AN IMPORTANT ASPECT OF A PERSON'S FREEDOM AND AUTONOMY."** [#having-privacy-does-not-mean-keeping-everything-secret-it-means-having-the-freedom-to-choose-what-to-reveal-to-whom-what-to-make-public-and-what-to-keep-secret-the-right-to-privacy-is-a-decision-right-it-enables-each-person-to-decide-where-they-want-to-be-on-the-spectrum-between-secrecy-and-transparency-in-each-situation-it-is-an-important-aspect-of-a-persons-freedom-and-autonomy]
**The illustrative case:**
> **"Someone with a RARE MEDICAL CONDITION might be VERY HAPPY to provide their private medical data to researchers if it might help develop treatments. HOWEVER, THIS PERSON MUST HAVE A CHOICE OVER WHO MAY ACCESS THIS DATA AND FOR WHAT PURPOSE. If information about their condition could HINDER THEIR ACCESS TO MEDICAL INSURANCE OR EMPLOYMENT, this person would probably be MUCH MORE CAUTIOUS."**
**The transfer — this is the chapter's sharpest reframe:**
> ### **"When data is extracted from people through surveillance infrastructure, PRIVACY RIGHTS ARE NOT NECESSARILY ERODED BUT RATHER *TRANSFERRED TO THE DATA COLLECTOR*. Companies that acquire data essentially say, 'TRUST US TO DO THE RIGHT THING WITH YOUR DATA' — WHICH MEANS THAT THE RIGHT TO DECIDE WHAT TO REVEAL AND WHAT TO KEEP SECRET IS TRANSFERRED FROM THE INDIVIDUAL TO THE COMPANY."** [#when-data-is-extracted-from-people-through-surveillance-infrastructure-privacy-rights-are-not-necessarily-eroded-but-rather-transferred-to-the-data-collector-companies-that-acquire-data-essentially-say-trust-us-to-do-the-right-thing-with-your-data--which-means-that-the-right-to-decide-what-to-reveal-and-what-to-keep-secret-is-transferred-from-the-individual-to-the-company]
>
> **"The companies in turn CHOOSE TO KEEP MUCH OF THE OUTCOME SECRET, BECAUSE TO REVEAL IT WOULD BE PERCEIVED AS CREEPY AND WOULD HARM THEIR BUSINESS MODEL (which relies on KNOWING MORE ABOUT PEOPLE THAN OTHER COMPANIES DO). Intimate information is revealed ONLY INDIRECTLY — in the form of TOOLS FOR TARGETING ADVERTISEMENTS TO SPECIFIC GROUPS (such as those SUFFERING FROM A PARTICULAR ILLNESS)."**
>
> ### **"EVEN IF PARTICULAR USERS CANNOT BE PERSONALLY REIDENTIFIED FROM THE BUCKET OF PEOPLE TARGETED BY A PARTICULAR AD, THEY HAVE LOST THEIR AGENCY ABOUT THE DISCLOSURE OF SOME INTIMATE INFORMATION. IT IS NOT THE USER WHO DECIDES WHAT IS REVEALED TO WHOM ON THE BASIS OF THEIR PERSONAL PREFERENCES — IT IS THE COMPANY THAT EXERCISES THE PRIVACY RIGHT WITH THE GOAL OF MAXIMIZING ITS PROFIT."** [#even-if-particular-users-cannot-be-personally-reidentified-from-the-bucket-of-people-targeted-by-a-particular-ad-they-have-lost-their-agency-about-the-disclosure-of-some-intimate-information-it-is-not-the-user-who-decides-what-is-revealed-to-whom-on-the-basis-of-their-personal-preferences--it-is-the-company-that-exercises-the-privacy-right-with-the-goal-of-maximizing-its-profit]
**On "creepiness" as a managed perception, and what engineers owe:**
> **"Many companies want to AVOID BEING PERCEIVED AS CREEPY, AVOIDING THE QUESTION OF HOW INTRUSIVE THEIR DATA COLLECTION ACTUALLY IS AND INSTEAD FOCUSING ON MANAGING USER PERCEPTIONS. And even these perceptions are often managed poorly — SOMETHING MAY BE FACTUALLY CORRECT, BUT IF IT TRIGGERS PAINFUL MEMORIES, THE USER MAY NOT WANT TO BE REMINDED ABOUT IT."**
>
> ### **"With any kind of data, WE SHOULD EXPECT THE POSSIBILITY THAT IT IS WRONG, UNDESIRABLE, OR INAPPROPRIATE IN SOME WAY, AND WE NEED TO BUILD MECHANISMS FOR HANDLING THOSE FAILURES. Whether something is 'undesirable' or 'inappropriate' is DOWN TO HUMAN JUDGMENT; ALGORITHMS ARE OBLIVIOUS TO SUCH NOTIONS UNLESS WE EXPLICITLY PROGRAM THEM TO RESPECT HUMAN NEEDS. AS ENGINEERS, WE MUST BE HUMBLE, ACCEPTING AND PLANNING FOR SUCH FAILINGS."** [#with-any-kind-of-data-we-should-expect-the-possibility-that-it-is-wrong-undesirable-or-inappropriate-in-some-way-and-we-need-to-build-mechanisms-for-handling-those-failures-whether-something-is-undesirable-or-inappropriate-is-down-to-human-judgment-algorithms-are-oblivious-to-such-notions-unless-we-explicitly-program-them-to-respect-human-needs-as-engineers-we-must-be-humble-accepting-and-planning-for-such-failings]
**And the limit of privacy settings:**
> **"Privacy settings are A STARTING POINT for handing back some control. HOWEVER, REGARDLESS OF THE SETTING, THE SERVICE ITSELF STILL HAS UNFETTERED ACCESS TO THE DATA and is free to use it in any way permitted by the privacy policy. Even if the service promises NOT TO SELL the data to third parties, IT USUALLY GRANTS ITSELF UNRESTRICTED RIGHTS TO PROCESS AND ANALYZE THE DATA INTERNALLY, OFTEN GOING MUCH FURTHER THAN WHAT IS OVERTLY VISIBLE TO USERS."**
>
> **"This kind of LARGE-SCALE TRANSFER OF PRIVACY RIGHTS FROM INDIVIDUALS TO CORPORATIONS IS HISTORICALLY UNPRECEDENTED. Surveillance has always existed, BUT IT USED TO BE EXPENSIVE AND MANUAL, NOT SCALABLE AND AUTOMATED. Trust relationships have always existed — between a patient and their doctor, a defendant and their attorney — BUT IN THESE CASES THE USE OF DATA HAS BEEN STRICTLY GOVERNED BY ETHICAL, LEGAL, AND REGULATORY CONSTRAINTS."**
#### 2.5 Data as assets and power [#25-data-as-assets-and-power]
**The reframe from "exhaust" to "labor":**
**Who wants the data, beyond the company that collected it:**
> **"DATA BROKERS OPERATING IN SECRECY, purchasing, aggregating, analyzing, and reselling people's personal data. STARTUPS ARE VALUED BY THEIR USER NUMBERS, OR 'EYEBALLS' — THAT IS, BY THEIR SURVEILLANCE CAPABILITIES."**
>
> **"GOVERNMENTS want it too, and they may seek to obtain it by means of SECRET DEALS, COERCION, LEGAL COMPULSION, OR SIMPLY THEFT. WHEN A COMPANY GOES BANKRUPT, THE PERSONAL DATA IT HAS COLLECTED IS ONE OF THE ASSETS THAT GET SOLD. And because data is difficult to secure, BREACHES HAPPEN DISCONCERTINGLY OFTEN."**
> ### **"These observations have led critics to say that data is NOT JUST AN ASSET, BUT A 'TOXIC ASSET,' or at least 'HAZARDOUS MATERIAL.' MAYBE DATA IS NOT THE NEW GOLD, OR THE NEW OIL, BUT RATHER THE NEW URANIUM."** [#these-observations-have-led-critics-to-say-that-data-is-not-just-an-asset-but-a-toxic-asset-or-at-least-hazardous-material-maybe-data-is-not-the-new-gold-or-the-new-oil-but-rather-the-new-uranium]
**The five ways it goes wrong even with good intentions:**
> **"Even if we think we are capable of preventing abuse, whenever we collect data WE NEED TO BALANCE THE BENEFITS WITH THE RISK OF IT FALLING INTO THE WRONG HANDS:"**
>
> 1. **"Computer systems may be COMPROMISED BY CRIMINALS OR HOSTILE FOREIGN INTELLIGENCE SERVICES"**
> 2. **"Data may be LEAKED BY INSIDERS"**
> 3. **"The company may FALL INTO THE HANDS OF UNSCRUPULOUS MANAGEMENT THAT DOES NOT SHARE OUR VALUES"**
> 4. **"The country may be TAKEN OVER BY A REGIME THAT HAS NO QUALMS ABOUT COMPELLING US TO HAND OVER THE DATA"**
> 5. **Bankruptcy — the data is sold as an asset**
>
> ### **"When collecting data, WE NEED TO CONSIDER NOT JUST TODAY'S POLITICAL ENVIRONMENT, BUT ALL POSSIBLE FUTURE GOVERNMENTS. There is NO GUARANTEE that every government elected in the future will respect human rights and civil liberties. As Bruce Schneier observes: 'IT IS POOR CIVIC HYGIENE TO INSTALL TECHNOLOGIES THAT COULD SOMEDAY FACILITATE A POLICE STATE.'"** [#when-collecting-data-we-need-to-consider-not-just-todays-political-environment-but-all-possible-future-governments-there-is-no-guarantee-that-every-government-elected-in-the-future-will-respect-human-rights-and-civil-liberties-as-bruce-schneier-observes-it-is-poor-civic-hygiene-to-install-technologies-that-could-someday-facilitate-a-police-state]
**On power:**
> **"'KNOWLEDGE IS POWER,' as the old adage goes. And furthermore: 'TO SCRUTINIZE OTHERS WHILE AVOIDING SCRUTINY ONESELF IS ONE OF THE MOST IMPORTANT FORMS OF POWER.' This is why totalitarian governments want surveillance: IT GIVES THEM THE POWER TO CONTROL THE POPULATION. Although today's technology companies are NOT OVERTLY SEEKING POLITICAL POWER, the data and knowledge they have accumulated — MUCH OF IT SURREPTITIOUSLY, OUTSIDE OF PUBLIC OVERSIGHT — NEVERTHELESS GIVES THEM A LOT OF POWER OVER OUR LIVES."**
#### 2.6 Remembering the Industrial Revolution [#26-remembering-the-industrial-revolution]
**Bruce Schneier's framing, which is the moral center of the chapter:**
> ### **"DATA IS THE POLLUTION PROBLEM OF THE INFORMATION AGE, AND PROTECTING PRIVACY IS THE ENVIRONMENTAL CHALLENGE. ALMOST ALL COMPUTERS PRODUCE INFORMATION. IT STAYS AROUND, FESTERING. HOW WE DEAL WITH IT — HOW WE CONTAIN IT AND HOW WE DISPOSE OF IT — IS CENTRAL TO THE HEALTH OF OUR INFORMATION ECONOMY. JUST AS WE LOOK BACK TODAY AT THE EARLY DECADES OF THE INDUSTRIAL AGE AND WONDER HOW OUR ANCESTORS COULD HAVE IGNORED POLLUTION IN THEIR RUSH TO BUILD AN INDUSTRIAL WORLD, OUR GRANDCHILDREN WILL LOOK BACK AT US DURING THESE EARLY DECADES OF THE INFORMATION AGE AND JUDGE US ON HOW WE ADDRESSED THE CHALLENGE OF DATA COLLECTION AND MISUSE."** [#data-is-the-pollution-problem-of-the-information-age-and-protecting-privacy-is-the-environmental-challenge-almost-all-computers-produce-information-it-stays-around-festering-how-we-deal-with-it--how-we-contain-it-and-how-we-dispose-of-it--is-central-to-the-health-of-our-information-economy-just-as-we-look-back-today-at-the-early-decades-of-the-industrial-age-and-wonder-how-our-ancestors-could-have-ignored-pollution-in-their-rush-to-build-an-industrial-world-our-grandchildren-will-look-back-at-us-during-these-early-decades-of-the-information-age-and-judge-us-on-how-we-addressed-the-challenge-of-data-collection-and-misuse]
>
> **"WE SHOULD TRY TO MAKE THEM PROUD."**
#### 2.7 Legislation and self-regulation [#27-legislation-and-self-regulation]
**What GDPR demands, and the direct conflict:**
> **GDPR: personal data must be "COLLECTED FOR SPECIFIED, EXPLICIT AND LEGITIMATE PURPOSES AND NOT FURTHER PROCESSED IN A MANNER THAT IS INCOMPATIBLE WITH THOSE PURPOSES" and be "ADEQUATE, RELEVANT AND LIMITED TO WHAT IS NECESSARY."**
>
> ### **"However, THIS PRINCIPLE OF DATA MINIMIZATION RUNS DIRECTLY COUNTER TO THE PHILOSOPHY OF BIG DATA, WHICH IS TO MAXIMIZE DATA COLLECTION, TO COMBINE IT WITH OTHER DATASETS, AND TO EXPERIMENT AND EXPLORE IN ORDER TO GENERATE NEW INSIGHTS. EXPLORATION MEANS USING DATA FOR UNFORESEEN PURPOSES — WHICH THE GDPR STATES IS THE OPPOSITE OF THE 'SPECIFIED AND EXPLICIT' PURPOSES FOR WHICH THE DATA MUST HAVE BEEN COLLECTED."** [#however-this-principle-of-data-minimization-runs-directly-counter-to-the-philosophy-of-big-data-which-is-to-maximize-data-collection-to-combine-it-with-other-datasets-and-to-experiment-and-explore-in-order-to-generate-new-insights-exploration-means-using-data-for-unforeseen-purposes--which-the-gdpr-states-is-the-opposite-of-the-specified-and-explicit-purposes-for-which-the-data-must-have-been-collected]
>
> **"While this regulation has had SOME EFFECT ON THE ONLINE ADVERTISING INDUSTRY, IT HAS BEEN WEAKLY ENFORCED AND DOES NOT SEEM TO HAVE LED TO MUCH OF A CHANGE IN CULTURE AND PRACTICES ACROSS THE WIDER TECH INDUSTRY."**
**The counter-argument, given fairly:**
> **"Companies broadly OPPOSE REGULATION as being a burden and a hindrance to innovation. TO SOME EXTENT, THAT OPPOSITION IS JUSTIFIED. For example, sharing medical data creates clear risks to privacy BUT ALSO POTENTIAL OPPORTUNITIES: HOW MANY DEATHS COULD BE PREVENTED IF DATA ANALYSIS WERE ABLE TO HELP US ACHIEVE BETTER DIAGNOSTICS OR FIND BETTER TREATMENTS? OVERREGULATION MAY PREVENT SUCH BREAKTHROUGHS. IT IS DIFFICULT TO BALANCE THE POTENTIAL OPPORTUNITIES WITH THE RISKS."**
**The call to action:**
> ### **"FUNDAMENTALLY, WE NEED A CULTURE SHIFT IN THE TECH INDUSTRY WITH REGARD TO PERSONAL DATA. WE SHOULD STOP REGARDING USERS AS METRICS TO BE OPTIMIZED, AND REMEMBER THAT THEY ARE HUMANS WHO DESERVE RESPECT, DIGNITY, AND AGENCY."** [#fundamentally-we-need-a-culture-shift-in-the-tech-industry-with-regard-to-personal-data-we-should-stop-regarding-users-as-metrics-to-be-optimized-and-remember-that-they-are-humans-who-deserve-respect-dignity-and-agency]
>
> * **"We should SELF-REGULATE our data collection and processing practices in order to ESTABLISH AND MAINTAIN THE TRUST of the people who depend on our software."**
> * **"We should take it upon ourselves to EDUCATE END USERS ABOUT HOW THEIR DATA IS USED RATHER THAN KEEPING THEM IN THE DARK."**
> * **"We should allow each individual to MAINTAIN THEIR PRIVACY (i.e., their control over their own data) and NOT STEAL THAT CONTROL FROM THEM THROUGH SURVEILLANCE."**
>
> ### **"OUR INDIVIDUAL RIGHT TO CONTROL OUR DATA IS LIKE THE NATURAL ENVIRONMENT OF A NATIONAL PARK: IF WE DON'T EXPLICITLY PROTECT AND CARE FOR IT, IT WILL BE DESTROYED. IT WILL BE THE TRAGEDY OF THE COMMONS, AND WE WILL ALL BE WORSE OFF FOR IT. UBIQUITOUS SURVEILLANCE IS NOT INEVITABLE. WE ARE STILL ABLE TO STOP IT."** [#our-individual-right-to-control-our-data-is-like-the-natural-environment-of-a-national-park-if-we-dont-explicitly-protect-and-care-for-it-it-will-be-destroyed-it-will-be-the-tragedy-of-the-commons-and-we-will-all-be-worse-off-for-it-ubiquitous-surveillance-is-not-inevitable-we-are-still-able-to-stop-it]
>
> ### **"As a FIRST STEP, WE SHOULD NOT RETAIN DATA FOREVER, BUT PURGE IT AS SOON AS IT IS NO LONGER NEEDED, AND MINIMIZE WHAT WE COLLECT IN THE FIRST PLACE. DATA YOU DON'T HAVE IS DATA THAT CAN'T BE LEAKED, STOLEN, OR COMPELLED BY GOVERNMENTS TO BE HANDED OVER."** [#as-a-first-step-we-should-not-retain-data-forever-but-purge-it-as-soon-as-it-is-no-longer-needed-and-minimize-what-we-collect-in-the-first-place-data-you-dont-have-is-data-that-cant-be-leaked-stolen-or-compelled-by-governments-to-be-handed-over]
>
> ### **"AS PEOPLE WORKING IN TECHNOLOGY, IF WE DON'T CONSIDER THE SOCIETAL IMPACT OF OUR WORK, WE'RE NOT DOING OUR JOB."** [#as-people-working-in-technology-if-we-dont-consider-the-societal-impact-of-our-work-were-not-doing-our-job]
***
# 14.5 Self-test (/docs/ddia/doing-right-thing/self-test)
Why is "a technology is not good or bad in itself" *not* a defence for engineers?
What does the book say ethics is, and what does it say ethics is *not*?
Why are organizations structurally biased toward "no" in automated decisions, and why is that rational for them but harmful in aggregate?
Define "algorithmic prison." Contrast it with the presumption of innocence.
Give the fair case *for* algorithmic decision-making over human judgment.
Explain the mechanism by which a model amplifies bias. Why doesn't excluding protected attributes fix it?
What does "machine learning is like money laundering for bias" mean?
Why does the book say moral imagination is required, and that only humans can provide it?
Give three accountability questions raised by automated decisions.
Contrast a traditional credit score with ML-based scoring on: the question answered, the inputs, and the possibility of correction.
Why does a correct probability distribution not justify a decision about an individual? Give the life-expectancy example.
Give the dual-use example: the same analytics capability used for good and for predation.
Trace the credit-score/employment feedback loop step by step.
Why is the German gas station example unsettling? What did nobody program?
What is systems thinking, and what is the key question it asks of a data system?
Describe the four-stage drift from "service for the user" to "surveillance."
Perform the data→surveillance substitution on a sentence from your own workplace.
Why does the book say the surveillance infrastructure exceeds what totalitarian regimes could achieve — and what is the one difference?
Give two reasons people accept corporate surveillance, and explain why each is a position of privilege.
Give three examples where surveillance is already used for consequential decisions, not just recommendations.
Give the four reasons "the user consented" fails as a defence. Which one concerns *non-users*?
What does GDPR require of consent, word for word on the key phrases? Name two other lawful bases.
Define privacy as a *decision right*. Why is "privacy is dead" a misunderstanding?
Explain the transfer of privacy rights to the data collector. Why do companies keep the outcome secret?
Why is "you can't be reidentified from that ad bucket" not a sufficient answer?
What is the limitation of user-facing privacy settings?
Contrast "data exhaust" with "data as labor." Which does the book endorse, and why?
Give five ways data can end up somewhere you didn't intend, even with good intentions.
Why "the new uranium" rather than "the new oil"?
What does Schneier's "poor civic hygiene" line mean for a data-retention decision you make today?
Draw the Industrial Revolution parallel: what were the harms, what were the safeguards, and what did they cost?
State Schneier's pollution framing. What is the environmental challenge of the information age?
Why does GDPR's purpose limitation directly conflict with the philosophy of big data? Has enforcement changed the culture?
Give the strongest argument *against* data protection regulation, as the book states it.
What is the single concrete first step the book recommends, and what is the one-line justification for it?
Name three of your own technical decisions from earlier chapters that carry an ethical dimension, and say what it is.
**Reflective question (no right answer):** take a system you've actually built or worked on. Who is its customer — the user, or someone else? What data does it collect that isn't needed to deliver the service to the user? Who is harmed if that data leaks, and who is harmed if it doesn't exist? What would you change if you started again, and what stops you changing it now?
# 14.6 Terminology introduced here (/docs/ddia/doing-right-thing/terminology-introduced-here)
# 14.3 Where this chapter collides with the rest of the book (/docs/ddia/doing-right-thing/where-chapter-collides-rest)
**This is the chapter's real value: it retroactively reframes technical decisions you already made.**
| Technical decision | The ethical dimension it carries |
| ---------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Ch 1 — "we store data because we think its value exceeds the cost"** | **The cost includes breach liability, compliance fines, and the SAFETY RISK TO USERS when data reveals criminalized behavior** (abortion travel, sexuality). **Data minimization is a design constraint, not a compliance chore** |
| **Ch 3 — event sourcing keeps every event forever** | **Directly conflicts with GDPR erasure. Crypto-shredding moves the problem, doesn't solve it** |
| **Ch 4 — vector embeddings and semantic search** | **The model learns whatever bias is in the corpus** — §1.2 |
| **Ch 5 — data outlives code** | **So does personal data. The 5-year-old row is still about a real person** |
| **Ch 6 — derived data can be recreated from the source** | **Which means a deletion must propagate to EVERY derived copy**, and you must know where they all are |
| **Ch 7 — shard per tenant** | **Turns GDPR export and deletion into "operations on their shard"** — one of the seven advantages listed there, and the most underrated |
| **Ch 11 — batch inference at scale** | **Where §1.1's "algorithmic prison" is manufactured, millions of decisions at a time** |
| **Ch 12 — immutability is a virtue** | **§2.6's "it stays around, festering." Immutability and the right to be forgotten are in direct tension** |
| **Ch 13 — auditability and provenance** | **The same machinery that lets you debug a pipeline lets you EXPLAIN A DECISION TO A JUDGE** (§1.3) and prove you deleted what you said you deleted |
***
# 5.5 Avro (/docs/ddia/encoding-evolution/avro)
Started **2009 as a Hadoop subproject, because Protocol Buffers was not a good fit for Hadoop's use cases.**
**Two schema languages:** **Avro IDL** (for human editing) and a **JSON-based** one (machine-readable). Like protobuf, they specify only fields and types — **no complex validation rules.**
```txt
record Person {
string userName;
union { null, long } favoriteNumber = null;
array interests;
}
```
```json
{ "type": "record", "name": "Person",
"fields": [
{"name": "userName", "type": "string"},
{"name": "favoriteNumber", "type": ["null", "long"], "default": null},
{"name": "interests", "type": {"type": "array", "items": "string"}}
] }
```
#### 5.1 The encoding — 32 bytes, and NOTHING self-describing [#51-the-encoding--32-bytes-and-nothing-self-describing]
**Notice: the schema has NO TAG NUMBERS.**
> **The encoding is simply VALUES CONCATENATED TOGETHER.** To parse it, **you go through the fields in the order they appear in the schema and use the schema to determine each field's datatype.**
>
> **⇒ The binary data can be decoded correctly ONLY IF the reading code uses the EXACT SAME SCHEMA as the writing code. Any mismatch means incorrectly decoded data.**
So how does Avro evolve at all? **This is the clever part.**
#### 5.2 Writer's schema and reader's schema [#52-writers-schema-and-readers-schema]
**Schema resolution — how differences are reconciled:**
#### 5.3 Evolution rules [#53-evolution-rules]
| | Meaning in Avro |
| -------------------------- | ------------------------------------------------------ |
| **Forward compatibility** | **The WRITER can use a NEWER schema than the reader** |
| **Backward compatibility** | **The WRITER can use an OLDER schema than the reader** |
> **THE RULE: you may add or remove only a field THAT HAS A DEFAULT VALUE.**
* Add a field **with** a default → a reader on the new schema reading old data **fills in the default.** ✔
* Add a field **without** a default → **new readers can't read data written by old writers → breaks BACKWARD compatibility.** ✗
* Remove a field **without** a default → **old readers can't read data written by new writers → breaks FORWARD compatibility.** ✗
**Nulls are explicit.** In some languages `null` is an acceptable default for any variable — **not in Avro.** To allow null you must use a **union type**: `union { null, long, string } field;`. **You can use `null` as a default only if it is the FIRST branch of the union.** More verbose than nullable-by-default, but **it helps prevent bugs by being explicit about what can and cannot be null.**
**Two asymmetric cases worth remembering:**
| Change | Backward | Forward |
| ---------------------------------------------------------------------------------------------- | -------- | -------- |
| **Change a field name** (reader's schema declares an **alias** matching the old writer's name) | ✔ | ✗ |
| **Add a branch to a union type** | ✔ | ✗ |
| **Change a datatype** (where Avro can convert) | possible | possible |
#### 5.4 But how does the reader know the writer's schema? [#54-but-how-does-the-reader-know-the-writers-schema]
**You can't include the whole schema with every record** — it would likely be much bigger than the encoded data, **negating all the space savings.** Three answers depending on context:
| Context | Mechanism |
| -------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Large file with lots of records** (millions, all one schema) | **Include the schema once at the beginning of the file.** Avro specifies a file format for this: **object container files** |
| **Database with individually written records** (different records written at different times with different schemas) | **Include a version number at the start of every encoded record** and keep a list of schema versions. Reader extracts the version, fetches the corresponding writer's schema, decodes. **This is how Confluent's Schema Registry for Kafka and LinkedIn's Espresso work** |
| **Records over a network connection** | **Negotiate the schema version at connection setup**, use it for the connection's lifetime. **The Avro RPC protocol works like this** |
> **A database of schema versions is useful in any case: it acts as documentation and gives you a chance to CHECK SCHEMA COMPATIBILITY.** Version number = a simple incrementing integer or **a hash of the schema.**
#### 5.5 Why "no tag numbers" matters: dynamically generated schemas [#55-why-no-tag-numbers-matters-dynamically-generated-schemas]
**The scenario:** dump a relational database's contents to a file in a binary format.
***
# 5.10 Decision cheat sheet (/docs/ddia/encoding-evolution/decision-cheat-sheet)
**Which encoding?**
| Situation | Choose |
| ----------------------------------------------------------------------- | ---------------------------------- |
| Public API, cross-organization, must be debuggable | **JSON + OpenAPI** |
| Internal service-to-service, high volume, typed languages | **Protobuf + gRPC** |
| Streaming/analytics, many records per file, schemas generated from data | **Avro + schema registry** |
| Columnar analytics files | **Parquet** |
| Anything at all | **Never a language-native format** |
**Is this change safe?**
Ask both directions separately. *Can new code read old data?* (backward) *Can old code read new data?* (forward) A change is only safe if the answer to **both** is yes for the duration of your rolling upgrade window — which, for public APIs, may be **indefinite**.
**Protobuf or Avro?**
**Protobuf** when the reader shouldn't need the writer's schema (RPC, self-contained messages) and schemas are hand-authored. **Avro** when schemas are **dynamically generated** or there's one schema per file/topic, and you're willing to run a registry.
**RPC or messaging?**
RPC when the caller **needs the answer now** and the callee's availability is acceptable. Messaging when you want **buffering, redelivery, fan-out, and decoupling** — and can tolerate asynchrony.
**How do I version a public API?**
Pick one mechanism (URL path is the least surprising), **instrument usage per version and per client from day one**, publish a deprecation policy, and expect to run old versions **far longer than you planned**.
**Do I need durable execution?**
Only when a multi-step process spans services or third parties **and** partial completion is unacceptable. Otherwise a retry with an idempotency key is simpler and less brittle. If you adopt it, commit to determinism and to versioning-by-deployment.
***
# 5.1 Encoding and decoding (/docs/ddia/encoding-evolution/encoding-decoding)
Programs work with data in **(at least) two representations:**
> ⚠️ **Terminology clash:** *serialization* also means something completely different in transactions (Ch 8). **The book uses "encoding" throughout to avoid overloading the word.**
**When you don't need it:** when a database **operates directly on compressed data loaded from disk** (Ch 4 §7.6), and with **zero-copy formats designed to be used both at runtime and on disk/network without an explicit conversion step — Cap'n Proto and FlatBuffers.**
***
# 5.14 Forward links (/docs/ddia/encoding-evolution/forward-links)
| Concept here | Where it's developed |
| ------------------------------------------------ | ------------------------------------------- |
| Why you can't know whether a request got through | **Ch 9** — Trouble with Distributed Systems |
| Idempotence and exactly-once processing | **Ch 12** — Stream Processing |
| Distributed transactions vs. workflows | **Ch 8** — Transactions |
| Coordination services (etcd, ZooKeeper) | **Ch 10** — Consistency and Consensus |
| Message brokers compared in depth | **Ch 12** — Stream Processing |
| Avro object container files in batch jobs | **Ch 11** — Batch Processing |
| Determinism as a system property | **Ch 9** |
| Event sourcing needing indefinite retention | **Ch 3** §3, **Ch 12** |
# 5. Encoding and Evolution (/docs/ddia/encoding-evolution)
> "Everything changes and nothing stands still." — Heraclitus
**The premise:** applications change. A change to features usually requires a change to the data they store. **But in a large application, code changes cannot happen instantaneously:**
* **Server-side:** you want a **rolling upgrade** (staged rollout) — deploy to a few nodes at a time, monitor, work through all nodes. **No downtime → more frequent releases → better evolvability.**
* **Client-side:** **you're at the mercy of the user**, who may not install the update for a long time.
**Therefore old and new versions of code, and old and new data formats, coexist in the system at the same time.**
***
# 5.3 JSON, XML, CSV, and binary variants (/docs/ddia/encoding-evolution/json-xml-csv-binary)
#### 3.1 The honest list of flaws [#31-the-honest-list-of-flaws]
| Format | Problems |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **XML** | **Too verbose and unnecessarily complicated** |
| **XML & CSV** | **Cannot distinguish a number from a string that happens to consist of digits** (except via an external schema) |
| **JSON** | Distinguishes strings and numbers, **but not integers from floating-point, and doesn't specify precision** |
| **CSV** | **No schema at all** — the application defines what each row and column means; adding a row or column must be handled manually. **Also quite vague** (what if a value contains a comma or newline?). Escaping rules are formally specified but **not all parsers implement them correctly** |
| **JSON & XML** | Good Unicode string support, **but no binary strings.** Workaround: **Base64**, indicated by the schema — **hacky, and increases data size by about a third** |
| **XML Schema & JSON Schema** | **Powerful and thus quite complicated to learn and implement.** Since correct interpretation depends on schema info, **applications not using schemas must hardcode encoding/decoding logic** |
#### 3.2 The 2⁵³ problem — worth knowing exactly [#32-the-2-problem--worth-knowing-exactly]
> **Integers greater than 2⁵³ cannot be exactly represented in an IEEE 754 double-precision float**, so they become inaccurate when parsed by a language that uses floating point — **like JavaScript.**
**The canonical real-world instance:** X uses a **64-bit number to identify each post.** The JSON returned by the API **includes post IDs TWICE — once as a JSON number and once as a decimal string** — to work around incorrect parsing by JavaScript applications.
**Practical rule:** any ID that could exceed 2⁵³ (Snowflake IDs, bigint PKs, nanosecond timestamps) **must be transported as a string in JSON.**
#### 3.3 JSON Schema [#33-json-schema]
Widely adopted: in **OpenAPI** specs, in **schema registries** (Confluent Schema Registry, Red Hat Apicurio), and in databases (**PostgreSQL's `pg_jsonschema`**, **MongoDB's `$jsonSchema` validator**).
Offers standard primitives (`string`, `number`, `integer`, `object`, `array`, `boolean`, `null`) **plus a separate validation specification** overlaying constraints — e.g. a `port` field with minimum 1 and maximum 65,535.
**Open vs closed content models:**
| Model | `additionalProperties` | Meaning |
| -------------------------- | ---------------------- | -------------------------------------------------------------------- |
| **Open** (**the default**) | `true` | **Any field not defined in the schema may exist, with any datatype** |
| **Closed** | `false` | **Only explicitly defined fields are allowed** |
> **Consequence: JSON Schemas are usually a definition of what ISN'T permitted (invalid values on defined fields) rather than what IS permitted.**
**Example — a map from integer-like keys to strings**, which JSON can't express natively (JSON objects always use string keys):
```json
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"patternProperties": { "^[0-9]+$": { "type": "string" } },
"additionalProperties": false
}
```
JSON Schema also supports **conditional if/else logic, named types, references to remote schemas**, and more. **All of which makes for a very powerful schema language — and unwieldy definitions.** It can be challenging to **resolve remote schemas, reason about conditional rules, or evolve schemas compatibly.** Same concerns apply to XML Schema.
#### 3.4 Binary variants of JSON/XML [#34-binary-variants-of-jsonxml]
A **profusion**: MessagePack, CBOR, BSON, BJSON, UBJSON, BISON, Hessian, Smile (JSON); WBXML, Fast Infoset (XML). **Adopted in various niches — more compact, sometimes faster to parse — but none as widely adopted as the textual versions.**
Some extend the datatype set (integers vs floats, binary strings) **but keep the JSON/XML data model unchanged.** Crucially:
> **Since they don't prescribe a schema, they must include ALL the object field names within the encoded data.**
**The running example record:**
```json
{ "userName": "Martin", "favoriteNumber": 1337, "interests": ["daydreaming", "hacking"] }
```
**MessagePack encoding, byte by byte:**
*(If an object has more than 15 fields — too many for four bits — it gets a different type indicator and the count is encoded in two or four bytes.)*
**The verdict:**
> **66 vs 81 bytes is only a little less. It's not clear whether such a small space reduction (and perhaps a parsing speedup) is worth the loss of human-readability.** The schema-driven formats do **twice as well** — because they drop the field names entirely.
***
# 5.2 Language-specific formats — and why not to use them (/docs/ddia/encoding-evolution/language-specific-formats-not)
`java.io.Serializable`, Python `pickle`, Ruby `Marshal`, Kryo for Java.
Convenient — in-memory objects saved and restored with minimal code. **But four deep problems:**
| Problem | Consequence |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Tied to one programming language** | Reading in another language is difficult. **You commit to your current language for potentially a long time and preclude integrating with other organizations' systems** |
| **Decoding must instantiate arbitrary classes** | **A frequent source of security problems.** If an attacker can get your application to decode an arbitrary byte sequence, **they can instantiate arbitrary classes**, which often allows **remote code execution** |
| **Versioning is an afterthought** | Built for quick-and-easy encoding; **they neglect forward and backward compatibility** |
| **Efficiency is an afterthought** | **Java's built-in serialization is notorious for bad performance and bloated encoding** |
> **Generally a bad idea for anything other than very transient purposes.**
*(In practice: "deserialization of untrusted data" is a top-10 vulnerability class — Java gadget chains, Python `pickle.loads`, PHP `unserialize`, Ruby YAML. If you take one operational rule from this section: **never deserialize a language-native format from a source you don't control.**)*
***
# 5.6 The merits of schemas (/docs/ddia/encoding-evolution/merits-schemas)
Protocol Buffers' and Avro's schema languages are **much simpler than XML Schema or JSON Schema**, which support detailed validation rules ("must match this regex", "must be between 0 and 100"). **Being simpler to implement and use, they've gained support across a wide range of languages.**
**These ideas are not new.** **ASN.1** — a schema definition language **first standardized in 1984** — used to define network protocols; **its binary encoding (DER) is still used to encode SSL certificates (X.509).** It **supports schema evolution using tag numbers, similar to Protocol Buffers.** But it's **very complex and badly documented — probably not a good choice for new applications.**
Also worth noting: **most relational databases have a proprietary binary network protocol**, with a vendor-supplied **driver (ODBC/JDBC)** decoding responses into in-memory data structures.
**Four properties of schema-driven binary encodings:**
1. **Much more compact than "binary JSON" variants**, since they **omit field names from the encoded data**
2. **The schema is valuable documentation — and because the schema is REQUIRED for decoding, you can be sure it is up to date** (manually maintained documentation easily diverges from reality)
3. **A database of schemas lets you check forward and backward compatibility of changes BEFORE anything is deployed**
4. **Code generation from the schema enables compile-time type checking** for statically typed languages
> **Schema evolution gives the same flexibility as schemaless/schema-on-read JSON databases, while also providing better guarantees about your data and better tooling.**
>
> **Still, keep the number of concurrent schema formats to a minimum to keep operations simple.**
***
# 5.7 Modes of Dataflow (/docs/ddia/encoding-evolution/modes-dataflow)
**Compatibility is a relationship between one process that encodes data and another that decodes it.** Who encodes, who decodes? Four modes:
#### 7.1 Dataflow through databases [#71-dataflow-through-databases]
**Backward compatibility is clearly necessary** — otherwise your future self can't decode what you previously wrote.
**But forward compatibility is ALSO often required**, because several processes access the database at once: different applications/services, or **multiple instances of the same service**, and during a rolling upgrade **some instances run new code and some run old.** So **a value written by newer code may be read by older code still running.**
**"Data outlives code."**
> A database allows any value to be updated at any time — **you may have values written five milliseconds ago and others written five years ago.** When you deploy a new version of a server-side application, you may **entirely replace the old version within a few minutes.** **The same is NOT true of database contents: the five-year-old data is still there, in the original encoding, unless you have explicitly rewritten it.**
**Migration is possible but expensive on a large dataset, so most databases defer it — asynchronously, best-effort:**
| System | How it defers |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **LSM-tree storage engines** | **Rewrite data using the latest format during COMPACTION** (Ch 4 §2.3) |
| **Most relational databases** | **Simple schema changes (add a column with a null default) without rewriting existing data.** When an old row is read, **the database fills in nulls for columns missing from the encoded data on disk** |
> **Schema evolution thus allows the entire database to APPEAR as if it were encoded with a single schema, even though the underlying storage contains records encoded with various historical schema versions.**
**More complex changes still require rewriting**, often at the application level — **changing a single-valued attribute to multivalued, or moving data into a separate table.** **Maintaining forward and backward compatibility across such migrations remains a research problem.**
**Archival storage.** A periodic snapshot (for backup or warehouse loading) **is typically encoded using the LATEST schema, even if the source contained a mixture of eras — since you're copying anyway, you might as well encode consistently.** Because the dump is **written in one go and thereafter immutable**, **Avro object container files are a good fit**, and it's a good opportunity to write an **analytics-friendly column-oriented format like Parquet.**
#### 7.2 Dataflow through services: REST and RPC [#72-dataflow-through-services-rest-and-rpc]
Clients connect to servers; **the API exposed by the server is a service.**
**The web works this way** — browsers GET HTML/CSS/JS/images and POST data; the API is a standardized set of protocols and formats (HTTP, URLs, SSL/TLS, HTML). **Because browsers, servers, and site authors mostly agree on these standards, you can use any browser to access any website (at least in theory!).**
But browsers aren't the only client — **native mobile/desktop apps and client-side JavaScript** also make HTTP requests, where **the response is not HTML for a human but data convenient for further processing (most often JSON)**, and **the API on top of HTTP is application-specific.**
**Services vs databases:**
> Both allow clients to submit and query data. **But databases allow ARBITRARY queries using query languages; services expose an APPLICATION-SPECIFIC API allowing only inputs and outputs predetermined by the business logic.** This restriction provides **encapsulation** — services can impose **fine-grained restrictions on what clients can and cannot do.**
**Why compatibility matters here:** the design goal of microservices is **independently deployable and evolvable services**, each **owned by one team able to release frequently without coordinating with other teams.** **So we must expect old and new versions of servers and clients running at the same time.**
> **As long as APIs remain compatible, teams are free to modify their systems any way they'd like — this property makes internal migrations of data, services, or even entire systems much easier.**
**Three contexts for web services:**
1. A client app on a user's device → a service over HTTP, **typically over the public internet**
2. **One service → another service in the same organization**, often within the same private network
3. **One service → a service owned by a DIFFERENT organization**, usually via the internet — **data exchange between organizations' backends**, e.g. credit card processing, or **OAuth for shared access to user data**
**REST** is the most popular design philosophy — building on HTTP's principles: **simple data formats, URLs for identifying resources, and HTTP features for cache control, authentication, and content type negotiation.**
**IDLs for services:** clients must know which endpoint to query and what data format to send/expect. **The two most popular service IDLs: OpenAPI (Swagger)** for JSON web services, and **Protocol Buffers** for gRPC.
```yaml
openapi: 3.0.0
info: { title: "Ping, Pong", version: 1.0.0 }
servers: [ { url: http://localhost:8080 } ]
paths:
/ping:
get:
summary: Given a ping, returns a pong message
responses:
'200':
description: A pong
content:
application/json:
schema:
type: object
properties:
message: { type: string, example: "Pong!" }
```
```python
from fastapi import FastAPI
from pydantic import BaseModel
app = FastAPI(title="Ping, Pong", version="1.0.0")
class PongResponse(BaseModel):
message: str = "Pong!"
@app.get("/ping", response_model=PongResponse,
summary="Given a ping, returns a pong message")
async def ping():
return PongResponse()
```
**Two directions of coupling between definition and code:**
**Service frameworks** (Spring Boot, FastAPI, gRPC) let developers focus on business logic **while the framework handles routing, metrics, caching, authentication.**
#### 7.3 The problems with RPC — the classic critique [#73-the-problems-with-rpc--the-classic-critique]
**A long line of predecessors, all with serious problems:**
| Technology | Problem |
| ----------------- | ---------------------------------------------------------------------------------------------- |
| **EJB, Java RMI** | **Limited to Java** |
| **DCOM** | **Limited to Microsoft platforms** |
| **CORBA** | **Excessively complex; does not provide backward or forward compatibility** |
| **SOAP / WS-\*** | Aims at cross-vendor interoperability but **plagued by complexity and compatibility problems** |
**All are based on RPC (introduced in the 1970s), which tries to make a remote network service call look the same as a local function call — "location transparency."**
> **Although this seems convenient at first, the approach is FUNDAMENTALLY FLAWED.** A network request is very different from a local function call:
| # | Local function call | Network request |
| - | ----------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 | **Predictable** — succeeds or fails based on parameters **under your control** | **Unpredictable**, for reasons **entirely outside your control**: request/response lost, remote machine slow or unavailable. **Network problems are common, so applications must anticipate them (e.g. retries)** |
| 2 | Returns a result, throws an exception, or **never returns** (infinite loop / crash) | **A third outcome: returns WITHOUT A RESULT because of a TIMEOUT. You simply don't know what happened** — no way of knowing whether the request got through |
| 3 | No such problem | **Retrying a "failed" request may perform the action MULTIPLE TIMES**, if the request got through and only the response was lost — **unless you build deduplication (idempotence) into the protocol** |
| 4 | **Takes about the same time each call** | **Much slower, and latency is wildly variable** — \<1 ms at good times, **many seconds** when the network is congested or the service overloaded, **for exactly the same work** |
| 5 | Can efficiently pass **references (pointers) to local memory** | **All parameters must be encoded into bytes.** Fine for immutable primitives; **quickly problematic with larger amounts of data and mutable objects** |
| 6 | Single process, single language — no type translation | **Client and service may be in different languages**, so the framework must translate datatypes. **This can end up ugly** — recall **JavaScript's 2⁵³ problem** |
> **There's no point trying to make a remote service look too much like a local object, because it's a FUNDAMENTALLY DIFFERENT THING. Part of the appeal of REST is that it treats state transfer over a network as a process DISTINCT from a function call.**
#### 7.4 Load balancing, service discovery, and service meshes [#74-load-balancing-service-discovery-and-service-meshes]
**The problem: a client must know the address of the service it's connecting to — service discovery.** The simplest approach (hardcode IP and port) works **until the server goes offline, moves, or becomes overloaded — then the client must be manually reconfigured.**
**Load balancing** = spreading requests across the multiple instances usually running for availability and scalability.
| Solution | How it works | Trade-off |
| ----------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Hardware load balancers** | Specialized equipment in datacenters. Clients connect to a single host:port; connections routed to one of the servers. **Detects network failures to a downstream server and shifts traffic** | Requires an appliance |
| **Software load balancers** (NGINX, HAProxy) | Same behavior, **as applications on a standard machine** | — |
| **DNS** | Multiple IP addresses per domain name; **the client's network layer picks which to use** | **DNS is designed to propagate changes over LONGER periods and to CACHE entries.** If servers start/stop/move frequently, **clients might see stale IPs with no server running on them** |
| **Service discovery systems** (etcd, ZooKeeper) | A **centralized registry**. A new instance **registers itself** (host, port) with **metadata: shard ownership, datacenter location**, and **periodically sends a HEARTBEAT** to signal it's still available. The client queries the registry for endpoints, then **connects directly** | **Supports a much more dynamic environment than DNS**, and the extra metadata **enables smarter load-balancing decisions** by clients |
| **Service meshes** (Istio, Linkerd) | **Combines software load balancing and service discovery.** Deployed as an **in-process client library or a "sidecar" container on BOTH client and server.** The client app connects to its **own local** load balancer, which connects to the **server's** load balancer, which routes to the local server process | **Complicated** — but: because everything is routed through **local** connections, **connection encryption is handled entirely at the load-balancer level, shielding clients and servers from SSL certificates and TLS.** Also **sophisticated observability**: which services call each other in real time, failure detection, traffic load |
**Choosing:** very dynamic environments with **Kubernetes** → often a **service mesh**. **Specialized infrastructure (databases, messaging) might require purpose-built load balancers.** **Simpler deployments are best served with software load balancers.**
#### 7.5 Encoding and evolution for RPC [#75-encoding-and-evolution-for-rpc]
> **A simplifying assumption vs databases: it is reasonable to assume ALL SERVERS WILL BE UPDATED FIRST and all clients second. Thus you need BACKWARD compatibility only on REQUESTS, and FORWARD compatibility on RESPONSES.**
**Compatibility properties are inherited from the encoding:**
* **gRPC (Protocol Buffers) and Avro RPC** evolve per their format's rules
* **RESTful APIs** most commonly use JSON responses and JSON/URI-encoded/form-encoded requests. **Adding optional request parameters and adding new fields to response objects are usually considered compatible changes**
**The hard part is organizational:**
> **RPC is often used across ORGANIZATIONAL BOUNDARIES, so the service provider often has NO CONTROL over its clients and cannot force them to upgrade. Compatibility needs to be maintained for a long time, PERHAPS INDEFINITELY.** If a breaking change is required, **the provider frequently ends up maintaining multiple versions of the API side by side.**
**There is no agreement on how API versioning should work.** Common approaches for REST:
* **Version number in the URL** (`/v2/users`)
* **Version in the HTTP `Accept` header**
* For services using **API keys**: **store the client's requested version on the server**, updatable through a **separate administrative interface**
#### 7.6 Durable execution and workflows [#76-durable-execution-and-workflows]
**The setup:** a payment processor charges a credit card and deposits funds into a bank account. Services: fraud detection, credit card integration, bank integration.
**Workflow definitions** may be written in a general-purpose language, a **DSL**, or a markup language such as **BPEL**; graphical notations like **BPMN** exist.
**Three families of workflow engine:**
| Family | Examples | Purpose |
| --------------------- | ----------------------------- | ----------------------------------------------------------------------- |
| Data orchestration | **Airflow, Dagster, Prefect** | Integrate with data systems, orchestrate **ETL** tasks |
| Graphical / business | **Camunda, Orkes** | **BPMN** notation so **non-engineers** can define and execute workflows |
| **Durable execution** | **Temporal, Restate** | Exactly-once semantics for workflows |
**Why durable execution exists:**
> We want to process each payment **exactly once**. A failure mid-workflow could result in **a credit card charge but no corresponding bank deposit.** In a service-based architecture **you can't simply wrap the two tasks in a database transaction.** Moreover, **you might be interacting with third-party payment gateways you have limited control over.**
**How it works:**
> **If a task fails, the framework re-executes it — but SKIPS any RPC calls or state changes that succeeded before the failure. It will PRETEND to make the call, but instead RETURN THE RESULTS FROM THE PREVIOUS CALL.**
>
> **This is possible because durable execution frameworks LOG ALL RPCs AND STATE CHANGES TO DURABLE STORAGE, like a write-ahead log.**
```python
@workflow.defn
class PaymentWorkflow:
@workflow.run
async def run(self, payment: PaymentRequest) -> PaymentResult:
is_fraud = await workflow.execute_activity(
check_fraud, payment, start_to_close_timeout=timedelta(seconds=15))
if is_fraud:
return PaymentResultFraudulent
credit_card_response = await workflow.execute_activity(
debit_credit_card, payment, start_to_close_timeout=timedelta(seconds=15))
# ...
```
**Three real challenges — all consequences of "replay must be deterministic":**
1. **External services must STILL provide an idempotent API.** Developers must remember to **use unique IDs for these APIs to prevent duplicate execution.**
2. **Code changes are brittle.** Because the framework **logs each RPC call in order, it expects subsequent executions to make the same RPC calls in the same order.** **You might introduce undefined behavior simply by REORDERING FUNCTION CALLS.**
> **Instead of modifying an existing workflow's code, it is safer to deploy a NEW VERSION separately, so re-executions of existing invocations continue using the old version and only new invocations use the new code.**
3. **Nondeterministic code is problematic** — **random number generators, system clocks.** Frameworks **provide their own deterministic implementations, but you have to remember to use them.** Some offer static analysis (**Temporal's Workflow Check**) to detect introduced nondeterminism.
> **Making code deterministic is a powerful idea but tricky to do robustly.** (Returns in Ch 9.)
*Note how this is the same constraint as event-sourcing projections in Ch 3 §3.3 — replayability demands determinism, everywhere it appears.*
#### 7.7 Event-driven architectures [#77-event-driven-architectures]
**A request is called an event or message.** Two structural differences from RPC:
1. **The sender usually does not wait for the recipient to process the event**
2. **Events are typically not sent via a direct network connection but via an intermediary — a MESSAGE BROKER** (event broker, message queue, message-oriented middleware) **which stores the message temporarily**
**Five advantages over direct RPC:**
| Advantage | Why |
| ------------------------------- | ------------------------------------------------------------------------------------------------- |
| **Buffering** | Acts as a buffer **if the recipient is unavailable or overloaded — improving system reliability** |
| **Redelivery** | **Automatically redelivers messages to a process that has crashed, preventing message loss** |
| **No service discovery needed** | Senders **don't need to connect directly to the recipient's IP address** |
| **Fan-out** | **The same message can be sent to several recipients** |
| **Logical decoupling** | **The sender just publishes and doesn't care who consumes** |
**Communication is asynchronous** — the sender sends and forgets. **You can implement a synchronous RPC-like model by having the sender wait for a response on a separate channel.**
**The landscape:** formerly commercial enterprise software (**TIBCO, IBM WebSphere, webMethods**), then open source (**RabbitMQ, ActiveMQ, HornetQ, NATS, Redpanda, Apache Kafka**), more recently cloud services (**Amazon Kinesis, Azure Service Bus, Google Cloud Pub/Sub**).
**Two distribution patterns:**
**Encoding:** brokers **typically don't enforce any data model — a message is just bytes with metadata**, so any encoding works. **Common approach: Protocol Buffers, Avro, or JSON, with a SCHEMA REGISTRY deployed alongside the broker to store valid schema versions and check compatibility.** **AsyncAPI** is the messaging equivalent of OpenAPI.
**Durability varies.** Many write messages to disk so they survive a broker crash/restart. **Unlike databases, many brokers automatically DELETE messages after consumption.** **Some can be configured to store messages indefinitely — which you'd require for event sourcing** (Ch 3 §3).
> ⚠️ **If a consumer REPUBLISHES messages to another topic, be careful to PRESERVE UNKNOWN FIELDS** — otherwise you reproduce the data-loss problem of §0 in the message pipeline.
#### 7.8 Distributed actor frameworks [#78-distributed-actor-frameworks]
**The actor model** is a concurrency model for a single process. **Rather than dealing directly with threads (and race conditions, locking, deadlock), logic is encapsulated in ACTORS.** Each actor typically represents one client or entity, may have **local state not shared with any other actor**, and communicates by **sending and receiving asynchronous messages.**
> **Message delivery is NOT GUARANTEED — in certain error scenarios, messages will be lost.** Since **each actor processes only one message at a time**, it needn't worry about threads, and **each actor can be scheduled independently.**
**Distributed actor frameworks** — **Akka, Orleans, Erlang/OTP** — use this model to **scale across multiple nodes**. **The same message-passing mechanism is used whether sender and recipient are on the same node or different nodes**; across nodes the message is transparently encoded, sent, and decoded.
> **Location transparency works BETTER in the actor model than in RPC, because the actor model ALREADY ASSUMES messages may be lost, even within a single process.** Although network latency is higher than in-process, **there is less of a fundamental mismatch between local and remote communication.**
*(This is the sharp counterpoint to §7.3: location transparency isn't inherently wrong — it's wrong when the local model promises more reliability than the network can deliver. Actors work because they promise less.)*
**But:** a distributed actor framework is essentially **a message broker + the actor programming model in one framework** — **and if you want rolling upgrades, you still have to worry about forward and backward compatibility**, since messages may flow from a new-version node to an old-version node and vice versa.
***
# 5.9 Production failure catalog for this chapter (/docs/ddia/encoding-evolution/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ------------------------------------------------------ | ------------------------------------------------------------------------ |
| A field silently disappears after an edit | **Unknown-field loss** during a rolling upgrade (§0) |
| Large IDs are wrong by a few digits in the browser | **2⁵³** — JSON number parsed as an IEEE 754 double |
| RCE from a "harmless" data endpoint | Deserializing a **language-native format** from untrusted input |
| Decoding produces garbage after a schema change | **Reused Protobuf tag number** |
| Big values truncated after widening a type | Old code still holding a **32-bit variable** |
| "Set to zero" indistinguishable from "not set" | **proto3 default-value ambiguity** |
| Every consumer breaks at once after a producer deploy | Added an Avro field **without a default** |
| Nothing can decode after a registry outage | **Schema registry** is on every consumer's critical path and unbacked-up |
| Old readers break after a field rename | Rename is **backward but not forward** compatible in Avro |
| gRPC traffic all lands on one backend | **HTTP/2 multiplexing** vs an L4 load balancer |
| Server burns CPU on abandoned requests | **Deadline not propagated** |
| Duplicate side effects after a retry | Timeout gave no information; **no idempotence in the protocol** |
| Deprecated API version can never be removed | **No per-version usage instrumentation**; no control over clients |
| Fields dropped as messages pass through a pipeline | Consumer **republished** after decoding into an old model |
| Workflow throws `NonDeterministicError` after a deploy | **Reordered or added activity calls**, breaking replay |
| Payment charged twice | Third party **not idempotent**, no idempotency key |
| Clients see stale IPs after a failover | **DNS caching** used for a dynamic service population |
***
# 5.4 Protocol Buffers (/docs/ddia/encoding-evolution/protocol-buffers)
Binary encoding library from Google; **similar to Apache Thrift** (originally Facebook) — most of this section applies to Thrift too.
**Requires a schema**, written in the **interface definition language (IDL)**:
```protobuf
syntax = "proto3";
message Person {
string user_name = 1;
int64 favorite_number = 2;
repeated string interests = 3;
}
```
**A code generation tool** produces classes implementing the schema in various languages; your application calls the generated code to encode/decode.
> **The schema language is very simple compared to JSON Schema: it defines fields and their types, but does NOT support other restrictions on possible values.**
#### 4.1 The encoding — 33 bytes [#41-the-encoding--33-bytes]
**Compare with MessagePack:** no `userName`, `favoriteNumber`, `interests` strings anywhere. **Field tags are aliases for fields — a compact way of indicating which field we mean without spelling out the name.**
#### 4.2 Schema evolution rules [#42-schema-evolution-rules]
**An encoded record is just the concatenation of its encoded fields.** Each field is identified by tag number and annotated with a datatype. **If a field value is not set, it is simply omitted.**
***
# 5.12 Self-test (/docs/ddia/encoding-evolution/self-test)
Define backward and forward compatibility. Which is usually harder, and precisely why?
An older client calls a newer service. Which compatibility do you need on the request, and which on the response? Now reverse the versions.
Draw the three-step sequence by which an old node **silently deletes** a field added by a new node. What property must the decoder have to prevent it?
Give three reasons never to use `pickle`/`java.io.Serializable` for persisted or transmitted data.
Why does X's API return post IDs twice? What is the exact numeric boundary involved?
What is an "open content model" in JSON Schema, and why does it mean schemas define what *isn't* permitted rather than what is?
MessagePack saves 19% over JSON; Protobuf saves 59%. Where does the extra saving come from?
In Protobuf, why can you rename a field but never change its tag? Why must you `reserved` a deleted tag?
Explain how Protobuf achieves forward compatibility for an unknown field. What specific piece of information makes skipping possible?
Why is an Avro-encoded record undecodable without a schema, when a Protobuf one is decodable (structurally) without one?
State the single Avro rule for adding/removing fields. Derive from it why adding a field without a default breaks backward compatibility.
Name the three mechanisms by which an Avro reader learns the writer's schema, and the context each suits.
Why is Avro friendlier than Protobuf to **dynamically generated** schemas? What manual step does Protobuf require?
"Data outlives code." Give two mechanisms by which databases avoid rewriting old data on a schema change.
List the six ways a network request differs from a local function call. Which one makes retries dangerous?
Why does location transparency work better in the actor model than in RPC?
Compare DNS-based load balancing with a service-discovery system on two axes: change propagation speed, and metadata richness.
What simplifying assumption can you make about RPC deployments that you *cannot* make about databases? What does it let you skip?
Explain how a durable execution framework avoids double-charging a credit card. Why does it still require the payment gateway to be idempotent?
you run a public JSON API used by 400 external customers, an internal gRPC mesh, and a Kafka event backbone feeding a warehouse. A core entity gains a new required field and one field is renamed. Write the migration plan — encoding by encoding, with the compatibility direction you're relying on at each step and the order in which you deploy.
# 5.8 Technology deep dives (/docs/ddia/encoding-evolution/technology-deep-dives)
***
#### 8.1 Protocol Buffers / gRPC [#81-protocol-buffers--grpc]
**Problem it solves.** Compact, fast, cross-language encoding with **explicitly defined forward/backward compatibility**, plus generated type-safe client and server code.
**Why wasn't JSON enough?** \~2.5× larger on the wire, no schema by default (so no compatibility checking and no generated types), ambiguous number semantics, no binary strings, and slower to parse.
**Why wasn't Avro enough?** Protobuf's tag-based encoding means **the reader doesn't need the writer's schema at all** — each message is self-delimiting. That makes it far better suited to **request/response over a network** where negotiating a writer schema per call would be a burden.
**How it works internally.** Tag+wire-type packed into a key byte; varint integers; length-delimited strings/bytes/submessages; `repeated` = repeated occurrences of the same tag. **Unknown fields are skippable because the wire type tells the parser the length** — and proto3 (since 3.5) **retains** unknown fields rather than dropping them, which is what makes forward compatibility safe.
gRPC adds: **HTTP/2 transport** (multiplexed streams, header compression), **four call types** (unary, server-streaming, client-streaming, bidirectional), deadlines propagated through the call chain, and per-call metadata.
**Deployment.** `.proto` files in a **shared repo or a schema registry**, with generated code published as versioned artifacts per language. **`buf` (or `protolock`) in CI to enforce compatibility rules** — this is the single highest-value practice, and it's the mechanical equivalent of the rules in §4.2.
**Monitoring.** Per-method latency and status-code distribution (gRPC codes, not HTTP); **deadline-exceeded rate** (the clearest signal of a too-tight timeout or a slow dependency); message sizes (p99 — a 4 MB default max is a real limit people hit); **unknown-field rate** if you can instrument it, as a leading indicator of version skew; connection churn.
**Scaling.** HTTP/2 multiplexes over **one long-lived TCP connection**, which defeats connection-level L4 load balancers — you need **L7/gRPC-aware balancing** or client-side balancing with a service mesh. This surprises nearly everyone the first time.
**Backup.** N/A for the wire format; **the `.proto` files and the tag-reservation history are the artifact to preserve.** Losing the record of which tags were once used is how you get silent corruption years later.
**What actually breaks in production.**
* **Reusing a retired field tag.** The nightmare case: old persisted data has tag 7 as a string; new code declares tag 7 as an int64. Decoding produces garbage or crashes, and there's no error at deploy time. **`reserved 7;` exists precisely to prevent this** — use it every single time you delete a field.
* **`int32` → `int64` widening**, then **old code truncating** a large value it reads back. Silent data corruption.
* **proto2 `required` fields** — the field that can never be removed, on either side, forever. This is why proto3 removed the concept.
* **Default-value ambiguity in proto3:** an unset field and a field explicitly set to `0`/`""`/`false` are indistinguishable on the wire. "Set the discount to 0" and "don't change the discount" become the same message. Requires wrapper types or `optional` (reintroduced in proto3 3.15).
* **A gRPC deadline not propagated**, so the server keeps working on a request whose caller gave up.
* **Load balancing on L4** with HTTP/2 → all traffic pinned to whichever backend won the connection race.
***
#### 8.2 Avro + a schema registry (Confluent Schema Registry, Apicurio) [#82-avro--a-schema-registry-confluent-schema-registry-apicurio]
**Problem it solves.** Encode millions of records compactly where **the schema is known per-file or per-topic rather than per-message**, and support **dynamically generated schemas** from a database or ETL job.
**Why wasn't Protobuf enough?** **Tag numbers must be assigned by hand**, which makes generating a schema from a relational table (or a changing CSV, or an inferred JSON shape) an administrative chore that's dangerous to automate — **the generator must never reassign a previously used tag.** Avro identifies fields by **name**, so regeneration is safe.
**Why wasn't JSON-per-record enough?** Field names repeated in every one of a billion records; no compatibility enforcement.
**How it works internally.** §5: values concatenated with no field identifiers; schema resolution matches writer↔reader **by name**; defaults fill absent fields; unions for nullability.
**The registry's role:** each Kafka message is prefixed with a **magic byte + 4-byte schema ID**; the consumer fetches the writer's schema by ID (and caches it), then resolves against its own reader schema.
**Deployment.** Registry as an HA service (Confluent's is itself backed by a Kafka topic); **compatibility level set per subject**; schema evolution gated in CI by a `test-compatibility` call before merge.
**Monitoring.** Registry availability and p99 latency (**it is on the critical path of every consumer's cold start**); cache hit rate in clients; **schema version count per subject** (rapid growth means something is generating schemas per-deploy, usually a bug); compatibility-check rejection rate.
**Scaling.** Clients cache schemas indefinitely, so steady-state load is near zero — **the load is a thundering herd at consumer restart.**
**Backup.** **Back up the `_schemas` topic / registry store.** Losing the registry makes every historical message on every topic undecodable. This is a genuine single point of catastrophic failure, and it's routinely un-backed-up.
**What actually breaks.**
* **Adding a field without a default.** Breaks backward compatibility; the registry rejects it if configured correctly, and silently poisons every consumer if it isn't.
* **Registry down at consumer start** → the whole consumer fleet can't decode anything. Cache warm-up and a fallback are worth designing.
* **`NONE` compatibility level** set "temporarily" during an incident and never restored.
* **Renaming a field** — backward compatible via aliases, **not forward compatible.** Old readers break. Almost nobody remembers this asymmetry.
* **A default of `null` not first in the union**, so it's rejected — or worse, a union reordered so branch indexes shift and old data decodes as the wrong type.
* **Two producers writing the same topic with divergent schema lineages.**
***
#### 8.3 REST/OpenAPI services [#83-restopenapi-services]
**Problem it solves.** Cross-organization, long-lived, human-inspectable APIs over ubiquitous HTTP infrastructure.
**Why wasn't RPC enough?** §7.3's six reasons, plus: **REST treats state transfer over a network as a process distinct from a function call**, which sets honest expectations. And HTTP brings caching, content negotiation, auth, and intermediaries for free.
**Why wasn't gRPC enough for public APIs?** Browsers can't speak raw gRPC without a proxy; binary payloads are undebuggable with `curl`; and external partners overwhelmingly expect JSON.
**How it works internally.** URLs identify resources; HTTP verbs and status codes carry semantics; `Cache-Control`/`ETag` for caching; `Accept` for content negotiation. OpenAPI describes it; tooling generates docs, SDKs, mocks, and **compatibility checks**.
**Deployment.** Gateway (auth, rate limiting, quotas) → service. Versioning strategy chosen **once, deliberately** — URL path, `Accept` header, or server-stored per-API-key version.
**Monitoring.** Per-endpoint p99 and error rate; **usage per API version and per client** — the only data that tells you when a deprecated version can actually be removed; deprecation-header adoption; payload sizes; **`Sunset`/`Deprecation` header** delivery.
**Scaling.** Stateless services behind L7 balancers; CDN for cacheable GETs; per-client rate limits.
**What actually breaks.**
* **A "compatible" change that isn't.** Tightening a validation rule, narrowing an enum, changing a nullable field to non-nullable, or **changing a number to a string** — all break clients while looking innocuous in a diff.
* **The forgotten v1** you can never turn off, because three enterprise customers still call it and nobody instrumented per-version usage.
* **Clients that break on unknown fields.** Adding a response field is supposed to be safe — it isn't, if the client uses a strict deserializer. This is the §0 trap on the API surface, and it's why "additive changes are safe" is only true if clients are lenient.
* **2⁵³ in JSON** — a bigint ID silently mangled by a JavaScript client.
* **Breaking change shipped without a deprecation window**, because the provider had no control over the clients and no way to force an upgrade.
***
#### 8.4 Message brokers (Kafka, RabbitMQ, SQS, Pub/Sub) [#84-message-brokers-kafka-rabbitmq-sqs-pubsub]
**Problem it solves.** Decouple sender from recipient in time, in location, and in fan-out — buffering when the recipient is down, redelivering on crash, and letting many consumers read the same message.
**Why wasn't direct RPC enough?** If the recipient is down, the request fails; the sender must know the recipient's address; one message can't reach five consumers; and a slow consumer applies backpressure directly to the user-facing request path.
**How it works internally.** Ch 12 covers this properly. Briefly: **RabbitMQ-style** brokers track per-message acknowledgment and delete on ack; **Kafka-style log brokers** append to a partitioned, ordered, retained log and track a consumer **offset** — which is what makes replay and event sourcing possible.
**Deployment.** Broker cluster with replication; a **schema registry alongside** (§8.2); dead-letter queues; consumer groups sized to partition count.
**Monitoring.** **Consumer lag per partition** — the metric; broker disk usage vs retention; rebalance frequency (frequent rebalances mean consumers are timing out); DLQ depth and age; **redelivery rate**; end-to-end latency from produce to process.
**Scaling.** Partition count bounds consumer parallelism and **is painful to increase** (it changes key→partition mapping). Choose it with room to grow.
**Backup.** Topics with infinite retention are a system of record and need real backups (mirroring to a second cluster, or tiered storage to object storage). Topics with 7-day retention are a buffer, not a store — know which each of yours is.
**What actually breaks.**
* **Unknown-field loss on republish** — the §7.7 warning, made concrete: a consumer decodes into an old model, transforms, and republishes, silently dropping fields added by a newer producer.
* **Poison messages** looping forever without a DLQ, blocking the partition behind them.
* **Retention expiry before a stalled consumer catches up** → permanent data loss with no error, just a jump in the offset.
* **Ordering assumptions across partitions** — order is per-partition only.
* **A consumer that isn't idempotent**, meeting at-least-once delivery.
* **Rebalance storms** when processing time exceeds `max.poll.interval.ms`, so the group never stabilizes and throughput goes to zero.
***
#### 8.5 Durable execution engines (Temporal, Restate) — and workflow orchestrators (Airflow, Dagster) [#85-durable-execution-engines-temporal-restate--and-workflow-orchestrators-airflow-dagster]
**Problem it solves.** Exactly-once multi-step business processes spanning services and third parties, where **you cannot wrap the steps in a database transaction.**
**Why weren't distributed transactions enough?** Ch 8 shows 2PC is rarely usable across microservices and third-party APIs, and it runs counter to service independence.
**Why wasn't a retry loop enough?** Retrying step 3 must not re-charge the card in step 2. You need to know **what already succeeded**, durably, across process crashes.
**How it works internally.** Every RPC and state change is **appended to a durable event history (a write-ahead log)**. On replay, the framework **re-executes the workflow code from the top**, but **intercepts each activity call**: if the history shows it already completed, it **returns the recorded result instead of calling out again.** The workflow function is therefore a **deterministic reduction over its own history** — which is precisely why nondeterminism is fatal.
**Deployment.** Worker fleet polling task queues + a service cluster backed by a database (Cassandra/MySQL/Postgres for Temporal). **Workflow code versioned, never edited in place** — deploy a new version so in-flight executions keep the old one.
**Monitoring.** Workflow-execution latency and failure rate; **activity retry counts**; **stuck/blocked workflows** (the ones that silently accumulate); task-queue backlog and worker poll-success rate; **nondeterminism-error count** — a nonzero value means a deploy broke replay for in-flight executions; workflow history length (there are hard limits, and a long-running loop hits them).
**Scaling.** Workers scale horizontally and statelessly; the state store is the bottleneck. Long-running workflows must use `continueAsNew` to bound history size.
**Backup.** The **event history is the system of record** — back it up like a database, because losing it strands every in-flight business process in an unknowable state.
**What actually breaks.**
* **Nondeterminism introduced by an innocent edit.** Reorder two activity calls, or add an `if` before an existing one, and **replay of in-flight workflows diverges from history** → `NonDeterministicError`. This is the #1 Temporal production incident and it's a *deploy-time* self-inflicted wound.
* **`datetime.now()` / `random()` / `uuid4()` inside workflow code**, producing different values on replay.
* **A non-idempotent third party.** The framework guarantees it won't *re-issue* a recorded call — it cannot guarantee the gateway didn't process a call whose response was lost. **You still need idempotency keys.**
* **Unbounded history** on a long-lived workflow, hitting the event-count limit.
* **Treating Airflow like Temporal.** Airflow orchestrates *idempotent batch tasks* on a schedule; it does not give you exactly-once semantics for stateful business transactions. Conflating the two is a common and expensive design error.
***
# 5.13 Terminology introduced here (/docs/ddia/encoding-evolution/terminology-introduced-here)
# 5.0 The two compatibilities — memorize this (/docs/ddia/encoding-evolution/two-compatibilities-memorize)
| Direction | Definition | Difficulty |
| -------------------------- | -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Backward compatibility** | **Newer code can read data written by OLDER code** | **Normally not hard.** As the author of the newer code, **you know the format written by older code**, so you can explicitly handle it (worst case, keep the old code around to read old data) |
| **Forward compatibility** | **OLDER code can read data written by NEWER code** | **Trickier** — it requires **older code to IGNORE additions made by a newer version** |
**For APIs, work it out from who is newer:**
| Scenario | Requirement |
| -------------------------------- | ---------------------------------------------------------------------- |
| **Older client → newer service** | **Backward compat on the REQUEST**, **forward compat on the RESPONSE** |
| **Newer client → older service** | **Forward compat on the REQUEST**, **backward compat on the RESPONSE** |
#### The forward-compatibility trap: unknown-field loss [#the-forward-compatibility-trap-unknown-field-loss]
> **This is the single most under-appreciated bug class in schema evolution.** It doesn't error, it doesn't log — it just deletes data on a round-trip through an old node during a rolling upgrade.
**How the data models differ on this to begin with:**
* **Relational**: all data conforms to **one schema**; it can be changed via migrations (`ALTER`), but **exactly one schema is in force at any point in time.**
* **Schema-on-read ("schemaless")**: no enforcement, so **the database can contain a mixture of older and newer data formats written at different times.**
***
# 5.11 Worked examples (/docs/ddia/encoding-evolution/worked-examples)
**① Size comparison, same record.**
JSON 81 B → MessagePack 66 B (−19%) → Protobuf 33 B (−59%) → Avro 32 B (−60%).
The step from JSON to a binary-JSON variant buys almost nothing; **the step to a schema-driven format halves it**, because field names leave the payload. At 1 billion records/day, that's \~49 GB/day saved on the wire versus JSON.
**② Varint cost.** `favorite_number = 1337` → 2 bytes. `favorite_number = 1_000_000_000_000` → 6 bytes. `favorite_number = -1` in a plain `int64` field → **10 bytes**, because negative numbers set the high bits. This is why protobuf has `sint64` (zigzag encoding), where −1 costs 1 byte. Choosing `int64` for a field that holds negative deltas is a real, measurable waste.
**③ Compatibility matrix for one change.** You add `string email = 4;` to `Person`.
| | Old code reads | New code reads |
| -------------------------- | ----------------------------------------- | ------------------------------------------------------------------------------------------ |
| **Old data** (no field 4) | fine | `email` = `""` (default) — **can you distinguish "no email" from "empty email"? No.** |
| **New data** (has field 4) | **skips tag 4 by length, preserves it** ✔ | fine |
**④ The same change in Avro.** Add `email` **with** `"default": ""` → both directions fine. Add it **without** a default → new reader + old data = **failure**, because there's nothing to fill in. This asymmetry is the whole Avro rule in one line.
**⑤ Rolling upgrade window.** 100 nodes, 2-minute health check per node, deployed 10 at a time = **\~20 minutes of mixed versions**. During those 20 minutes, every write from a new node may be read by an old node. Now consider a **mobile client**: the equivalent window is **months to years**. That's the real reason forward compatibility isn't optional.
**⑥ Durable execution replay cost.** A workflow with 50 activities that fails at activity 49 replays **all 49 prior activities from history** on the retry — cheap, because they're history reads, not RPCs. But if the history has 20,000 events (a loop), replay itself becomes the bottleneck, and you need `continueAsNew`.
***
# 13.5 Aiming for Correctness (/docs/ddia/philosophy-streaming-systems/aiming-correctness)
> **"With STATELESS services that only read data, it's NOT A BIG DEAL if something goes wrong; you can fix the bug and restart. STATEFUL systems are not so simple. THEY ARE DESIGNED TO REMEMBER THINGS FOREVER, SO IF SOMETHING GOES WRONG, THE EFFECTS ALSO POTENTIALLY LAST FOREVER."**
**The unflattering state of the art:**
> **"For approximately four decades, ATOMICITY, ISOLATION, AND DURABILITY have been the tools of choice. HOWEVER, THOSE FOUNDATIONS ARE WEAKER THAN THEY SEEM: witness the confusion of weak isolation levels."**
>
> **"CONSISTENCY IS OFTEN TALKED ABOUT BUT POORLY DEFINED. Some people assert that we should 'EMBRACE WEAK CONSISTENCY' for the sake of better availability, WHILE LACKING A CLEAR IDEA OF WHAT THAT MEANS IN PRACTICE."**
>
> ### **"For a topic that is so important, OUR UNDERSTANDING AND OUR ENGINEERING METHODS ARE SURPRISINGLY FLAKY. It is VERY DIFFICULT TO DETERMINE WHETHER IT IS SAFE to run a particular application using a particular isolation level or replication configuration. Often, SIMPLE SOLUTIONS APPEAR TO WORK CORRECTLY WHEN CONCURRENCY IS LOW AND THERE ARE NO FAULTS, BUT TURN OUT TO HAVE MANY SUBTLE BUGS IN MORE DEMANDING CIRCUMSTANCES."** [#for-a-topic-that-is-so-important-our-understanding-and-our-engineering-methods-are-surprisingly-flaky-it-is-very-difficult-to-determine-whether-it-is-safe-to-run-a-particular-application-using-a-particular-isolation-level-or-replication-configuration-often-simple-solutions-appear-to-work-correctly-when-concurrency-is-low-and-there-are-no-faults-but-turn-out-to-have-many-subtle-bugs-in-more-demanding-circumstances]
>
> **"Kyle Kingsbury's JEPSEN experiments have highlighted THE STARK DISCREPANCIES BETWEEN SOME PRODUCTS' CLAIMED SAFETY GUARANTEES AND THEIR ACTUAL BEHAVIOR. Even if infrastructure products were free from problems, APPLICATION CODE WOULD STILL NEED TO CORRECTLY USE THE FEATURES THEY PROVIDE, WHICH IS ERROR-PRONE IF THE CONFIGURATION IS HARD TO UNDERSTAND."**
#### 5.1 The end-to-end argument [#51-the-end-to-end-argument]
> ### **"JUST BECAUSE AN APPLICATION USES A DATA SYSTEM THAT PROVIDES COMPARATIVELY STRONG SAFETY PROPERTIES, SUCH AS SERIALIZABLE TRANSACTIONS, THAT DOES NOT MEAN THE APPLICATION IS GUARANTEED TO BE FREE FROM DATA LOSS OR CORRUPTION. If an application has a bug that causes it to write incorrect data or delete data, SERIALIZABLE TRANSACTIONS AREN'T GOING TO SAVE YOU."** [#just-because-an-application-uses-a-data-system-that-provides-comparatively-strong-safety-properties-such-as-serializable-transactions-that-does-not-mean-the-application-is-guaranteed-to-be-free-from-data-loss-or-corruption-if-an-application-has-a-bug-that-causes-it-to-write-incorrect-data-or-delete-data-serializable-transactions-arent-going-to-save-you]
>
> **"This is an argument in favor of IMMUTABLE AND APPEND-ONLY DATA, because it is easier to recover from such mistakes if YOU REMOVE THE ABILITY OF FAULTY CODE TO DESTROY GOOD DATA."**
**The four-layer duplicate-suppression failure — trace it carefully:**
**The fix — a request ID passed end to end:**
```sql
ALTER TABLE requests ADD UNIQUE (request_id);
BEGIN TRANSACTION;
INSERT INTO requests
(request_id, from_account, to_account, amount)
VALUES('0286FDB8-D7E1-423F-B40B-792B3608036C', 4321, 1234, 11.00);
UPDATE accounts SET balance = balance + 11.00 WHERE account_id = 1234;
UPDATE accounts SET balance = balance - 11.00 WHERE account_id = 4321;
COMMIT;
```
> **"Generate a unique identifier for each request (such as a UUID) and include it AS A HIDDEN FORM FIELD in the client application, or CALCULATE A HASH OF ALL THE RELEVANT FORM FIELDS to derive the request ID. IF THE BROWSER SUBMITS TWICE, THE TWO REQUESTS WILL HAVE THE SAME REQUEST ID."**
>
> **"This relies on A UNIQUENESS CONSTRAINT. RELATIONAL DATABASES CAN GENERALLY MAINTAIN A UNIQUENESS CONSTRAINT CORRECTLY, EVEN AT WEAK ISOLATION LEVELS — whereas an APPLICATION-LEVEL CHECK-THEN-INSERT MAY FAIL under nonserializable isolation."**
>
> **Bonus: "the `requests` table acts as A KIND OF EVENT LOG, useful for event sourcing or CDC. The updates to the balances DON'T HAVE TO HAPPEN IN THE SAME TRANSACTION, since they are REDUNDANT AND COULD BE DERIVED FROM THE REQUEST EVENT in a downstream consumer — as long as the event is processed exactly once, WHICH CAN AGAIN BE ENFORCED USING THE REQUEST ID."**
**The principle, stated in 1984 by Saltzer, Reed, and Clark:**
> ### **"THE FUNCTION IN QUESTION CAN COMPLETELY AND CORRECTLY BE IMPLEMENTED ONLY WITH THE KNOWLEDGE AND HELP OF THE APPLICATION STANDING AT THE ENDPOINTS OF THE COMMUNICATION SYSTEM. THEREFORE, PROVIDING THAT QUESTIONED FUNCTION AS A FEATURE OF THE COMMUNICATION SYSTEM ITSELF IS NOT POSSIBLE. (Sometimes an incomplete version provided by the communication system MAY BE USEFUL AS A PERFORMANCE ENHANCEMENT.)"** [#the-function-in-question-can-completely-and-correctly-be-implemented-only-with-the-knowledge-and-help-of-the-application-standing-at-the-endpoints-of-the-communication-system-therefore-providing-that-questioned-function-as-a-feature-of-the-communication-system-itself-is-not-possible-sometimes-an-incomplete-version-provided-by-the-communication-system-may-be-useful-as-a-performance-enhancement]
**It generalizes to three things:**
| Function | Low-level mechanism | Why it's insufficient |
| ------------------------- | --------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Duplicate suppression** | TCP sequence numbers; stream processor exactly-once | **Can't prevent a user resubmitting a timed-out form** |
| **Integrity checking** | **Checksums in Ethernet, TCP, TLS** | **"They cannot detect corruption due to BUGS IN THE SOFTWARE at the sending and receiving ends, or CORRUPTION ON THE DISKS."** ⇒ **you need END-TO-END CHECKSUMS** |
| **Encryption** | **WiFi password** protects against local snooping; **TLS** protects against network attackers | **"Neither protects against COMPROMISES OF THE SERVER. Only END-TO-END ENCRYPTION AND AUTHENTICATION can protect against all these things"** |
> **"Although the low-level features CANNOT PROVIDE THE DESIRED END-TO-END FEATURES BY THEMSELVES, THEY ARE STILL USEFUL, SINCE THEY REDUCE THE PROBABILITY OF PROBLEMS AT HIGHER LEVELS. HTTP requests would often get mangled if we didn't have TCP putting packets back in order. WE JUST NEED TO REMEMBER THAT THE LOW-LEVEL RELIABILITY FEATURES ARE NOT BY THEMSELVES SUFFICIENT."**
**The uncomfortable conclusion:**
> **"That is a shame, because FAULT-TOLERANCE MECHANISMS ARE HARD TO GET RIGHT. It would be really nice to WRAP UP THE HIGH-LEVEL FAULT-TOLERANCE MACHINERY IN AN ABSTRACTION so that application code needn't worry about it — BUT IT SEEMS THAT WE HAVE NOT YET FOUND THE RIGHT ONE."**
>
> **"Transactions collapse a wide range of issues down to TWO POSSIBLE OUTCOMES: COMMIT OR ABORT. That is A HUGE SIMPLIFICATION — BUT IT IS NOT ENOUGH. Transactions are EXPENSIVE, especially with heterogeneous storage. WHEN WE REFUSE TO USE DISTRIBUTED TRANSACTIONS BECAUSE THEY ARE TOO EXPENSIVE, WE END UP HAVING TO REIMPLEMENT FAULT-TOLERANCE MECHANISMS IN APPLICATION CODE. As numerous examples have shown, REASONING ABOUT CONCURRENCY AND PARTIAL FAILURE IS DIFFICULT AND COUNTERINTUITIVE, AND SO MOST APPLICATION-LEVEL MECHANISMS DO NOT WORK CORRECTLY. THE CONSEQUENCE IS LOST OR CORRUPTED DATA."**
#### 5.2 Enforcing constraints without distributed transactions [#52-enforcing-constraints-without-distributed-transactions]
**Uniqueness requires consensus — and here's how to get it from a log:**
**Multishard atomicity WITHOUT atomic commit — the money-transfer worked example:**
#### 5.3 Timeliness vs Integrity — the chapter's most important distinction [#53-timeliness-vs-integrity--the-chapters-most-important-distinction]
**The credit card statement example, which makes it concrete:**
> **"It is NOT SURPRISING if a transaction you made within the last 24 hours DOES NOT YET APPEAR. It is normal that these systems have a certain lag. WE KNOW THAT BANKS RECONCILE AND SETTLE TRANSACTIONS ASYNCHRONOUSLY, AND TIMELINESS IS NOT VERY IMPORTANT HERE."**
>
> **"HOWEVER, IT WOULD BE VERY BAD IF THE STATEMENT BALANCE WAS NOT EQUAL TO THE SUM OF THE TRANSACTIONS PLUS THE PREVIOUS BALANCE (an error in the sums), OR IF A TRANSACTION WAS CHARGED TO YOU BUT NOT PAID TO THE MERCHANT (DISAPPEARING MONEY). SUCH PROBLEMS WOULD BE VIOLATIONS OF THE INTEGRITY OF THE SYSTEM."**
**Why this matters so much for dataflow:**
> **"ACID transactions usually provide BOTH timeliness AND integrity. Thus, if you approach correctness from the point of view of ACID, THE DISTINCTION IS FAIRLY INCONSEQUENTIAL."**
>
> ### **"AN INTERESTING PROPERTY OF EVENT-BASED DATAFLOW SYSTEMS IS THAT THEY DECOUPLE TIMELINESS AND INTEGRITY. When processing asynchronously, THERE IS NO GUARANTEE OF TIMELINESS unless you explicitly build consumers that wait. HOWEVER, INTEGRITY IS IN FACT CENTRAL TO STREAMING SYSTEMS. Exactly-once semantics IS A MECHANISM FOR PRESERVING INTEGRITY."** [#an-interesting-property-of-event-based-dataflow-systems-is-that-they-decouple-timeliness-and-integrity-when-processing-asynchronously-there-is-no-guarantee-of-timeliness-unless-you-explicitly-build-consumers-that-wait-however-integrity-is-in-fact-central-to-streaming-systems-exactly-once-semantics-is-a-mechanism-for-preserving-integrity]
**The four mechanisms that give integrity without distributed transactions:**
1. **"REPRESENTING THE CONTENT OF THE WRITE OPERATION AS A SINGLE MESSAGE, which can easily be written atomically — an approach that FITS VERY WELL WITH EVENT SOURCING"**
2. **"DERIVING ALL OTHER STATE UPDATES FROM THAT SINGLE MESSAGE VIA DETERMINISTIC DERIVATION FUNCTIONS, similarly to stored procedures"**
3. **"PASSING A CLIENT-GENERATED REQUEST ID THROUGH ALL THESE LEVELS OF PROCESSING, enabling END-TO-END DUPLICATE SUPPRESSION AND IDEMPOTENCE"**
4. **"MAKING MESSAGES IMMUTABLE AND ALLOWING DERIVED DATA TO BE REPROCESSED from time to time, WHICH MAKES IT EASIER TO RECOVER FROM BUGS"**
#### 5.4 Loosely interpreted constraints — the business-reality argument [#54-loosely-interpreted-constraints--the-business-reality-argument]
> **"Enforcing a uniqueness constraint requires consensus, typically implemented by funneling all events through a single node. THIS LIMITATION IS UNAVOIDABLE if we want the traditional form. HOWEVER, MANY REAL APPLICATIONS HAVE A BUSINESS REQUIREMENT TO ALLOW VIOLATIONS OF WHAT YOU MIGHT THINK OF AS HARD CONSTRAINTS:"**
| Case | The reality |
| ---------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Overselling stock** | **"You can order in more stock, apologize for the delay, and offer a discount. THIS IS THE SAME AS WHAT YOU'D HAVE TO DO IF A FORKLIFT TRUCK RAN OVER SOME OF THE ITEMS IN YOUR WAREHOUSE. THUS, THE APOLOGY WORKFLOW ALREADY NEEDS TO BE PART OF YOUR BUSINESS PROCESSES ANYWAY, AND A HARD CONSTRAINT MIGHT BE UNNECESSARY"** |
| **Overbooking flights and hotels** | **"The constraint of 'one person per seat' is DELIBERATELY VIOLATED FOR BUSINESS REASONS, and compensation processes (refunds, upgrades, a complimentary room at a neighboring hotel) are put in place. EVEN IF NO OVERBOOKING OCCURRED, APOLOGY AND COMPENSATION PROCESSES WOULD BE NEEDED to deal with flights canceled because of BAD WEATHER OR STAFF GOING ON STRIKE"** |
| **Overdrafts** | **"The bank can charge an overdraft fee and ask them to pay back what they owe. BY LIMITING TOTAL WITHDRAWALS PER DAY, THE RISK TO THE BANK IS BOUNDED"** |
| **Cross-organization integration** | **"Inconsistencies WILL INEVITABLY ARISE, and correction mechanisms are necessary. Settlement of payments between banks is an example"** |
> **A change to correct a mistake is a COMPENSATING TRANSACTION. "The cost of the apology varies, but IT IS OFTEN QUITE LOW; YOU CAN'T UNSEND AN EMAIL, BUT YOU CAN SEND A FOLLOW-UP EMAIL WITH A CORRECTION. If you accidentally charge a credit card twice, you can refund one, and the cost is just the processing fees and perhaps a customer complaint. Once money has been paid out of an ATM you can't directly get it back — although in principle you can SEND DEBT COLLECTORS."**
>
> ### **"IF THE COST OF THE APOLOGY IS ACCEPTABLE, THE TRADITIONAL MODEL OF CHECKING ALL CONSTRAINTS BEFORE EVEN WRITING THE DATA IS UNNECESSARILY RESTRICTIVE. It may well be reasonable to GO AHEAD WITH A WRITE OPTIMISTICALLY AND CHECK THE CONSTRAINT AFTER THE FACT. You can still ensure that VALIDATION OCCURS BEFORE TAKING ACTIONS THAT WOULD BE EXPENSIVE TO RECOVER FROM, BUT THAT DOESN'T IMPLY YOU MUST VALIDATE BEFORE YOU EVEN WRITE THE DATA."** [#if-the-cost-of-the-apology-is-acceptable-the-traditional-model-of-checking-all-constraints-before-even-writing-the-data-is-unnecessarily-restrictive-it-may-well-be-reasonable-to-go-ahead-with-a-write-optimistically-and-check-the-constraint-after-the-fact-you-can-still-ensure-that-validation-occurs-before-taking-actions-that-would-be-expensive-to-recover-from-but-that-doesnt-imply-you-must-validate-before-you-even-write-the-data]
>
> **"These applications DO require integrity. You would not want to lose a reservation or have money disappear. BUT THEY DON'T REQUIRE TIMELINESS ON THE ENFORCEMENT OF THE CONSTRAINT."**
#### 5.5 Coordination-avoiding data systems [#55-coordination-avoiding-data-systems]
**The two observations, combined:**
> ### **THE APOLOGY CALCULUS: "Coordination and constraints REDUCE THE NUMBER OF APOLOGIES YOU HAVE TO MAKE FOR INCONSISTENCIES, BUT POTENTIALLY ALSO REDUCE THE PERFORMANCE AND AVAILABILITY OF YOUR SYSTEM, AND THUS POTENTIALLY INCREASE THE NUMBER OF APOLOGIES YOU HAVE TO MAKE FOR OUTAGES. YOU CANNOT REDUCE THE NUMBER OF APOLOGIES TO ZERO, BUT YOU CAN AIM TO FIND THE BEST TRADE-OFF — THE SWEET SPOT WITH NEITHER TOO MANY INCONSISTENCIES NOR TOO MANY AVAILABILITY PROBLEMS."** [#the-apology-calculus-coordination-and-constraints-reduce-the-number-of-apologies-you-have-to-make-for-inconsistencies-but-potentially-also-reduce-the-performance-and-availability-of-your-system-and-thus-potentially-increase-the-number-of-apologies-you-have-to-make-for-outages-you-cannot-reduce-the-number-of-apologies-to-zero-but-you-can-aim-to-find-the-best-trade-off--the-sweet-spot-with-neither-too-many-inconsistencies-nor-too-many-availability-problems]
#### 5.6 Trust, but Verify [#56-trust-but-verify]
> **"Traditionally, system models take A BINARY APPROACH toward faults: we assume that some things can happen and that other things CAN NEVER happen. IN REALITY, IT IS MORE A QUESTION OF PROBABILITIES. The question is WHETHER VIOLATIONS OF OUR ASSUMPTIONS HAPPEN OFTEN ENOUGH THAT WE MAY ENCOUNTER THEM IN PRACTICE."**
>
> **"We have seen that data can become corrupted IN MEMORY, ON DISK, AND ON THE NETWORK. MAYBE THIS IS SOMETHING WE SHOULD BE PAYING MORE ATTENTION TO? IF YOU ARE OPERATING AT LARGE ENOUGH SCALE, EVEN VERY UNLIKELY THINGS DO HAPPEN."**
**Software bugs in the systems you trust most:**
> **"Even widely used database software has bugs — PAST VERSIONS OF MySQL HAVE FAILED TO CORRECTLY MAINTAIN UNIQUENESS CONSTRAINTS, AND POSTGRESQL'S SERIALIZABLE ISOLATION LEVEL HAS EXHIBITED WRITE SKEW ANOMALIES IN THE PAST — even though MySQL and PostgreSQL are ROBUST AND WELL-REGARDED databases BATTLE-TESTED BY MANY PEOPLE FOR MANY YEARS. IN LESS MATURE SOFTWARE, THE SITUATION IS LIKELY TO BE MUCH WORSE."**
>
> **"When it comes to APPLICATION code, we have to assume MANY MORE BUGS, since most applications don't receive anywhere near the amount of review and testing that database code does. MANY APPLICATIONS DON'T EVEN CORRECTLY USE THE FEATURES DATABASES OFFER FOR PRESERVING INTEGRITY, SUCH AS FOREIGN-KEY OR UNIQUENESS CONSTRAINTS."**
>
> ### **"ACID consistency is based on the idea that the database starts in a consistent state and a transaction transforms it to another consistent state. HOWEVER, THIS NOTION MAKES SENSE ONLY IF WE ASSUME THE TRANSACTION IS FREE FROM BUGS. IF THE APPLICATION USES THE DATABASE INCORRECTLY — FOR EXAMPLE, USING A WEAK ISOLATION LEVEL UNSAFELY — THE INTEGRITY OF THE DATABASE CANNOT BE GUARANTEED."** [#acid-consistency-is-based-on-the-idea-that-the-database-starts-in-a-consistent-state-and-a-transaction-transforms-it-to-another-consistent-state-however-this-notion-makes-sense-only-if-we-assume-the-transaction-is-free-from-bugs-if-the-application-uses-the-database-incorrectly--for-example-using-a-weak-isolation-level-unsafely--the-integrity-of-the-database-cannot-be-guaranteed]
**What mature systems actually do:**
> **"Large-scale storage systems such as HDFS AND AMAZON S3 DO NOT FULLY TRUST DISKS. These systems run BACKGROUND PROCESSES THAT CONTINUALLY READ BACK FILES, COMPARE THEM TO OTHER REPLICAS, AND MOVE FILES FROM ONE DISK TO ANOTHER, in order to mitigate the risk of SILENT CORRUPTION."**
>
> ### **"IF YOU WANT TO BE SURE THAT YOUR DATA IS STILL THERE, YOU HAVE TO READ IT AND CHECK. By the same argument, IT IS IMPORTANT TO TRY RESTORING FROM YOUR BACKUPS FROM TIME TO TIME — OTHERWISE YOU MAY FIND OUT THAT YOUR BACKUP IS BROKEN WHEN IT IS TOO LATE AND YOU HAVE ALREADY LOST DATA. DON'T JUST BLINDLY TRUST THAT IT IS ALL WORKING."** [#if-you-want-to-be-sure-that-your-data-is-still-there-you-have-to-read-it-and-check-by-the-same-argument-it-is-important-to-try-restoring-from-your-backups-from-time-to-time--otherwise-you-may-find-out-that-your-backup-is-broken-when-it-is-too-late-and-you-have-already-lost-data-dont-just-blindly-trust-that-it-is-all-working]
>
> **"NOT MANY SYSTEMS CURRENTLY HAVE THIS KIND OF 'TRUST, BUT VERIFY' APPROACH OF CONTINUALLY AUDITING THEMSELVES. MANY ASSUME THAT CORRECTNESS GUARANTEES ARE ABSOLUTE AND MAKE NO PROVISION FOR THE POSSIBILITY OF RARE DATA CORRUPTION. In the future we may see more SELF-VALIDATING OR SELF-AUDITING SYSTEMS."**
**Designing for auditability:**
> **"If a transaction mutates several objects, THE UNDERLYING REASON CAN BE DIFFICULT TO TELL AFTER THE FACT. Even if you capture the transaction logs, THE INSERTIONS, UPDATES, AND DELETIONS DO NOT NECESSARILY GIVE A CLEAR PICTURE OF WHY those mutations were performed. THE INVOCATION OF THE APPLICATION LOGIC THAT DECIDED ON THOSE MUTATIONS IS TRANSIENT AND CANNOT BE REPRODUCED."**
>
> **"By contrast, EVENT-BASED SYSTEMS CAN PROVIDE BETTER AUDITABILITY. User input is represented as a SINGLE IMMUTABLE EVENT, and any resulting state updates are DERIVED from it. The derivation can be made DETERMINISTIC AND REPEATABLE."**
>
> **"Being explicit about dataflow MAKES THE PROVENANCE OF DATA MUCH CLEARER. For the event log, we can USE HASHES to check that the event storage has not been corrupted. For any derived state, we can RERUN THE BATCH AND STREAM PROCESSORS to check whether we get the same result, OR EVEN RUN A REDUNDANT DERIVATION IN PARALLEL."**
>
> **"A deterministic and well-defined dataflow also makes it easier to DEBUG AND TRACE the execution to determine WHY it did something — A KIND OF TIME-TRAVEL DEBUGGING CAPABILITY."**
**And the end-to-end argument, one more time:**
> ### **"CHECKING THE INTEGRITY OF DATA SYSTEMS IS BEST DONE IN AN END-TO-END FASHION. THE MORE SYSTEMS WE CAN INCLUDE IN AN INTEGRITY CHECK, THE FEWER OPPORTUNITIES THERE ARE FOR CORRUPTION TO GO UNNOTICED. IF WE CAN CHECK THAT AN ENTIRE DERIVED DATA PIPELINE IS CORRECT END TO END, THEN ANY DISKS, NETWORKS, SERVICES, AND ALGORITHMS ALONG THE PATH ARE IMPLICITLY INCLUDED IN THE CHECK."** [#checking-the-integrity-of-data-systems-is-best-done-in-an-end-to-end-fashion-the-more-systems-we-can-include-in-an-integrity-check-the-fewer-opportunities-there-are-for-corruption-to-go-unnoticed-if-we-can-check-that-an-entire-derived-data-pipeline-is-correct-end-to-end-then-any-disks-networks-services-and-algorithms-along-the-path-are-implicitly-included-in-the-check]
>
> **"Having continuous end-to-end integrity checks GIVES YOU INCREASED CONFIDENCE ABOUT THE CORRECTNESS OF YOUR SYSTEMS, WHICH IN TURN ALLOWS YOU TO MOVE FASTER. LIKE AUTOMATED TESTING, AUDITING INCREASES THE CHANCES THAT BUGS WILL BE FOUND QUICKLY. IF YOU ARE NOT AFRAID OF MAKING CHANGES, YOU CAN MUCH BETTER EVOLVE AN APPLICATION."**
**Tools — and the blockchain connection, treated soberly:**
> **"A transaction log can be made TAMPER-PROOF by periodically signing it with a hardware security module, BUT THAT DOES NOT GUARANTEE THAT THE RIGHT TRANSACTIONS WENT INTO THE LOG IN THE FIRST PLACE."**
>
> **"BLOCKCHAINS are SHARED APPEND-ONLY LOGS WITH CRYPTOGRAPHIC CONSISTENCY CHECKS; THE TRANSACTIONS THEY STORE ARE EVENTS, AND SMART CONTRACTS ARE BASICALLY STREAM PROCESSORS. The difference from the consensus protocols of Ch 10 is that blockchains are BYZANTINE FAULT-TOLERANT — they still work if some nodes have corrupted data BECAUSE THE REPLICAS CONTINUALLY CHECK ONE ANOTHER'S INTEGRITY."**
>
> **"FOR MOST APPLICATIONS, BLOCKCHAINS HAVE TOO HIGH AN OVERHEAD TO BE USEFUL. HOWEVER, SOME OF THEIR CRYPTOGRAPHIC TOOLS CAN BE USED IN A LIGHTER-WEIGHT CONTEXT. MERKLE TREES are trees of hashes that can efficiently prove a record appears in a dataset. CERTIFICATE TRANSPARENCY uses cryptographically verified append-only logs and Merkle trees to check TLS/SSL certificates; IT AVOIDS NEEDING A CONSENSUS PROTOCOL BY HAVING A SINGLE LEADER PER LOG."**
***
# 13.12 Backward links — where each idea came from (/docs/ddia/philosophy-streaming-systems/backward-links-where-each)
| Idea here | Established in |
| ---------------------------------------------------------------------------------------- | ------------------------ |
| Systems of record vs derived data | **Ch 1** §1.8 |
| Evolvability; minimizing irreversibility | **Ch 2** §5.3 |
| Event sourcing and CQRS | **Ch 3** §3 |
| Indexes as derived structures; the write/read trade-off | **Ch 4** |
| Rolling upgrades; schema evolution | **Ch 5** |
| The dual-write race; multi-leader conflicts; read-your-writes | **Ch 6**, **Ch 12** §3.1 |
| Sharding and secondary indexes (term- vs document-partitioned) | **Ch 7** §5 |
| Weak isolation levels; write skew; 2PC and XA's failures | **Ch 8** |
| System models; what we assume can and cannot fail; fencing | **Ch 9** |
| Total order broadcast = consensus; state machine replication; uniqueness needs consensus | **Ch 10** |
| Immutable inputs; reprocessing; batch fault tolerance | **Ch 11** |
| CDC, log compaction, exactly-once, time-dependence of joins | **Ch 12** |
| Ethics, regulation, and the limits of deletion | **Ch 14** |
# 13.1 Data Integration (/docs/ddia/philosophy-streaming-systems/data-integration)
**The starting observation:**
> **"If you have a problem such as 'I want to store some data and look it up again later,' THERE IS NO ONE RIGHT SOLUTION, but many approaches each appropriate in different circumstances. A software implementation typically has to PICK ONE. It's hard enough to get ONE code path robust and performing well; TRYING TO SATISFY TOO MANY USE CASES WITH MANY FEATURES IS LIKELY TO LEAD TO POOR IMPLEMENTATIONS OF THOSE FEATURES compared to specialized tools."**
>
> ### **"EVERY PIECE OF SOFTWARE — EVEN A SO-CALLED 'GENERAL-PURPOSE' DATABASE — IS DESIGNED FOR A PARTICULAR USAGE PATTERN."** [#every-piece-of-software--even-a-so-called-general-purpose-database--is-designed-for-a-particular-usage-pattern]
**Two challenges, in order:**
1. **Map software products to the circumstances they fit.** *"Vendors are UNDERSTANDABLY RELUCTANT to tell you about the kinds of workloads for which their software is POORLY SUITED, but hopefully the previous chapters have equipped you with QUESTIONS TO ASK that will help you READ BETWEEN THE LINES."*
2. **Even with a perfect map, "in complex applications data is used in various ways, and ONE PIECE OF SOFTWARE IS UNLIKELY TO BE SUITABLE FOR ALL OF THEM. Therefore YOU INEVITABLY END UP HAVING TO COBBLE TOGETHER SEVERAL PIECES OF SOFTWARE."**
*(The canonical example: an OLTP database plus a full-text search index. **"Some databases include full-text indexing, which can be sufficient for simple applications, but more sophisticated search requires specialist information-retrieval tools. Conversely, SEARCH INDEXES ARE GENERALLY NOT VERY SUITABLE AS A DURABLE SYSTEM OF RECORD."**)*
#### 1.1 Reasoning about dataflows [#11-reasoning-about-dataflows]
> **"When copies of the same data must be maintained in several storage systems, YOU NEED TO BE VERY CLEAR ABOUT THE INPUTS AND OUTPUTS. WHERE IS DATA WRITTEN FIRST, AND WHICH REPRESENTATIONS ARE DERIVED FROM WHICH SOURCES?"**
> ### **"If it is possible to FUNNEL ALL USER INPUT THROUGH A SINGLE SYSTEM THAT DECIDES ON AN ORDERING FOR ALL WRITES, it becomes much easier to derive other representations by PROCESSING THE WRITES IN THE SAME ORDER. This is an application of STATE MACHINE REPLICATION. WHETHER YOU USE CDC OR AN EVENT SOURCING LOG IS LESS IMPORTANT THAN THE PRINCIPLE OF DECIDING ON A TOTAL ORDER."** [#if-it-is-possible-to-funnel-all-user-input-through-a-single-system-that-decides-on-an-ordering-for-all-writes-it-becomes-much-easier-to-derive-other-representations-by-processing-the-writes-in-the-same-order-this-is-an-application-of-state-machine-replication-whether-you-use-cdc-or-an-event-sourcing-log-is-less-important-than-the-principle-of-deciding-on-a-total-order]
>
> **And: "updating a derived data system based on an event log can often be made DETERMINISTIC AND IDEMPOTENT, MAKING IT QUITE EASY TO RECOVER FROM FAULTS."**
#### 1.2 Derived data vs distributed transactions [#12-derived-data-vs-distributed-transactions]
| | **Distributed transactions** | **Log-based derived data** |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Mechanism** | **An ATOMIC COMMIT PROTOCOL ensures changes are applied atomically** | **Correctness through DETERMINISTIC RETRY AND IDEMPOTENCE** |
| **The biggest difference** | **"After a value is written, YOU CAN IMMEDIATELY READ THE UP-TO-DATE VALUE"** | **"Updated ASYNCHRONOUSLY, so they DO NOT BY DEFAULT GUARANTEE THAT READS ARE UP TO DATE"** |
| **Verdict** | **"Used successfully in environments willing to absorb their performance and operational costs. HOWEVER, XA HAS POOR FAULT TOLERANCE AND PERFORMANCE CHARACTERISTICS, WHICH SEVERELY LIMIT ITS USEFULNESS. It might be possible to create a better protocol, BUT GETTING IT WIDELY ADOPTED AND INTEGRATED WITH EXISTING TOOLS WOULD BE CHALLENGING, AND IT IS UNLIKELY TO HAPPEN SOON"** | **"In the absence of widespread support for a good distributed transaction protocol, LOG-BASED DERIVED DATA IS THE MOST PROMISING APPROACH for integrating different data systems"** |
> **The intellectual honesty worth quoting: "Guarantees such as reading your own writes ARE USEFUL, and IT IS NOT PRODUCTIVE TO TELL EVERYONE 'EVENTUAL CONSISTENCY IS INEVITABLE — SUCK IT UP AND LEARN TO DEAL WITH IT' (at least not without good guidance on how to deal with it)."**
#### 1.3 The limits of total ordering [#13-the-limits-of-total-ordering]
**Total order is feasible in small systems — "as demonstrated by the popularity of databases with single-leader replication, which construct precisely such a log." But four limits emerge at scale:**
#### 1.4 Ordering events to capture causality — the unfriending problem [#14-ordering-events-to-capture-causality--the-unfriending-problem]
> **"If NO CAUSAL LINK exists between events, the lack of a total order IS NOT A BIG PROBLEM, since concurrent events can be ordered arbitrarily. Some cases are easy — multiple updates of the SAME OBJECT can be totally ordered by ROUTING ALL UPDATES FOR THAT OBJECT ID TO THE SAME LOG SHARD. HOWEVER, CAUSAL DEPENDENCIES SOMETIMES ARISE IN MORE SUBTLE WAYS."**
> **"Unfortunately, THIS PROBLEM DOESN'T SEEM TO HAVE A SIMPLE ANSWER."** Three starting points:
>
> 1. **LOGICAL TIMESTAMPS "provide total ordering WITHOUT COORDINATION, so they may help when total order broadcast is not feasible. HOWEVER, THEY STILL REQUIRE RECIPIENTS TO HANDLE EVENTS DELIVERED OUT OF ORDER, and they require ADDITIONAL METADATA to be passed around"**
> 2. **"If you can LOG AN EVENT TO RECORD THE STATE OF THE SYSTEM THAT THE USER SAW BEFORE MAKING A DECISION and give it a unique identifier, then ANY LATER EVENTS CAN REFERENCE THAT EVENT IDENTIFIER in order to record the causal dependency"**
> 3. **CONFLICT RESOLUTION ALGORITHMS "help with processing events delivered in an unexpected order. They are USEFUL FOR MAINTAINING STATE, BUT THEY DO NOT HELP IF ACTIONS HAVE EXTERNAL SIDE EFFECTS (such as sending a notification)"**
>
> **"Perhaps patterns will emerge in the future that allow causal dependencies to be captured efficiently WITHOUT FORCING ALL EVENTS THROUGH THE BOTTLENECK OF TOTAL ORDER BROADCAST."**
#### 1.5 Batch and stream, together [#15-batch-and-stream-together]
> **"The MAIN FUNDAMENTAL DIFFERENCE is that stream processors operate on UNBOUNDED datasets, whereas batch inputs are of a KNOWN, FINITE SIZE."**
>
> **"Batch processing has A QUITE STRONG FUNCTIONAL FLAVOR (even if the code is not written in a functional language). It encourages DETERMINISTIC, PURE FUNCTIONS whose output depends only on the input and that have NO SIDE EFFECTS other than the explicit outputs, treating INPUTS AS IMMUTABLE AND OUTPUTS AS APPEND-ONLY. Stream processing is similar, but it EXTENDS OPERATORS TO ALLOW A MANAGED, FAULT-TOLERANT STATE."**
**Why asynchrony is the point, not a compromise:**
> ### **"In principle, derived data systems COULD be maintained SYNCHRONOUSLY, just as a relational database updates secondary indexes synchronously within the same transaction. HOWEVER, ASYNCHRONY IS WHAT MAKES SYSTEMS BASED ON EVENT LOGS ROBUST. IT ALLOWS A FAULT IN ONE PART OF THE SYSTEM TO BE CONTAINED LOCALLY, WHEREAS DISTRIBUTED TRANSACTIONS ABORT IF ANY ONE PARTICIPANT FAILS, SO THEY TEND TO AMPLIFY FAILURES BY SPREADING THEM TO THE REST OF THE SYSTEM."** [#in-principle-derived-data-systems-could-be-maintained-synchronously-just-as-a-relational-database-updates-secondary-indexes-synchronously-within-the-same-transaction-however-asynchrony-is-what-makes-systems-based-on-event-logs-robust-it-allows-a-fault-in-one-part-of-the-system-to-be-contained-locally-whereas-distributed-transactions-abort-if-any-one-participant-fails-so-they-tend-to-amplify-failures-by-spreading-them-to-the-rest-of-the-system]
*(And Ch 7's point again: **"a sharded system with secondary indexes needs to either send writes to multiple shards (term-partitioned) or send reads to all shards (document-partitioned). SUCH CROSS-SHARD COMMUNICATION IS ALSO MOST RELIABLE AND SCALABLE IF THE INDEX IS MAINTAINED ASYNCHRONOUSLY."**)*
**Reprocessing for application evolution — and the railway analogy:**
> **"WITHOUT REPROCESSING, SCHEMA EVOLUTION IS LIMITED TO SIMPLE CHANGES like adding a new optional field. WITH REPROCESSING, IT IS POSSIBLE TO RESTRUCTURE A DATASET INTO A COMPLETELY DIFFERENT MODEL to better serve new requirements."**
**Unifying batch and stream — the kappa architecture:**
> **"An early proposal was the LAMBDA ARCHITECTURE, which HAD A NUMBER OF PROBLEMS AND HAS FALLEN OUT OF USE. More recent systems allow batch computations (reprocessing historical data) and stream computations (processing events as they arrive) to be implemented IN THE SAME SYSTEM — sometimes known as the KAPPA ARCHITECTURE."**
**Three features required:**
1. **"The ability to REPLAY HISTORICAL EVENTS THROUGH THE SAME PROCESSING ENGINE that handles the stream of recent events"** — log-based brokers can replay; some stream processors can read from a DFS or object store
2. **"EXACTLY-ONCE SEMANTICS — ensuring the output is the same as if no faults had occurred. As with batch processing, this requires DISCARDING THE PARTIAL OUTPUTS OF ANY FAILED TASKS"**
3. **"Tools for WINDOWING BY EVENT TIME, NOT BY PROCESSING TIME, SINCE PROCESSING TIME IS MEANINGLESS WHEN REPROCESSING HISTORICAL EVENTS"** *(Apache Beam provides such an API, runnable on Flink or Google Cloud Dataflow)*
***
# 13.8 Decision cheat sheet (/docs/ddia/philosophy-streaming-systems/decision-cheat-sheet)
**How do I keep N systems in sync?**
**Designate one system of record. Funnel all input through it. Derive everything else from its ordered change log.** Never let the application write to two stores. Whether it's CDC or event sourcing matters far less than **deciding on a total order**.
**Distributed transaction or derived data?**
**Should I unbundle?**
**Only when no single piece of software satisfies all your requirements.** Unbundling is about **breadth, not depth**. If one database does everything you need, use it. Every additional component brings a learning curve, config quirks, and operational surprises.
**Where do I draw the write-path / read-path boundary?**
Precompute (write path) when the query set is **small and known**; compute on read when it's **large or unbounded**. Split it — cache the common queries, index the rest — and **draw the boundary differently for outliers** (the celebrity case). This is a dial, not a binary.
**Do I need a hard constraint here?**
**And note the asymmetry:** coordination reduces apologies-for-inconsistency but **increases apologies-for-outage.** Optimize the total.
**Timeliness or integrity?**
**Integrity always. Timeliness where it's worth paying for.** Violations of timeliness are temporary and self-healing; **violations of integrity are permanent and require explicit repair.**
**How do I make an operation exactly-once?**
**Mint a request ID on the client before the first attempt. Propagate it through every hop. Enforce a uniqueness constraint at the durable end.** Then make the derivation **deterministic** so replay is safe. No framework can do this for you — it is an **end-to-end** property by definition.
**What must I audit?**
**Anything you'd be unable to reconstruct if it were silently wrong.** At minimum: a periodic **source-vs-derived reconciliation**, a **restore-from-backup test**, and an **invariant check** (debits = credits, counts match). Prefer **end-to-end** checks over per-component ones — they implicitly cover every disk, network, service, and algorithm on the path.
***
# 13.3 Designing Applications Around Dataflow (/docs/ddia/philosophy-streaming-systems/designing-applications-around-dataflow)
#### 3.1 The spreadsheet that data systems still haven't caught up to [#31-the-spreadsheet-that-data-systems-still-havent-caught-up-to]
> **"Spreadsheets have POWERFUL DATAFLOW PROGRAMMING CAPABILITIES: you put a formula in one cell, and WHENEVER ANY INPUT CHANGES, THE RESULT IS AUTOMATICALLY RECALCULATED. THIS IS EXACTLY WHAT WE WANT AT A DATA SYSTEM LEVEL."**
>
> ### **"MOST DATA SYSTEMS STILL HAVE SOMETHING TO LEARN FROM THE FEATURES THAT VISICALC ALREADY HAD IN 1979."** [#most-data-systems-still-have-something-to-learn-from-the-features-that-visicalc-already-had-in-1979]
>
> **The difference: "today's data systems need to be FAULT-TOLERANT, SCALABLE, AND CAPABLE OF STORING DATA DURABLY. They also need to INTEGRATE DISPARATE TECHNOLOGIES written by different groups of people over time. IT IS UNREALISTIC TO EXPECT ALL SOFTWARE TO BE DEVELOPED USING ONE PARTICULAR LANGUAGE, FRAMEWORK, OR TOOL."**
#### 3.2 Application code as a derivation function [#32-application-code-as-a-derivation-function]
| Derived dataset | The transformation function |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Secondary index** | **"For each row, PICK OUT THE VALUES in the columns being indexed and SORT by those values."** *(Built into databases — you invoke it by "merely running `CREATE INDEX`")* |
| **Full-text search index** | **"LANGUAGE DETECTION, WORD SEGMENTATION, STEMMING OR LEMMATIZATION, SPELLING CORRECTION, AND SYNONYM IDENTIFICATION, followed by building a data structure for efficient lookups."** *(Basic linguistic features may be built in, but "MORE SOPHISTICATED FEATURES OFTEN REQUIRE DOMAIN-SPECIFIC TUNING")* |
| **ML model** | **"Derived from the TRAINING DATA by applying various FEATURE EXTRACTION AND STATISTICAL ANALYSIS functions. When applied to new input, its output is derived from that input AND its learned parameters (and hence, INDIRECTLY, from the training data)."** *("Feature engineering is NOTORIOUSLY APPLICATION-SPECIFIC")* |
| **Cache** | **"An aggregation of data IN THE FORM IN WHICH IT IS GOING TO BE DISPLAYED IN A UI. Populating it thus requires KNOWLEDGE OF WHAT FIELDS ARE REFERENCED IN THE UI; CHANGES IN THE UI MAY REQUIRE UPDATING THE DEFINITION OF HOW THE CACHE IS POPULATED AND REBUILDING IT."** |
> **"When the function is NOT a standard cookie-cutter function, CUSTOM CODE IS REQUIRED. THIS CUSTOM CODE IS WHERE MANY DATABASES STRUGGLE. Although relational databases support triggers, stored procedures, and UDFs, THEY HAVE BEEN SOMEWHAT OF AN AFTERTHOUGHT IN DATABASE DESIGN."**
#### 3.3 The separation of Church and state [#33-the-separation-of-church-and-state]
> **"In theory, databases COULD be deployment environments for arbitrary application code, LIKE AN OPERATING SYSTEM. However, in practice they have turned out to be POORLY SUITED. They do not fit well with the requirements of modern application development: DEPENDENCY AND PACKAGE MANAGEMENT, VERSION CONTROL, ROLLING UPGRADES, EVOLVABILITY, MONITORING, METRICS, CALLS TO NETWORK SERVICES, AND INTEGRATION WITH EXTERNAL SYSTEMS."**
>
> **"On the other hand, KUBERNETES, DOCKER, MESOS, YARN are designed SPECIFICALLY for running application code. BY FOCUSING ON DOING ONE THING WELL, they are able to do it MUCH BETTER than a database that provides execution of UDFs as one of its many features."**
>
> ### **"The trend has been to keep STATELESS APPLICATION LOGIC SEPARATE FROM STATE MANAGEMENT: NOT PUTTING APPLICATION LOGIC IN THE DATABASE AND NOT PUTTING PERSISTENT STATE IN THE APPLICATION. As people in the functional programming community like to joke, 'WE BELIEVE IN THE SEPARATION OF CHURCH AND STATE.'"** [#the-trend-has-been-to-keep-stateless-application-logic-separate-from-state-management-not-putting-application-logic-in-the-database-and-not-putting-persistent-state-in-the-application-as-people-in-the-functional-programming-community-like-to-joke-we-believe-in-the-separation-of-church-and-state]
>
> *(The joke explained: **Alonzo Church** created the **lambda calculus**, which **has no mutable state** — "so one could say that mutable state is separate from Church's work.")*
**The passivity problem:**
> **"In this typical model, the database acts as A KIND OF MUTABLE SHARED VARIABLE that can be accessed synchronously over the network. HOWEVER, IN MOST PROGRAMMING LANGUAGES YOU CANNOT SUBSCRIBE TO CHANGES IN A MUTABLE VARIABLE — YOU CAN ONLY READ IT PERIODICALLY. Unlike in a spreadsheet, READERS OF THE VARIABLE DON'T GET NOTIFIED IF THE VALUE CHANGES."** *(You can implement it yourself — the **observer pattern** — but most languages don't have it built in.)*
>
> **"DATABASES HAVE INHERITED THIS PASSIVE APPROACH TO MUTABLE DATA. If you want to find out whether the content has changed, OFTEN YOUR ONLY OPTION IS TO POLL. SUBSCRIBING TO CHANGES IS ONLY JUST BEGINNING TO EMERGE AS A FEATURE."**
#### 3.4 Dataflow: state changes and code, collaborating [#34-dataflow-state-changes-and-code-collaborating]
> ### **"Instead of treating a database as A PASSIVE VARIABLE THAT IS MANIPULATED BY THE APPLICATION, we think much more about THE INTERPLAY AND COLLABORATION BETWEEN STATE, STATE CHANGES, AND CODE THAT PROCESSES THEM. APPLICATION CODE RESPONDS TO STATE CHANGES IN ONE PLACE BY TRIGGERING STATE CHANGES IN ANOTHER PLACE."** [#instead-of-treating-a-database-as-a-passive-variable-that-is-manipulated-by-the-application-we-think-much-more-about-the-interplay-and-collaboration-between-state-state-changes-and-code-that-processes-them-application-code-responds-to-state-changes-in-one-place-by-triggering-state-changes-in-another-place]
**Two properties log-based brokers must provide:**
1. **"THE ORDER of state changes is often important — if several views are derived from an event log, THEY NEED TO PROCESS THE EVENTS IN THE SAME ORDER so that they remain consistent with one another"**
2. **"FAULT TOLERANCE is essential — LOSING JUST A SINGLE MESSAGE CAUSES THE DERIVED DATASET TO GO PERMANENTLY OUT OF SYNC with its data source"**
> **"Stable message ordering and fault-tolerant processing are QUITE STRINGENT DEMANDS, BUT THEY ARE MUCH LESS EXPENSIVE AND MORE OPERATIONALLY ROBUST THAN DISTRIBUTED TRANSACTIONS."**
>
> **"Like UNIX TOOLS CHAINED BY PIPES, stream operators can be composed to build large systems around dataflow. EACH OPERATOR TAKES STREAMS OF STATE CHANGES AS INPUT AND PRODUCES OTHER STREAMS OF STATE CHANGES AS OUTPUT."**
#### 3.5 Stream processors vs services — the currency example [#35-stream-processors-vs-services--the-currency-example]
**The organizational similarity:** *"Composing stream operators into dataflow systems has a lot of similar characteristics to microservices — the advantage is primarily ORGANIZATIONAL SCALABILITY THROUGH LOOSE COUPLING. However, the underlying communication mechanism is very different: ONE-DIRECTIONAL, ASYNCHRONOUS MESSAGE STREAMS rather than synchronous request/response."*
***
# 13. A Philosophy of Streaming Systems (/docs/ddia/philosophy-streaming-systems)
> "If the highest aim of a captain was to preserve his ship, he would keep it in port forever." — St. Thomas Aquinas (paraphrased)
> **This chapter is MORE OPINIONATED than previous chapters, presenting A DEEP DIVE INTO ONE PARTICULAR PHILOSOPHY rather than comparing multiple approaches.**
**The arc:**
***
# 13.4 Observing Derived State — the write path and the read path (/docs/ddia/philosophy-streaming-systems/observing-derived-state-write)
**The full-text search spectrum, which shows the boundary is a dial:**
> **"Shifting the boundary was in fact the topic of the social networking example in Ch 2. In that example we also saw how THE BOUNDARY MIGHT BE DRAWN DIFFERENTLY FOR CELEBRITIES COMPARED TO ORDINARY USERS. AFTER 500 PAGES, WE HAVE COME FULL CIRCLE!"**
#### 4.1 Extending the write path to the end-user device [#41-extending-the-write-path-to-the-end-user-device]
> **"In the past, web browsers were STATELESS CLIENTS that could do useful things only when you had an internet connection (just about the only thing you could do offline was SCROLL UP AND DOWN in a page you had previously loaded). However, single-page JavaScript apps now have A LOT OF STATEFUL CAPABILITIES."**
>
> ### **"When we move away from the assumption of stateless clients talking to a central database and toward state maintained ON END-USER DEVICES, A WORLD OF NEW OPPORTUNITIES OPENS UP. In particular, WE CAN THINK OF THE ON-DEVICE STATE AS A CACHE OF STATE ON THE SERVER. THE PIXELS ON THE SCREEN ARE A MATERIALIZED VIEW OF MODEL OBJECTS IN THE CLIENT APP; THE MODEL OBJECTS ARE A LOCAL REPLICA OF STATE IN A REMOTE DATACENTER."** [#when-we-move-away-from-the-assumption-of-stateless-clients-talking-to-a-central-database-and-toward-state-maintained-on-end-user-devices-a-world-of-new-opportunities-opens-up-in-particular-we-can-think-of-the-on-device-state-as-a-cache-of-state-on-the-server-the-pixels-on-the-screen-are-a-materialized-view-of-model-objects-in-the-client-app-the-model-objects-are-a-local-replica-of-state-in-a-remote-datacenter]
**Pushing state changes to clients:**
> **"If you load a typical web page and the data subsequently changes on the server, THE BROWSER DOES NOT FIND OUT UNTIL YOU RELOAD. The browser reads the data at ONLY ONE POINT IN TIME, ASSUMING IT IS STATIC. Thus THE STATE IN THE BROWSER IS A STALE CACHE that is not updated unless you explicitly poll. (HTTP-based feed subscription protocols LIKE RSS ARE REALLY JUST A BASIC FORM OF POLLING.)"**
>
> **"SERVER-SENT EVENTS (the EventSource API) and WEBSOCKETS provide channels by which a browser can keep AN OPEN TCP CONNECTION and the server can ACTIVELY PUSH MESSAGES."**
>
> ### **"Actively pushing state changes all the way to client devices means EXTENDING THE WRITE PATH ALL THE WAY TO THE END USER. When a client is first initialized, it will STILL NEED TO USE A READ PATH TO GET ITS INITIAL STATE, BUT THEREAFTER IT CAN RELY ON A STREAM OF STATE CHANGES. The ideas around stream processing are thus NOT RESTRICTED TO RUNNING IN A DATACENTER; WE CAN EXTEND THEM ALL THE WAY TO END-USER DEVICES."** [#actively-pushing-state-changes-all-the-way-to-client-devices-means-extending-the-write-path-all-the-way-to-the-end-user-when-a-client-is-first-initialized-it-will-still-need-to-use-a-read-path-to-get-its-initial-state-but-thereafter-it-can-rely-on-a-stream-of-state-changes-the-ideas-around-stream-processing-are-thus-not-restricted-to-running-in-a-datacenter-we-can-extend-them-all-the-way-to-end-user-devices]
**And offline is already solved:** *"The devices will be offline some of the time. BUT WE ALREADY SOLVED THAT PROBLEM: in Ch 12 we discussed how a consumer of a log-based broker can RECONNECT AFTER FAILING and ENSURE IT DOESN'T MISS ANY MESSAGES. THE SAME TECHNIQUE WORKS FOR INDIVIDUAL USERS, WHERE EACH DEVICE IS A SMALL SUBSCRIBER TO A SMALL STREAM OF EVENTS."*
**End-to-end event streams — and why we don't build everything this way:**
> **"Tools such as REACT and ELM already have the ability to UPDATE THE RENDERED UI IN RESPONSE TO CHANGES IN THE UNDERLYING STATE. It would be very natural to extend this programming model to ALSO ALLOW A SERVER TO PUSH STATE CHANGE EVENTS INTO THIS CLIENT-SIDE EVENT PIPELINE."**
>
> **"State changes could then flow through an END-TO-END WRITE PATH: from the interaction on ONE DEVICE, through event logs and various derived data systems and stream processors, ALL THE WAY TO THE USER INTERFACE ON ANOTHER DEVICE — with fairly low delay, say UNDER ONE SECOND END TO END."**
>
> ### **"Some applications — INSTANT MESSAGING AND ONLINE GAMES — ALREADY HAVE SUCH A 'REAL-TIME' ARCHITECTURE. WHY DON'T WE BUILD ALL APPLICATIONS THIS WAY?"** [#some-applications--instant-messaging-and-online-games--already-have-such-a-real-time-architecture-why-dont-we-build-all-applications-this-way]
>
> **"THE CHALLENGE IS THAT THE ASSUMPTION OF STATELESS CLIENTS AND REQUEST/RESPONSE INTERACTIONS IS DEEPLY INGRAINED IN OUR DATABASES, LIBRARIES, FRAMEWORKS, AND PROTOCOLS. Many datastores support operations where a request returns A SINGLE RESPONSE; FAR FEWER SUPPORT OPERATIONS WHERE A REQUEST RETURNS A STREAM OF RESPONSES OVER TIME."**
#### 4.2 Reads are events too [#42-reads-are-events-too]
> **"It is possible to represent READ REQUESTS AS STREAMS OF EVENTS and send BOTH the read events AND the write events through a stream processor. The processor responds to read events by EMITTING THE RESULT OF THE READ TO AN OUTPUT STREAM."**
>
> ### **"When both writes and reads are represented as events and routed to the same operator, WE ARE IN FACT PERFORMING A STREAM-TABLE JOIN BETWEEN THE STREAM OF READ QUERIES AND THE DATABASE."** [#when-both-writes-and-reads-are-represented-as-events-and-routed-to-the-same-operator-we-are-in-fact-performing-a-stream-table-join-between-the-stream-of-read-queries-and-the-database]
>
> **"This correspondence between SERVING REQUESTS and PERFORMING JOINS IS QUITE FUNDAMENTAL: A ONE-OFF READ REQUEST PASSES THROUGH THE JOIN OPERATOR, WHICH THEN IMMEDIATELY FORGETS THE REQUEST; A SUBSCRIBE REQUEST IS A PERSISTENT JOIN WITH PAST AND FUTURE EVENTS ON THE OTHER SIDE OF THE JOIN."**
**Why you might want a log of reads:**
> **"Recording a log of read events potentially has benefits with regard to TRACKING CAUSAL DEPENDENCIES AND DATA PROVENANCE. The log would allow you to RECONSTRUCT WHAT THE USER SAW BEFORE THEY MADE A PARTICULAR DECISION. For example, in an online shop, THE PREDICTED SHIPPING DATE AND THE INVENTORY STATUS SHOWN TO A CUSTOMER LIKELY AFFECT WHETHER THEY CHOOSE TO BUY. To analyze this connection, YOU NEED TO RECORD THE RESULT OF THE USER'S QUERY."** *(This is §1.4's option 2, made concrete.)*
>
> **The cost: "additional STORAGE AND I/O costs. Optimizing such systems is STILL AN OPEN RESEARCH PROBLEM — but IF YOU ALREADY LOG READ REQUESTS FOR OPERATIONAL PURPOSES, IT'S NOT A BIG CHANGE TO MAKE THAT LOG THE SOURCE OF THE REQUESTS INSTEAD."**
**Multishard data processing:** *"For single-shard queries this is perhaps OVERKILL. However, it opens the possibility of DISTRIBUTED EXECUTION OF COMPLEX QUERIES that need to combine data from several shards, TAKING ADVANTAGE OF THE INFRASTRUCTURE FOR MESSAGE ROUTING, SHARDING, AND JOINING THAT IS ALREADY PROVIDED."* Examples: **Storm's distributed RPC computing "the number of people who have seen a URL — the UNION OF THE FOLLOWER SETS of everyone who posted it"**; and **fraud prevention**, where assessing a purchase requires **"reputation scores of the user's IP address, email address, billing address, shipping address — each database itself sharded, so collecting the scores requires A SEQUENCE OF JOINS WITH DIFFERENTLY SHARDED DATASETS."**
***
# 13.7 Production failure catalog for this chapter (/docs/ddia/philosophy-streaming-systems/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ---------------------------------------------------------------- | -------------------------------------------------------------------- |
| Database and search index permanently disagree | **App writes to both** — neither is "in charge" of ordering |
| Ex-partner receives the message they shouldn't | **Causal dependency lost across two services** with no shared order |
| Ordering ambiguous between two regions | **Total order broadcast doesn't scale past one leader** |
| One failing component takes down the whole system | **Synchronous distributed transactions amplify local faults** |
| Schema migration is a terrifying all-or-nothing cutover | No **side-by-side derived views**; irreversibility |
| Reprocessing produces different numbers than the original run | **Nondeterministic derivation** (time-dependent join, external call) |
| Ten pieces of infrastructure, one small team, constant incidents | **Premature unbundling** — "a form of premature optimization" |
| $22 transferred instead of $11 | **Non-idempotent transaction + user retry** past every dedup layer |
| Two accounts created with the same username | **Uniqueness enforced without consensus**, or async multi-leader |
| Multishard transfer half-applied | No request-ID dedup, or a nondeterministic processor |
| Money stuck permanently "reserved" | Lost/undelivered downstream event with no sweeper |
| Credits and debits don't sum to zero | **Integrity violation** — permanent, needs explicit repair |
| Derived store silently drifted from the source | **No reconciliation / end-to-end integrity check** |
| Backup found to be broken during a real incident | **Never restore-tested** |
| Corruption present in every retained backup | Silent corruption undetected for longer than the retention window |
| Uniqueness constraint violated by the database itself | **A database bug** — MySQL has done this |
| "Serializable" isolation exhibited write skew | **A database bug** — PostgreSQL has done this |
| App uses weak isolation unsafely; DB "consistency" meaningless | **ACID consistency assumes bug-free transactions** |
| Cannot explain *why* a set of rows changed | Mutation log without the **intent**; application logic was transient |
| UI shows stale data until reload | **Read path only**; write path never extended to the client |
***
# 13.10 Self-test (/docs/ddia/philosophy-streaming-systems/self-test)
Why does "every piece of software, even a general-purpose database, is designed for a particular usage pattern" lead inevitably to composing multiple systems?
Draw the good and bad dataflow topologies for a database + search index. What exactly goes wrong in the bad one?
What principle matters more than the choice between CDC and event sourcing, and why?
Compare distributed transactions and log-based derived data on mechanism and on the one guarantee that genuinely differs.
Give the four limits of total ordering. Which one applies to microservices, and which to offline-capable clients?
Tell the unfriending story. Why is it fundamentally a join problem? Give the three partial mitigations and the limitation of each.
Why is asynchrony "what makes systems based on event logs robust"? Contrast with what distributed transactions do to a local fault.
Explain the railway gauge migration and map each step onto a data-system migration. What property makes it safe?
What was the lambda architecture, and what replaced it? List the three features required to unify batch and stream.
Walk through what `CREATE INDEX` does. Name the two other operations in this book that follow the same four steps.
Explain "the dataflow across an entire organization looks like one huge database." What are batch and stream processors, in that metaphor?
Distinguish federation from unbundling. Which tradition does each follow, and which problem is harder?
Give the two levels at which log-based integration provides loose coupling.
State the argument *against* unbundling. What is the goal of unbundling — and what is it explicitly not?
What did VisiCalc have in 1979 that data systems still lack? What three requirements make it hard to replicate?
Give four derived datasets and their derivation functions. Which are cookie-cutter and which require custom code?
Why are databases poor deployment environments for application code? Explain the Church-and-state joke.
Why is a database a "passive" mutable variable, and what would it take to make it active?
Explain the currency-conversion example. Why is the dataflow version both faster and more robust? What problem does it *not* remove?
Define the write path and the read path. Which is eager and which is lazy? What sits at the boundary?
Walk the full-text search spectrum from no-index to precompute-everything. Why is the far end impossible, and what's the practical middle?
What does it mean to say the pixels on screen are a materialized view? What does extending the write path to the device require, and why is offline already solved?
Explain "reads are events too." What join is being performed, and what distinguishes a one-off read from a subscription?
Why might you want to log read events, and what does it cost?
Trace the four layers at which duplicate suppression fails for a money transfer. At which layer does it actually break, and why can't the layer below fix it?
Write the request-ID transaction. Why does the uniqueness constraint work even at weak isolation, and what does the requests table give you for free?
State the end-to-end argument verbatim in substance. Apply it to duplicate suppression, integrity checking, and encryption.
If low-level reliability mechanisms can't provide end-to-end correctness, why keep them?
Describe enforcing unique usernames with a shared log. What general principle does it embody, and what replication model does it rule out?
Walk through the four steps of the multishard money transfer. **Where exactly does atomicity come from?** What are the three requirements?
What happens if the source processor crashes mid-request? Show why the outcome is still correct.
Define timeliness and integrity precisely. State the slogan. Which one is catastrophic when violated, and why?
Use the credit card statement to illustrate both. What would be a timeliness violation, and what an integrity violation?
Name the four mechanisms by which dataflow systems achieve integrity without distributed transactions.
Give three business situations where a "hard" constraint is deliberately violated. What is the forklift argument?
What is a compensating transaction? Give examples ordered by cost of apology.
State the two observations that combine into coordination-avoiding data systems. What guarantee do they keep, and what do they give up?
Explain the apology calculus. Why can't you drive apologies to zero?
Why does ACID consistency "make sense only if we assume the transaction is free from bugs"?
What do HDFS and S3 do that most systems don't? What is the corresponding advice about backups?
Why do event-based systems audit better than mutation logs? What can you check for the event log, and what for derived state?
Why is end-to-end integrity checking better than per-component checking? What does it implicitly cover?
What can a hardware-signed transaction log *not* guarantee? What do blockchains add, and why are they usually the wrong tool?
you're building a ride-hailing platform. Requirements: (a) a driver can accept at most one ride at a time; (b) surge pricing must be computed from live demand within 5 seconds; (c) the rider's app must show the car moving in real time and keep working through a tunnel; (d) daily payouts to drivers must be exactly correct, with a full audit trail; (e) the whole thing runs in 3 regions and must survive losing one. For each requirement: state whether you need **timeliness or integrity or both**, whether the constraint is **hard or loosely-interpretable**, which mechanism you'd use, where the **write path/read path boundary** sits, and what **audit** proves it's working. Identify the *one* place you'd accept synchronous coordination and justify why nothing else needs it.
# 13.6 Technology deep dives (/docs/ddia/philosophy-streaming-systems/technology-deep-dives)
***
#### 6.1 The end-to-end request ID (idempotency key) [#61-the-end-to-end-request-id-idempotency-key]
**Problem it solves.** Guarantee an operation takes effect exactly once **across every hop from the user's finger to durable storage** — the only place duplicate suppression can be complete.
**Why weren't the lower layers enough?** §5.1's four-layer trace: **TCP** covers one connection; **transactions** are tied to a connection and can't survive a client-side timeout; **2PC** fixes the coordinator↔database hop but not the human↔browser hop; **stream-processor exactly-once** covers only inside the framework.
**How it works internally.** The client (or a hash of the form fields) generates the ID **before the first attempt**, so a retry carries the same ID. Server-side, a **uniqueness constraint** on the ID column is the enforcement point — and crucially, **relational databases maintain uniqueness constraints correctly even at weak isolation levels**, whereas a hand-rolled check-then-insert does not (Ch 8's write skew). The row doubles as an **event log** for downstream derivation.
**Deployment.** ID minted client-side; propagated as a header through every service hop; stored with a **retention window** (long enough to outlive any client retry, e.g. 24–72 h); a scheduled purge job.
**Monitoring.** **Duplicate-suppression hit rate** (how often the constraint fires — a rising rate means clients are retrying more, i.e. something upstream is degrading); requests-table size and purge lag; **requests arriving with no ID** (a client that forgot to send one is a silent correctness hole).
**Scaling.** Shard by hash of the request ID so all attempts of one request land on one shard (§5.2).
**What actually breaks.**
* **The ID generated server-side**, so each retry gets a new one — the single most common way to implement this wrong. **It must be minted before the first attempt.**
* **Purging too aggressively**, so a slow client retry after the window duplicates.
* **The ID not propagated through an intermediate service**, breaking the chain at one hop.
* **A non-unique "unique" key** — e.g. hashing only *some* form fields, so two genuinely different requests collide and one is silently swallowed.
* **Trusting the database's uniqueness constraint** without knowing your database version's bug history (§5.6: **MySQL has failed to maintain uniqueness constraints**).
***
#### 6.2 Log-based multishard workflows (the money-transfer pattern) [#62-log-based-multishard-workflows-the-money-transfer-pattern]
**Problem it solves.** Multi-shard atomicity **without** an atomic commit protocol — §5.2's worked example.
**Why not 2PC?** Ch 8 §5.4's four problems, plus §5.2's throughput argument: **an atomic commit forces the transaction into a total order with respect to every other transaction on any participating shard**, so shards can no longer be processed independently.
**How it works internally.** The invariant is the sentence to memorize: **atomicity comes from the single atomic append of the initial request event.** Everything downstream is *eventual but inevitable*, made safe by (a) **strict per-shard log ordering**, (b) **at-least-once delivery**, (c) **deterministic processors**, and (d) **request-ID deduplication at every consumer.** Note the self-loop: the source processor **emits an event back into its own input log**, which is how the reserve→execute two-phase state change stays crash-safe without a transaction.
**Deployment.** One log shard per entity (account); one processor per shard with local state derived entirely from its log; a compacted state store; clients subscribe to the source shard's output for approval/decline.
**Monitoring.** **Reserved-but-not-executed balances** (money in limbo — the count and the *age* of the oldest; a growing age means a processor is stuck); per-shard consumer lag; duplicate-suppression rate; **a periodic reconciliation job asserting that debits and credits sum to zero** (§5.6's auditing, and the *only* thing that actually proves integrity).
**Scaling.** Linear in shards, because **no cross-shard coordination exists.** This is the whole point.
**What actually breaks.**
* **Nondeterminism in the processor** — a clock read, a random value, an external API call — so replay after a crash makes a *different* decision and the emitted events don't match. Every dataflow-correctness argument in this chapter rests on determinism.
* **A downstream consumer that forgets to deduplicate**, turning at-least-once into double-crediting.
* **Money stuck in "reserved"** because the outgoing event was lost or a processor is permanently stalled — this needs a timeout/sweeper, and it needs to be *monitored*, not assumed.
* **Reordering within a shard** (misconfigured partitioning, or parallel consumers on one partition).
* **No reconciliation job**, so an integrity violation goes undetected for months.
***
#### 6.3 The unbundled stack (Debezium + Kafka + Flink/Materialize + serving stores) [#63-the-unbundled-stack-debezium--kafka--flinkmaterialize--serving-stores]
**Problem it solves.** Compose specialized stores so that **writes stay in sync across all of them**, with faults contained.
**Why not federation alone?** §2.2: federation unifies **reads**; it **"does not have a good answer to synchronizing writes."**
**Why not distributed transactions?** §2.3: **no standardized protocol across systems written by different groups**, and synchronous coupling **"tends to escalate local faults into large-scale failures."**
**How it works internally.** CDC turns the system of record into the leader; the log imposes the total order; stream processors are the derivation functions; serving stores are the "index types." The whole thing is **`CREATE INDEX`, unbundled** (§2.1).
**Deployment.** System-of-record DB → Debezium → Kafka (compacted where views must be rebuildable) → Flink/Kafka Streams/Materialize → Elasticsearch / Redis / ClickHouse / a warehouse. Schema registry across the whole thing (Ch 5).
**Monitoring — the thing to instrument is *the pipeline*, not the pieces.**
* **End-to-end lag** per derived system (source commit → visible in the derived store) — the number users actually feel
* **Per-hop lag** so you can localize the stall
* **Divergence checks**: periodically compare row counts / checksums between the source and each derived store. §5.6: *"if we can check that an entire derived data pipeline is correct end to end, then any disks, networks, services, and algorithms along the path are implicitly included."*
* **Rebuild time** per derived system — your recovery budget, and it grows silently
* **Schema-change events** on the source (§Ch 12 §3.3: the schema is now a public API)
**Scaling.** Each hop scales independently — **that's the payoff of loose coupling**, at both the system and human level (§2.3).
**What actually breaks.**
* **The complexity itself.** §2.4 is blunt: **"building for scale that you don't need is wasted effort and may lock you into an inflexible design — in effect, a form of premature optimization."** Most teams that unbundle prematurely regret it.
* **Nobody owns the pipeline end to end** — each team monitors its own hop, and the *end-to-end* lag is nobody's metric.
* **Silent divergence** with no reconciliation job.
* **A source schema change** cascading into three downstream outages.
* **An unbounded rebuild time** discovered during an incident.
***
#### 6.4 Auditing and integrity verification (S3/HDFS scrubbing, Merkle trees, Certificate Transparency) [#64-auditing-and-integrity-verification-s3hdfs-scrubbing-merkle-trees-certificate-transparency]
**Problem it solves.** Detect corruption *before* it propagates — because §5.6 establishes that **hardware and software both fail in ways your system model excludes.**
**Why aren't checksums enough?** §5.1: **Ethernet/TCP/TLS checksums cannot detect corruption caused by software bugs at the endpoints, or corruption on disk.** Only end-to-end checks cover the whole path.
**How it works internally.** **Background scrubbing**: continually read back files, compare against replicas, migrate off suspect disks. **Merkle trees**: a tree of hashes giving O(log n) proofs that a record is in a dataset and O(log n) diffs between two replicas — which is also why anti-entropy repair (Ch 6) uses them. **Certificate Transparency**: an append-only Merkle log per CA, with **a single leader per log, which is how it avoids needing consensus.**
**Deployment.** A scheduled reconciliation/derivation-check job alongside the pipeline; **scheduled restore-from-backup tests** (§5.6: *"otherwise you may find out that your backup is broken when it is too late"*); optionally a redundant parallel derivation compared against the primary.
**Monitoring.** Scrub coverage (what fraction of data has been verified in the last N days — unverified data is *assumed* good, which is the failure mode); **divergence count and location**; **restore-test success and duration**; hash-mismatch alerts.
**What actually breaks.**
* **Backups that were never restore-tested** — the most common catastrophic failure in this entire book, and the cheapest to prevent.
* **Auditing that only checks the derived store against itself**, not against the source.
* **Corruption that predates every retained backup** — Ch 8's warning: *"if data has been corrupted for some time, replicas and recent backups may also be corrupted; you will need to restore from a historical backup."*
* **A reconciliation job that alerts but that nobody acts on**, which is the same as not having one.
* **Assuming blockchain-grade guarantees are needed** when a Merkle-tree diff and a nightly reconciliation would do — §5.6: *"for most applications, blockchains have too high an overhead to be useful."*
***
# 13.11 Terminology introduced here (/docs/ddia/philosophy-streaming-systems/terminology-introduced-here)
# 13.2 Unbundling Databases (/docs/ddia/philosophy-streaming-systems/unbundling-databases)
**The two philosophies, unresolved after 50 years:**
#### 2.1 `CREATE INDEX` is a batch job [#21-create-index-is-a-batch-job]
> **Think about what happens when you run `CREATE INDEX`: the database must**
>
> 1. **SCAN over a CONSISTENT SNAPSHOT of the table**
> 2. **Pick out the field values, SORT them, write out the index**
> 3. **PROCESS THE BACKLOG OF WRITES made since the snapshot was taken**
> 4. **CONTINUE to keep the index up to date whenever a transaction writes**
>
> ### **"This process is REMARKABLY SIMILAR TO SETTING UP A NEW FOLLOWER REPLICA, and also VERY SIMILAR TO BOOTSTRAPPING CDC in a streaming system. Whenever you run `CREATE INDEX`, THE DATABASE ESSENTIALLY REPROCESSES THE EXISTING DATASET AND DERIVES THE INDEX AS A NEW VIEW ONTO THE EXISTING DATA."** [#this-process-is-remarkably-similar-to-setting-up-a-new-follower-replica-and-also-very-similar-to-bootstrapping-cdc-in-a-streaming-system-whenever-you-run-create-index-the-database-essentially-reprocesses-the-existing-dataset-and-derives-the-index-as-a-new-view-onto-the-existing-data]
#### 2.2 The meta-database of everything [#22-the-meta-database-of-everything]
> ### **"In this light, THE DATAFLOW ACROSS AN ENTIRE ORGANIZATION STARTS LOOKING LIKE ONE HUGE DATABASE. Whenever a batch, stream, or ETL process transports data from one place and form to another, IT IS ACTING LIKE THE DATABASE SUBSYSTEM THAT KEEPS INDEXES OR MATERIALIZED VIEWS UP TO DATE."** [#in-this-light-the-dataflow-across-an-entire-organization-starts-looking-like-one-huge-database-whenever-a-batch-stream-or-etl-process-transports-data-from-one-place-and-form-to-another-it-is-acting-like-the-database-subsystem-that-keeps-indexes-or-materialized-views-up-to-date]
>
> **"Viewed like this, BATCH AND STREAM PROCESSORS ARE LIKE ELABORATE IMPLEMENTATIONS OF TRIGGERS, STORED PROCEDURES, AND MATERIALIZED VIEW MAINTENANCE ALGORITHMS. The derived data systems they maintain are LIKE DIFFERENT INDEX TYPES. Instead of implementing those facilities as features of a SINGLE INTEGRATED DATABASE PRODUCT, they are provided by VARIOUS PIECES OF SOFTWARE, RUNNING ON DIFFERENT MACHINES, ADMINISTERED BY DIFFERENT TEAMS."**
**Two avenues, and they are complementary:**
> **"Federated read-only querying requires MAPPING ONE DATA MODEL INTO ANOTHER, which takes some thought BUT IS ULTIMATELY QUITE A MANAGEABLE PROBLEM. KEEPING THE WRITES TO SEVERAL STORAGE SYSTEMS IN SYNC IS THE HARDER ENGINEERING PROBLEM."**
#### 2.3 Why log-based integration beats distributed transactions [#23-why-log-based-integration-beats-distributed-transactions]
> **"Transactions WITHIN a single storage or stream processing system are feasible, BUT WHEN DATA CROSSES THE BOUNDARY BETWEEN DIFFERENT TECHNOLOGIES, AN ASYNCHRONOUS EVENT LOG WITH IDEMPOTENT WRITES IS A MUCH MORE ROBUST AND PRACTICABLE APPROACH."**
>
> **"Distributed transactions ARE used within some stream processors to achieve exactly-once semantics, and this can work quite well. HOWEVER, WHEN A TRANSACTION WOULD NEED TO INVOLVE SYSTEMS WRITTEN BY DIFFERENT GROUPS OF PEOPLE, THE LACK OF A STANDARDIZED TRANSACTION PROTOCOL MAKES INTEGRATION MUCH HARDER."**
**Loose coupling manifests in two ways:**
| Level | Benefit |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **System** | **"Asynchronous event streams make the system AS A WHOLE more robust to outages or performance degradation of individual components. If a consumer runs slow or fails, THE EVENT LOG CAN BUFFER MESSAGES, allowing the producer and other consumers to CONTINUE RUNNING UNAFFECTED. The faulty consumer CAN CATCH UP WHEN IT IS FIXED, so it doesn't miss any data, AND THE FAULT IS CONTAINED. By contrast, THE SYNCHRONOUS INTERACTION OF DISTRIBUTED TRANSACTIONS TENDS TO ESCALATE LOCAL FAULTS INTO LARGE-SCALE FAILURES"** |
| **Human** | **"Unbundling allows software components and services to be DEVELOPED, IMPROVED, AND MAINTAINED INDEPENDENTLY BY DIFFERENT TEAMS. Specialization allows each team to FOCUS ON DOING ONE THING WELL, with well-defined interfaces. Event logs provide an interface POWERFUL ENOUGH to capture fairly strong consistency properties (because of durability and ordering), BUT ALSO GENERAL ENOUGH to be applicable to almost any kind of data"** |
#### 2.4 The honest caveat — don't unbundle prematurely [#24-the-honest-caveat--dont-unbundle-prematurely]
> **"If unbundling does become the way of the future, IT WILL NOT REPLACE DATABASES IN THEIR CURRENT FORM. They will still be needed AS MUCH AS EVER, for maintaining state in stream processors and to SERVE QUERIES for the output of batch and stream processors."**
>
> **"The COMPLEXITY OF RUNNING SEVERAL PIECES OF INFRASTRUCTURE CAN BE A PROBLEM. Each piece has a LEARNING CURVE and its own CONFIGURATION ISSUES AND OPERATIONAL QUIRKS, so IT IS WORTH DEPLOYING AS FEW MOVING PARTS AS POSSIBLE. A single integrated product may also achieve BETTER AND MORE PREDICTABLE PERFORMANCE on the workloads it's designed for."**
>
> ### **"BUILDING FOR SCALE THAT YOU DON'T NEED IS WASTED EFFORT AND MAY LOCK YOU INTO AN INFLEXIBLE DESIGN. IN EFFECT, IT IS A FORM OF PREMATURE OPTIMIZATION."** [#building-for-scale-that-you-dont-need-is-wasted-effort-and-may-lock-you-into-an-inflexible-design-in-effect-it-is-a-form-of-premature-optimization]
>
> ### **"THE GOAL OF UNBUNDLING IS NOT TO COMPETE WITH INDIVIDUAL DATABASES ON PERFORMANCE FOR PARTICULAR WORKLOADS; THE GOAL IS TO ALLOW YOU TO COMBINE SEVERAL DATABASES IN ORDER TO ACHIEVE GOOD PERFORMANCE FOR A MUCH WIDER RANGE OF WORKLOADS. IT'S ABOUT BREADTH, NOT DEPTH."** [#the-goal-of-unbundling-is-not-to-compete-with-individual-databases-on-performance-for-particular-workloads-the-goal-is-to-allow-you-to-combine-several-databases-in-order-to-achieve-good-performance-for-a-much-wider-range-of-workloads-its-about-breadth-not-depth]
>
> **"Thus, IF A SINGLE TECHNOLOGY DOES EVERYTHING YOU NEED, YOU'RE MOST LIKELY BEST OFF SIMPLY USING THAT PRODUCT. The advantages of unbundling come into the picture ONLY WHEN NO SINGLE PIECE OF SOFTWARE SATISFIES ALL YOUR REQUIREMENTS."**
***
# 13.9 Worked examples (/docs/ddia/philosophy-streaming-systems/worked-examples)
**① Why asynchrony contains faults, quantified.** A write must reach 5 systems, each with 99.9% availability.
* **Distributed transaction (all must commit):** availability = 0.999⁵ = **99.5%** → \~44 hours of downtime/year, and **any one system's outage stops all writes.**
* **Log-based derived (write to the log only):** the write path depends on 1 system = **99.9%**; a failing consumer only makes *its own* view stale, and it **catches up when fixed**.
This is §1.5's "asynchrony is what makes systems based on event logs robust," as a number.
**② The apology calculus.** An airline with 200 seats. Strict constraint: 0 overbookings, but \~8% no-show rate → **16 empty seats per flight**, \~$4,800 of lost revenue. Overbooking by 5%: \~10 extra seats sold (+$3,000), and the probability that all 210 show up is small; when it happens, compensation costs \~$800 per bumped passenger. **Expected value strongly favours the loose constraint** — and, critically, **the compensation process must exist anyway** for weather cancellations. This is why §5.4's argument isn't a hack; it's how the business already works.
**③ The four-layer duplicate trace.** Probability a $11 transfer becomes $22:
* P(client-side timeout after commit) ≈ 0.1% of requests on a poor mobile connection
* P(user retries | error shown) ≈ 60%
⇒ **\~0.06% of transfers double-charge** — 6 in 10,000. At 100,000 transfers/day that's **60 incorrect transfers per day**, every one of them a customer complaint. **A single request-ID column removes all of them.**
**④ Where the write/read boundary goes.** 10M documents, 1,000 distinct common queries, 50,000 queries/s of which 80% are the common ones.
* **No index:** read cost = scan 10M docs × 50,000/s. Impossible.
* **Index only:** write cost = update terms per document; read cost = per-query Boolean evaluation × 50,000/s.
* **Index + cache of the 1,000 common queries:** 40,000 q/s served from cache at near-zero cost; 10,000 q/s hit the index. **Write cost rises by 1,000 materialized-view updates per relevant document change.**
The right answer depends entirely on **write:read ratio** — and the celebrity insight is that **within one system, different keys may deserve different answers.**
**⑤ Integrity vs timeliness, priced.** A bank's ledger.
* **Timeliness violation:** a transaction doesn't appear for 24 h. Cost: **a support call**. Self-healing.
* **Integrity violation:** debits ≠ credits by $1. Cost: **a full audit, regulatory exposure, and manual reconstruction — and it does not self-heal.**
Even a 1-in-10⁶ integrity failure at 10M transactions/day is **10 per day**, each requiring human investigation. **This is why the reconciliation job is not optional.**
**⑥ Auditing coverage.** 500 TB across 3 replicas; a scrubber reading at 200 MB/s per node.
Full pass = 500e12 / (200e6 × 3) ≈ **833,000 s ≈ 9.6 days.** So **any given block is verified roughly every 10 days**; a corruption introduced today is detected in **\~5 days on average**. If your backup retention is 7 days, **you have \~2 days of margin** to restore a clean copy. **Lengthen retention or speed up scrubbing — and know which you're relying on.**
***
# 6.0 Backups vs replication — they are NOT the same thing (/docs/ddia/replication/backups-vs-replication-not)
> **Replicas quickly reflect writes from one node on other nodes. Backups store OLD SNAPSHOTS so you can go back in time.**
>
> **If you accidentally delete some data, replication doesn't help — the deletion is propagated to the replicas too.** You need a backup to restore it.
**They're complementary:**
* **Backups are often part of the process of setting up replication** (§1.2)
* **Archiving replication logs can be part of a backup process**
**Internal snapshots aren't enough either.** Some databases maintain immutable snapshots of past states — a kind of internal backup — **but this keeps old versions on the same storage medium as the current state.** With a large amount of data it's **cheaper to keep backups of old data in an object store optimized for infrequently accessed data, and store only the current state in primary storage.**
***
# 6.8 Decision cheat sheet (/docs/ddia/replication/decision-cheat-sheet)
**Which replication model?**
**Sync, semisync, or async?**
**Never all-synchronous** (one outage halts everything). **Semisynchronous** — one sync follower + async others — is the default correct answer: durability on two nodes, availability preserved. Pure async only when you can genuinely afford to lose the last few seconds of writes.
**Automatic or manual failover?**
Automatic if your failover manager does **real fencing** and your timeout is tuned. Manual if not — **some operations teams prefer manual even when automatic is supported**, and that is a defensible engineering position, not laziness.
**Which replication log format?**
**Logical/row-based**, unless you specifically need byte-identical standbys. It decouples you from the storage format, enables cross-version upgrades, and gives you CDC for free.
**How do I fix a replication-lag anomaly?**
* read-your-writes → read from the leader for *that user's own data*, or track a write timestamp/LSN
* monotonic reads → **pin each user to one replica by hash of user ID**
* consistent prefix → **co-locate causally related writes in one shard**, or track causality explicitly
**Choosing n, w, r?**
Start `n=3, w=2, r=2` (or `n=5, w=3, r=3`). Read-heavy with rare writes → `w=n, r=1` **only if you accept that one dead node stops all writes.** Multi-region → **`LOCAL_QUORUM` by default**, `EACH_QUORUM` only where correctness demands it. **And treat `w+r>n` as a probability adjustment, not a guarantee.**
**LWW, manual, or CRDT?**
LWW **only if you never update existing records** (insert-only with unique keys). Manual if conflicts are rare and the domain genuinely needs a human. **CRDT/OT for anything collaborative** — and accept that invariants over the merged state are not enforceable.
***
# 6.5 Detecting Concurrent Writes (/docs/ddia/replication/detecting-concurrent-writes)
**Conflicts might be detected as the writes happen — but not always; they could also be detected later, during read repair, hinted handoff, or anti-entropy.**
**To become eventually consistent, replicas must converge** — using any of §3.4's mechanisms: **LWW (Cassandra, ScyllaDB), manual resolution, or CRDTs (Riak).**
> **LWW is easy to implement. But a timestamp DOESN'T TELL YOU WHETHER TWO VALUES ARE ACTUALLY CONFLICTING (written concurrently) or not (written one after another). To resolve conflicts explicitly, the system must take more care to detect concurrent writes.**
#### 5.1 The happens-before relation [#51-the-happens-before-relation]
> **An operation A HAPPENS BEFORE another operation B if B knows about A, or depends on A, or builds upon A in some way.**
>
> ### **Two operations are CONCURRENT if neither happens before the other.** [#two-operations-are-concurrent-if-neither-happens-before-the-other]
**Three possibilities for any A and B: A happened before B, B happened before A, or they are concurrent.**
* **§3.2's overtaking example:** the insert **happens before** the increment, **because the value incremented by B is the value inserted by A** — B **builds upon** A, so B must have happened later. B is **causally dependent** on A.
* **§5's example:** the writes **are concurrent** — when each client starts, **it doesn't know another client is also operating on that key. There is no causal dependency.**
> **If one operation happened before another, the later one should OVERWRITE the earlier. If the operations are concurrent, we have a CONFLICT that needs to be resolved.**
**Concurrency, time, and relativity:**
> **It is NOT important whether they literally overlap in time.** Because of clock problems, **it's quite difficult to tell whether two things happened at exactly the same time** (Ch 9). **For defining concurrency, exact time doesn't matter — two operations are concurrent if they are both UNAWARE OF EACH OTHER, regardless of physical time.**
>
> People connect this to **special relativity**: information cannot travel faster than light, so **two events some distance apart cannot affect each other if the time between them is shorter than light's travel time.** **In computer systems, two operations might be concurrent EVEN THOUGH the speed of light would in principle have allowed one to affect the other** — if the network was slow or interrupted, **two operations can occur some time apart and still be concurrent, because network problems prevented one from knowing about the other.**
#### 5.2 The algorithm (single replica) [#52-the-algorithm-single-replica]
1. **The server maintains a version number per key**, increments it on every write, and **stores the new version number along with the value.**
2. **On read, the server returns ALL SIBLINGS** — all values not overwritten — **plus the latest version number.** **A client must READ a key BEFORE writing.**
3. **On write, the client must include the version number from the prior read, and must MERGE together all values it received in that read.** The write response **also returns all siblings**, allowing several writes to be chained.
4. **On receiving a write with a particular version number, the server can OVERWRITE all values with that version number OR BELOW** (it knows they've been merged into the new value) **but must KEEP all values with a HIGHER version number** (those are concurrent with the incoming write).
> **The server can determine whether two operations are concurrent JUST BY LOOKING AT VERSION NUMBERS. It does not need to interpret the value itself — so the value could be any data structure.**
>
> **A write WITHOUT a version number is concurrent with all other writes, so it will not overwrite anything — it will just be returned as one of the values on subsequent reads.**
#### 5.3 The shopping cart trace — follow this carefully [#53-the-shopping-cart-trace--follow-this-carefully]
> **In this example, the clients are NEVER fully up to date with the data on the server, since there is always another operation going on concurrently. But old versions DO get overwritten eventually, and NO WRITES ARE LOST.**
#### 5.4 Version vectors [#54-version-vectors]
**A single version number is not sufficient when there are MULTIPLE REPLICAS accepting writes concurrently.**
> **Instead, use a version number PER REPLICA as well as per key. Each replica increments its own version number when processing a write, and also keeps track of the version numbers it has seen from each of the other replicas. This information indicates which values to overwrite and which to keep as siblings.**
>
> **The collection of version numbers from all the replicas is called a VERSION VECTOR.**
**The most interesting variant is the DOTTED VERSION VECTOR, used in Riak 2.0.**
**Like the single version numbers, version vectors are sent from replicas to clients on read and must be sent back on write.** *(Riak encodes it as a string it calls **causal context**.)* **The version vector lets the database distinguish between overwrites and concurrent writes.**
> **The version vector also ensures that it is SAFE TO READ FROM ONE REPLICA AND SUBSEQUENTLY WRITE BACK TO ANOTHER. Doing so may create siblings, but NO DATA IS LOST as long as siblings are merged correctly.**
> ⚠️ **Version vector ≠ vector clock.** They're sometimes conflated. **The difference is subtle — in brief, when comparing the STATE OF REPLICAS, version vectors are the right data structure to use.**
***
# 6.12 Forward links (/docs/ddia/replication/forward-links)
| Concept here | Where it's developed |
| ------------------------------------------------------------- | ------------------------------------------- |
| Sharding a dataset too big for one machine | **Ch 7** — Sharding |
| Request routing to the current leader | **Ch 7** |
| Serializable transactions, write skew, phantoms | **Ch 8** — Transactions |
| Why timeouts can't distinguish slow from dead; clock problems | **Ch 9** — Trouble with Distributed Systems |
| Distributed locks, leases, and fencing tokens | **Ch 9** |
| Consensus, leader election, linearizability | **Ch 10** |
| Logical clocks and ID generators | **Ch 10** |
| Using shared logs / state machine replication | **Ch 10** |
| Change data capture in depth | **Ch 12** — Stream Processing |
| Detecting and resolving conflicts at scale | **Ch 13** |
# 6. Replication (/docs/ddia/replication)
> "The major difference between a thing that might go wrong and a thing that cannot possibly go wrong is that when a thing that cannot possibly go wrong goes wrong, it usually turns out to be impossible to get at or repair." — Douglas Adams
**Replication = keeping a copy of the same data on multiple machines connected via a network.**
**Why:**
* **Latency** — keep data geographically close to users
* **Availability & durability** — keep working even if parts have failed
* **Read throughput** — scale out the number of machines serving read queries
> **If the data doesn't change, replication is easy — copy it once and you're done. ALL the difficulty lies in handling CHANGES to replicated data.**
**Assumption for this chapter:** the dataset is small enough that **each machine holds a copy of the entire dataset.** Ch 7 relaxes this (sharding).
**Three families of algorithms — almost all distributed databases use one of these three:**
```txt
SINGLE-LEADER MULTI-LEADER LEADERLESS
───────────── ──────────── ──────────
client ──▶ [L] client ──▶ [L₁] [L₂] ◀── client client ──▶ [R] [R] [R]
│ │ │ ╲ ╱ ◀── parallel
▼ ▼ ▼ ╲ ╱ writes AND reads
[F][F][F] ╳ to several nodes
╱ ╲
ONE node orders writes each leader also acts as NO leader; no ordering
followers apply in the a FOLLOWER to the others imposed; clients detect
SAME order and correct stale nodes
```
> **The principles haven't changed much since the 1970s, because the fundamental constraints of networks have remained the same.** Nevertheless, concepts such as **eventual consistency still cause confusion.**
***
# 6.4 Leaderless Replication (/docs/ddia/replication/leaderless-replication)
**Abandon the leader entirely — any replica directly accepts writes from clients.**
**History:** some of the **earliest** replicated data systems were leaderless, **but the idea was mostly forgotten during the era of relational-database dominance.** It became fashionable again after **Amazon used it for its in-house Dynamo system in 2007.** **Riak, Cassandra, ScyllaDB** are open source Dynamo-inspired datastores — hence **"Dynamo-style."**
> ⚠️ **The original Dynamo was described in a paper but never released outside Amazon. The similarly named DynamoDB has a COMPLETELY DIFFERENT architecture: single-leader replication based on the Multi-Paxos consensus algorithm.**
In some implementations **the client sends writes to several replicas directly**; in others **a coordinator node does this on the client's behalf.** **Unlike a leader, that coordinator does NOT enforce a particular ordering of writes** — and **this difference has profound consequences.**
#### 4.1 Writing when a node is down [#41-writing-when-a-node-is-down]
```txt
3 replicas, one down for a reboot. NO FAILOVER — all replicas are equal.
user 1234 WRITE ──┬──▶ [replica 1] ok ─┐
├──▶ [replica 2] ok ─┼─ 2 of 3 acknowledged ⇒ SUCCESS
└──▶ [replica 3] ✗ ┘ (client simply IGNORES the missed replica)
(down)
… replica 3 comes back, MISSING the write …
user 2345 READ ───┬──▶ [replica 1] → version 7 ─┐
├──▶ [replica 2] → version 7 ─┼─ client takes the value with the
└──▶ [replica 3] → version 6 ─┘ GREATEST VERSION/TIMESTAMP
(even if returned by only ONE node)
│
└──▶ READ REPAIR: client writes version 7 back to replica 3
```
> **Read requests are also sent to SEVERAL NODES IN PARALLEL.** Every value written must be **tagged with a version number or timestamp** so the client can tell which responses are up to date.
#### 4.2 Three catch-up mechanisms [#42-three-catch-up-mechanisms]
| Mechanism | How it works | Coverage |
| ------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ |
| **Read repair** | A client reading from several nodes in parallel **detects stale responses and writes the newer value back to the stale replica** | **Works well for values that are READ OFTEN** |
| **Hinted handoff** | If a replica is unavailable, **another replica stores writes on its behalf as HINTS.** When the intended replica returns, **the hint-storing replica sends them over and deletes the hints** | **Covers values that are NEVER READ**, which read repair misses |
| **Anti-entropy** | A **background process periodically looks for differences between replicas and copies missing data** | **Does NOT copy writes in any particular order, and there may be a SIGNIFICANT DELAY before data is copied** |
#### 4.3 Quorums [#43-quorums]
> **With n replicas: every write must be confirmed by w nodes; every read must query at least r nodes.**
>
> ### **As long as w + r > n, we expect to get an up-to-date value when reading, because at least one of the r nodes we read from must be up to date.** [#as-long-as-w--r--n-we-expect-to-get-an-up-to-date-value-when-reading-because-at-least-one-of-the-r-nodes-we-read-from-must-be-up-to-date]
**Common choice:** **n odd (3 or 5), w = r = (n+1)/2 rounded up.** But you can vary them: **a workload with few writes and many reads may set w = n, r = 1 — faster reads, but ONE failed node causes ALL writes to fail.**
**Tolerance:**
| Config | Tolerates |
| ----------------- | ---------------------------------------- |
| w \< n | writes continue if a node is unavailable |
| r \< n | reads continue if a node is unavailable |
| **n=3, w=2, r=2** | **1 unavailable node** |
| **n=5, w=3, r=3** | **2 unavailable nodes** |
**Mechanics:** **reads and writes are normally sent to ALL n replicas in parallel; w and r determine how many you WAIT FOR.** If fewer than w or r are available, the operation returns an error. **A node could be unavailable for many reasons — crashed, powered down, disk full, network interruption — and we care only whether it returned a successful response.**
*Note: **there may be more than n nodes in the cluster, but any given value is stored on only n nodes** — which is what allows sharding (Ch 7).*
**Quorums need not be majorities.** **It matters only that the read and write sets OVERLAP in at least one node.** Majorities are common because **r = w = majority ensures w + r > n while tolerating up to ⌊n/2⌋ failures**, but other assignments allow flexibility in algorithm design.
**Deliberately violating the quorum condition (w + r ≤ n):** reads and writes still go to n nodes, but fewer successes are required.
* ✗ **More likely to read stale values**
* ✔ **Lower latency** (particularly beneficial with synchronous replication)
* ✔ **More highly available** — during a network interruption **there's a higher chance you can continue.** The database becomes unavailable **only when reachable replicas fall below w or r**
#### 4.4 The six ways quorums lie to you [#44-the-six-ways-quorums-lie-to-you]
> **Although quorums APPEAR to guarantee that a read returns the latest written value, in practice it is not so simple.**
> **Dynamo-style databases are generally optimized for use cases that can tolerate eventual consistency. The parameters w and r let you ADJUST THE PROBABILITY of stale reads — but it's wise NOT TO TAKE THEM AS ABSOLUTE GUARANTEES.**
#### 4.5 Monitoring staleness [#45-monitoring-staleness]
**Leader-based:** easy. **Writes are applied to leader and followers IN THE SAME ORDER, and each node has a position in the replication log.** Subtract the follower's position from the leader's ⇒ **replication lag**, exposed as a metric.
**Leaderless:** hard. **There is NO FIXED ORDER in which writes are applied.** **The number of hints a replica stores for handoff can be one measure of system health, but it's difficult to interpret usefully.**
> **Eventual consistency is a deliberately vague guarantee, but for OPERABILITY it's important to be able to QUANTIFY "eventual."**
#### 4.6 Single-leader vs leaderless performance [#46-single-leader-vs-leaderless-performance]
**Reading from the leader ensures up-to-date responses but has three performance problems:**
1. **Read throughput is limited by the leader's capacity**
2. **On leader failure you must wait for detection and failover.** Even a quick failover is noticed by users as increased response times; a long one means **the system is unavailable for its duration**
3. **Very sensitive to performance problems on the leader** — if the leader is slow (overload, resource contention), **increased response times immediately affect users**
**The leaderless advantage:**
> **Because there is no failover, and requests go to multiple replicas in parallel anyway, one replica becoming slow or unavailable has very little impact on response times — the client simply uses the responses from the faster replicas.** Using the fastest responses is called **REQUEST HEDGING**, and it **can significantly reduce tail latency.**
>
> **At its core, the resilience of a leaderless system comes from the fact that IT DOESN'T DISTINGUISH BETWEEN THE NORMAL CASE AND THE FAILURE CASE.**
**This is especially helpful for GRAY FAILURES** — a node that **isn't completely down but is running in a degraded state, unusually slow to handle requests** — or a node that's simply overloaded (e.g. **recovery via hinted handoff can cause a lot of additional load**). **A leader-based system has to DECIDE whether the situation is bad enough to warrant a failover (which can itself cause further disruption); in a leaderless system that question doesn't even arise.**
**But leaderless has its own performance problems:**
1. **One replica must still detect when another is unavailable to store hints, and the handoff must send them — putting additional load on the replicas AT A TIME WHEN THE SYSTEM IS ALREADY UNDER STRAIN**
2. **The more replicas, the bigger the quorums and the more responses to wait for.** Even waiting only for the fastest r or w, **a bigger r or w raises the chance of hitting a slow replica**, increasing overall response time. **In practice, quorums are seldom more than 4 of 7 or 5 of 9 nodes**
3. **A large-scale network interruption disconnecting a client from many replicas can make it impossible to form a quorum**
**Sloppy quorums** are the escape hatch: **allow any reachable replica to accept writes even if it's not one of the usual n replicas for that key.** *(Riak/Dynamo: "sloppy quorum"; Cassandra/ScyllaDB: consistency level `ANY`.)* **There is NO GUARANTEE that subsequent reads will see the written value, but depending on the application it may still be better than having the write fail.**
**The three-way summary:**
| | Resilience to network interruption | Staleness risk |
| ----------------------- | --------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| **Multi-leader** | **Greatest** — reads and writes need only ONE leader, which can be co-located with the client | **Reads can be ARBITRARILY out of date** |
| **Leaderless (quorum)** | Good | **Good fault tolerance AND a high likelihood of reading up-to-date data — a compromise** |
| **Single-leader** | Weakest | Strongest consistency available |
#### 4.7 Multi-region with leaderless [#47-multi-region-with-leaderless]
**Cassandra / ScyllaDB:** the client picks a node in its **local region — the coordinator node** — and sends the write there. **The coordinator forwards to all replicas in its OWN region AND TO ONE REPLICA IN EVERY OTHER REGION, which then forwards to the other replicas in that region. This optimization avoids making the cross-region request multiple times.**
**Consistency levels** determine how many responses are required: **a quorum across all regions, a separate quorum in each region, or a quorum only in the client's local region.** **A LOCAL quorum avoids waiting for slow cross-region requests but is more likely to return stale results.**
**Riak** keeps all client↔node communication **local to one region** (so **n describes replicas within one region**); **cross-region replication happens asynchronously in the background, in a style similar to multi-leader replication.**
***
# 6.3 Multi-Leader Replication (/docs/ddia/replication/multi-leader-replication)
**Also: active/active or bidirectional replication.** More than one node accepts writes; **each leader simultaneously acts as a follower to the other leaders.**
**Why the book only discusses the asynchronous variant:** with two leaders A and B, **if writes are synchronously replicated from A to B and the network between them is interrupted, you can't write to A until the connection is restored. Synchronous multi-leader is therefore very similar to single-leader** (equivalent to making B the leader and having A forward writes). **The rest of the section is asynchronous multi-leader, in which any leader can process writes even when its connection to the other leaders is interrupted.**
#### 3.1 Geographically distributed operation [#31-geographically-distributed-operation]
> **It rarely makes sense to use a multi-leader setup within a SINGLE region — the benefits rarely outweigh the added complexity.**
**Four-way comparison against single-leader in a multi-region deployment:**
| Dimension | Single-leader | Multi-leader |
| --------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------- |
| **Performance** | **Every write goes over the internet to the leader's region** — significant added latency, **might defeat the purpose of having multiple regions** | **Every write is processed in the LOCAL region** and replicated asynchronously. **The inter-region delay is hidden from users** |
| **Tolerance of regional outages** | Failover can promote a follower in another region | **Each region continues operating INDEPENDENTLY; replication catches up when the offline region returns** |
| **Tolerance of network problems** | **Very sensitive to the inter-region link** — a client must send its request over that link and wait for the response | **Can tolerate network problems better;** during a temporary interruption **each region's leader continues independently** |
| **Consistency** | **Can provide strong guarantees, e.g. serializable transactions** | **MUCH WEAKER — the biggest downside.** |
**The consistency limitation, stated precisely:**
> **You can't guarantee that a bank account won't go negative or that a username is unique; it's always possible for different leaders to process writes that are INDIVIDUALLY FINE (paying out some of the money in an account, registering a particular username) but that VIOLATE THE CONSTRAINT WHEN TAKEN TOGETHER with another write on another leader.**
>
> **This is simply a fundamental limitation of distributed systems. If you need to enforce such constraints, you're better off with a single-leader system.**
**Support:** MySQL, Oracle, SQL Server, YugabyteDB; as an external add-on in **Redis Enterprise, EDB Postgres Distributed, pglogical.**
> ⚠️ **As multi-leader replication is a RETROFITTED feature in many databases, there are often subtle configuration pitfalls and surprising interactions with other database features. Autoincrementing keys, triggers, and integrity constraints can be problematic. For this reason, multi-leader replication is often considered DANGEROUS TERRITORY that should be avoided if possible.**
#### 3.2 Replication topologies [#32-replication-topologies]
```txt
(a) CIRCULAR (b) STAR (c) ALL-TO-ALL
(most general)
L1 ──▶ L2 L2 L1 ◀──▶ L2
▲ │ ▲ ▲ ╲ ╱ ▲
│ ▼ L1 ◀──R──▶ L3 │ ╲ ╱ │
L4 ◀── L3 │ │ ╳ │
▼ │ ╱ ╲ │
each node receives L4 ▼ ╱ ╲ ▼
from ONE node and L4 ◀──▶ L3
forwards to ONE one designated ROOT
other node forwards to all others every leader sends its
(generalizable to a TREE) writes to every other
```
**Loop prevention (needed in circular and star, where a write passes through several nodes):** **each node has a unique identifier, and in the replication log each write is tagged with the identifiers of all nodes it has passed through. When a node receives a change tagged with its OWN identifier, it ignores it** — it knows it already processed it.
**Problems with each:**
| Topology | Problem |
| ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Circular / star** | **If just ONE node fails, it can interrupt the flow of replication messages between other nodes**, leaving them unable to communicate until it's fixed. The topology **could be reconfigured to work around the failed node, but in most deployments that must be done MANUALLY** |
| **All-to-all** | Better fault tolerance (**messages travel along different paths, avoiding a single point of failure**) — **but some network links may be faster than others**, so **replication messages can OVERTAKE others** |
**The overtaking problem — a causality violation:**
> **Simply attaching a TIMESTAMP to every write is NOT SUFFICIENT, because clocks cannot be trusted to be sufficiently in sync to correctly order these events** (Ch 9).
>
> **To order these events correctly, VERSION VECTORS can be used** (§4.6). **However, many multi-leader replication systems don't use good techniques for ordering updates, leaving them vulnerable to this issue. It's worth carefully reading the documentation and THOROUGHLY TESTING your database to ensure it really does provide the guarantees you believe it has.**
#### 3.3 Sync engines and local-first software [#33-sync-engines-and-local-first-software]
**The other major multi-leader use case: applications that must work while disconnected from the internet.** Calendar apps on phone, laptop, and other devices — you must be able to read and write **at any time, regardless of connectivity**, and changes sync when next online.
> **Every device has a local database replica that acts as a LEADER (it accepts writes), and there is an asynchronous multi-leader replication process (SYNC) between all your devices. The replication lag may be HOURS OR EVEN DAYS.**
>
> **Architecturally this is multi-leader replication between regions, taken to the extreme: each DEVICE is a "region," and the network connection between them is extremely unreliable.**
**Real-time collaboration is the same architecture.** Google Docs/Sheets, Figma, Linear. **What makes them responsive is that user input is immediately reflected in the UI without waiting for a network round-trip, and edits are shown to collaborators with low latency.**
> **Each browser tab that has opened the shared file is a REPLICA.** And note: **even if the app does NOT allow offline editing, the fact that multiple users can make edits without waiting for a server response ALREADY MAKES IT MULTI-LEADER.**
**Both offline editing and real-time collaboration need the same infrastructure:** capture the user's changes and **either send them immediately (online) or store locally for later (offline)**; **receive changes from collaborators, merge them into the local copy, and update the UI.** If multiple users changed the file concurrently, **conflict resolution logic is needed.**
**The vocabulary:**
| Term | Meaning |
| ----------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Sync engine** | A software library supporting this process. *The idea is old; the term has recently gained attention* |
| **Offline-first** | An app that allows the user to continue editing while offline |
| **Local-first** | Collaborative apps that are not only offline-first but **designed to keep working even if the DEVELOPER SHUTS DOWN ALL THEIR ONLINE SERVICES** — achieved with a sync engine using an **open standard sync protocol with multiple available providers.** **Git is a local-first collaboration system** (albeit without real-time collaboration) — you can sync via GitHub, GitLab, or any other host |
| **Netcode** | The multiplayer-video-game equivalent. **Techniques are quite specific to games and don't directly carry over** |
**Four advantages of sync engines** (vs the dominant model of keeping little state on the client and hitting a server for everything):
1. **Speed.** Local data means the UI responds much faster. **Some apps aim to respond to input in the NEXT FRAME — rendering within 16 ms on a 60 Hz display.**
2. **Offline.** **An app doesn't need a separate offline mode: being offline is the same as having a very large network delay.**
3. **Simpler programming model.** **Every service call requires error handling** (Ch 5 §7.3) — if an update request fails, the UI must reflect that error. **A sync engine lets the app read and write LOCAL data; these operations almost never fail, leading to a more DECLARATIVE programming style.**
4. **Real-time updates.** **A sync engine combined with a reactive programming model** is a good way to receive edit notifications and update the UI.
**The limitation:**
> **Sync engines work best when ALL the data the user may need is downloaded in advance and stored persistently on the client.** Fine for **all the files a user created** (one user doesn't generate that much data); **not suitable for the entire catalog of an ecommerce website.**
**History and implementations:** pioneered by **Lotus Notes in the 1980s** (without the term). Today: proprietary backends (**Google Firestore, Realm, Ditto**) and open source backends suitable for local-first software (**PouchDB/CouchDB, Automerge, Yjs**).
#### 3.4 Dealing with conflicting writes [#34-dealing-with-conflicting-writes]
> **The biggest problem with multi-leader replication — both in a geo-distributed server-side database and a local-first sync engine — is that concurrent writes on different leaders lead to conflicts that need to be resolved.**
```txt
USER 1 USER 2
sets title A → B sets title A → C
│ │
▼ ▼
[LEADER 1] ═══════ async replication ═══════ [LEADER 2]
│ │
└──────────────▶ CONFLICT DETECTED ◀─────────┘
(This problem does not occur in a single-leader database.)
```
> **Definition of concurrent:** the two writes are **concurrent because NEITHER WAS "AWARE" OF THE OTHER at the time it was originally made. It doesn't matter whether they literally happened at the same time** — if made while offline, they might have happened **some time apart. What matters is whether one write occurred in a state where the other had already taken effect.**
##### Strategy 1 — Conflict avoidance [#strategy-1--conflict-avoidance]
**Ensure all writes for a particular record go through the same leader.** Then **conflicts cannot occur even if the database as a whole is multi-leader.**
* **Not possible for a sync engine client being updated offline**, but **sometimes possible in geo-replicated server systems**
* Example: in an app where **a user can edit only their own data**, route requests from a particular user **always to the same region.** Different users have different "home" regions (picked by geographic proximity); **from any one user's point of view, the configuration is essentially single-leader**
* Another example: **odd/even autoincrement** — with two leaders, one generates only odd numbers, the other only even, so **they can't concurrently assign the same ID**
> ⚠️ **Conflict avoidance BREAKS DOWN if you allow the leader to be changed.** You might want to change the designated leader — a region is unavailable, or a user moved closer to a different region — **and there is now a risk the user performs a write WHILE THE CHANGE IS IN PROGRESS, leading to a conflict.**
##### Strategy 2 — Last write wins (LWW) [#strategy-2--last-write-wins-lww]
**Attach a timestamp to each write; always use the value with the greatest timestamp.** Ties broken by comparing the values (e.g. alphabetically earliest string).
> **The term is misleading: when two writes are CONCURRENT, which one is most recent is UNDEFINED, so the timestamp order of concurrent writes is essentially RANDOM.**
>
> **The real meaning of LWW: when the same record is concurrently written on different leaders, ONE OF THOSE WRITES IS RANDOMLY CHOSEN AS THE WINNER AND THE OTHERS ARE SILENTLY DISCARDED, even though they were successfully processed by their respective leaders. This achieves eventual consistency AT THE COST OF DATA LOSS.**
**When LWW is fine:** **if you can avoid conflicts — e.g. only inserting records with a unique key and never updating them.** If you update existing records, or different leaders may insert records with the same key, **you have to decide whether lost updates are acceptable.**
**The clock hazard:** with a **real-time clock (Unix timestamp)**, the system becomes **very sensitive to clock synchronization. If one node's clock is AHEAD of the others, your attempt to overwrite a value written by that node MAY BE IGNORED because it has a lower timestamp — even though it clearly occurred later.** Solvable with a **logical clock** (Ch 10).
##### Strategy 3 — Manual conflict resolution [#strategy-3--manual-conflict-resolution]
**Like a Git merge conflict** — but **it would be impractical for a conflict to stop the entire replication process until a human resolves it.** Instead databases **store all concurrently written values (siblings)** and **return all of them on the next read**; you resolve them **automatically in application code (e.g. concatenate B and C into B/C) or by asking the user**, then write back a new value. **Used by CouchDB.**
**Four problems:**
1. **The API of the database changes** — a title that was a string becomes **a set of strings usually containing one element.** Awkward in application code.
2. **Asking the user to merge is a lot of work** — for the developer (building conflict-resolution UI) and the user (**who may be confused about what they're being asked and why**). **In many cases it's better to merge automatically than to bother the user.**
3. **Naive automatic merging is surprising.** **The Amazon shopping-cart anomaly:**
```txt
START: cart = {Book, DVD, Soap}
DEVICE 1 removes Book → {DVD, Soap} ┐
├── MERGE BY SET UNION
DEVICE 2 removes DVD → {Book, Soap} ┘ ▼
{Book, DVD, Soap}
▲ ▲
BOTH REMOVED ITEMS REAPPEAR IN THE CART
```
4. **Resolution can itself introduce a new conflict.** If multiple nodes observe and concurrently resolve the same conflict, **one node may merge B and C into `B/C` and another into `C/B`. When that conflict is merged, you may get `B/C/C/B` or something similarly surprising.**
##### Strategy 4 — Automatic conflict resolution [#strategy-4--automatic-conflict-resolution]
> **Automatic conflict resolution ensures all replicas CONVERGE to the same state — all replicas that have processed the same set of writes have the same state, REGARDLESS OF THE ORDER in which the writes arrived. Combining eventual consistency with a convergence guarantee is STRONG EVENTUAL CONSISTENCY.**
**LWW is the simplest such algorithm.** More sophisticated merge algorithms exist per data type, **with the goal of preserving the intended effect of all updates as much as possible, and hence avoiding data loss:**
| Data type | Merge approach |
| --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Text** (wiki title/body) | **Detect which characters were inserted or deleted** between versions; **preserve all insertions and deletions made in any sibling.** Concurrent insertions at the same position are **ordered deterministically so all nodes get the same outcome** |
| **Collection** (ordered to-do list, unordered cart) | Merge like text, **tracking insertions and deletions.** This is what fixes the Amazon anomaly: **the algorithms track that Book and DVD were DELETED, so the merged result is `{Soap}`** |
| **Counter** (likes on a post) | **Tell how many increments and decrements happened on each sibling and add them together, so the result does not double-count and does not drop updates** |
| **Key-value map** | **Merge updates to the same key by applying one of the other algorithms to the values; updates to different keys are handled independently** |
**The limit:** **if you want to enforce that a list contains no more than five items, and multiple users concurrently add so there are more than five, your only option is to DROP SOME OF THE ITEMS.**
> **Nevertheless, automatic conflict resolution is sufficient to build many useful apps. And if you start from the requirement of building a collaborative offline-first or local-first app, CONFLICT RESOLUTION IS INEVITABLE, and automating it is often the best approach.**
##### CRDTs vs OT [#crdts-vs-ot]
Both perform automatic merges for all the types above; **different design philosophies and performance characteristics.**
**The worked example:** two replicas start with `ice`. One prepends `n` → `nice`; the other concurrently appends `!` → `ice!`. Both must converge to `nice!`.
**Where each is used:** **OT is most often used for real-time collaborative text editing (Google Docs). CRDTs are found in distributed databases (Redis Enterprise, Riak, Azure Cosmos DB).** **Sync engines for JSON can be implemented with either — CRDTs (Automerge, Yjs) or OT (ShareDB).** *It's possible to combine the advantages of both in one algorithm.*
Lists and arrays work the same way with list elements instead of characters; other datatypes like key-value maps **can be added quite easily.**
#### 3.5 Types of conflict — the subtle kind [#35-types-of-conflict--the-subtle-kind]
**Obvious conflict:** two writes concurrently modified the same field of the same record.
**Subtle conflict — the meeting room booking system:**
> The system **inserts a NEW RECORD for each booking** rather than updating a field. The application must ensure **each room is booked by only one group at any one time — no overlapping bookings.**
>
> **A conflict arises if two bookings are created for the same room at the same time. EVEN IF THE APPLICATION CHECKS AVAILABILITY BEFORE ALLOWING A BOOKING**, a conflict can arise if the two bookings are made close enough together that **both see the room as unbooked prior to inserting their record.**
**There isn't a quick ready-made answer.** More examples in Ch 8 (write skew / phantoms); scalable approaches to detecting and resolving in Ch 13.
***
# 6.2 Problems with Replication Lag (/docs/ddia/replication/problems-replication-lag)
**Read-scaling architecture:** many followers, reads distributed across them. **This removes load from the leader and lets reads be served by nearby replicas.**
> **This realistically works only with ASYNCHRONOUS replication. If you synchronously replicated to all followers, a single node failure or network outage would make the entire system unavailable for writing. And the more nodes you have, the likelier one is down — so a fully synchronous configuration would be very unreliable.**
**Eventual consistency:** reading from an async follower may return outdated information. **Run the same query on the leader and a follower at the same time and you may get different results.** **Stop writing and wait, and the followers eventually catch up.**
> **The term "eventually" is deliberately vague; in general, THERE IS NO LIMIT to how far a replica can fall behind.** Normally a fraction of a second — but **near capacity or with network problems, easily seconds or minutes.**
*(Note: **it's not only NoSQL databases that are eventually consistent — followers in an asynchronously replicated RELATIONAL database have the same characteristics.**)*
#### The three anomalies [#the-three-anomalies]
#### 2.1 Implementing read-after-write consistency [#21-implementing-read-after-write-consistency]
| Technique | How | Limitation |
| ------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Read modifiable things from the leader** | E.g. **always read the user's OWN profile from the leader, other users' profiles from a follower** | **Requires knowing whether something MIGHT have been modified without querying it.** Fails if most things are user-editable — **most reads would hit the leader, negating read scaling** |
| **Time-based** | **Track the time of the last update; for one minute after it, make all reads from the leader.** Also: **monitor replication lag and prevent queries on any follower more than one minute behind** | Coarse; wastes leader capacity |
| **Client remembers a timestamp** | The client remembers **the timestamp of its most recent write**; the system ensures the serving replica reflects updates **at least until that timestamp.** If not sufficiently up to date, **either route to another replica or WAIT until it catches up** | Timestamp is either **logical** (a log sequence number) or **the actual system clock — in which case CLOCK SYNCHRONIZATION becomes critical** (Ch 9) |
| **Cross-region** | **Any request that must be served by the leader must be routed to the region containing the leader** | Added latency and complexity |
**Cross-device read-after-write** adds two more problems:
1. **Remembering the user's last-update timestamp** breaks, because **code on one device doesn't know what happened on the other. The metadata must be CENTRALIZED.**
2. **Different devices may route to different regions** — desktop on home broadband vs mobile on cellular data take **completely different network routes.** If your approach requires reading from the leader, **you may first need to route all of a user's devices to the same region.**
#### 2.2 Implementing monotonic reads [#22-implementing-monotonic-reads]
> **Make sure each user always reads from THE SAME REPLICA** (different users can use different replicas) — e.g. **choose the replica by a HASH OF THE USER ID rather than randomly.**
>
> **However, if that replica fails, the user's queries must be rerouted to another replica.**
#### 2.3 Implementing consistent prefix reads [#23-implementing-consistent-prefix-reads]
**This is particularly a problem in SHARDED databases.** If the database always applies writes in the same order, **reads always see a consistent prefix and the anomaly can't happen.** But **in many distributed databases, different shards operate independently, so there is no global ordering of writes** — a reader **may see some parts of the database in an older state and some in a newer state.**
**Solutions:**
* **Make sure any writes that are causally related to each other go to the same shard** — but **in some applications that can't be done efficiently**
* **Explicitly track causal dependencies** (→ §4.5, happens-before)
#### 2.4 The pragmatic advice [#24-the-pragmatic-advice]
> **Think about how the application behaves if replication lag increases to several minutes or even hours. If the answer is "no problem," that's great.** If it's a bad experience, **design for a stronger guarantee.**
>
> **Pretending that replication is synchronous when in fact it is asynchronous is a recipe for problems down the line.**
And on where to solve it: **dealing with these issues in application code is complex and easy to get wrong.**
> **The simplest programming model is to choose a database providing STRONG CONSISTENCY (linearizability, Ch 10) and ACID transactions (Ch 8), so you can mostly ignore replication challenges and treat the database as if it had a single node.**
**The historical arc:** in the early 2010s the NoSQL movement argued these features limited scalability and that large-scale systems must embrace eventual consistency. **Since then, a number of databases provide strong consistency and transactions WHILE ALSO offering fault tolerance, high availability, and scalability — the NewSQL trend.**
**But weaker consistency still has legitimate reasons:** **stronger resilience in the face of network interruptions, and lower overheads compared to transactional systems.**
***
# 6.7 Production failure catalog for this chapter (/docs/ddia/replication/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| -------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------- |
| Data deleted by mistake is gone from every replica | **Replication is not backup** |
| Writes acknowledged to the client vanish after failover | **Async replication + promotion of a lagging follower** |
| Primary keys reused; private data shown to wrong users | **Promoted follower's autoincrement counter behind; cross-system (Redis) inconsistency** — the GitHub incident |
| Two nodes both accepting writes | **Split brain**; missing or broken **fencing** |
| Cluster failed over during a load spike, making it worse | **Timeout too short** → unnecessary failover |
| Both nodes shut down during split-brain protection | Poorly designed mutual-shutdown mechanism |
| Cannot upgrade the database without downtime | **WAL shipping** couples leader and follower to one storage format |
| User submits a form and it "didn't save" | **No read-after-write consistency** |
| A comment appears, then disappears on refresh | **No monotonic reads** — reads hitting different-lag replicas |
| An answer appears before the question | **No consistent prefix reads** across shards |
| Primary's disk fills with WAL | **Inactive replication slot / retained log for a dead follower** |
| Standby queries cancelled with "conflict with recovery" | Replay vs long-running read on the standby |
| Removed cart items reappear | **Set-union merge of siblings** — the Amazon anomaly |
| Merged value is `B/C/C/B` | **Concurrent conflict resolution creating a new conflict** |
| A username was registered twice in two regions | **Multi-leader cannot enforce cross-leader uniqueness** |
| An UPDATE arrives before its INSERT | **All-to-all topology overtaking**; timestamps insufficient |
| Writes silently dropped in Cassandra | **LWW + clock skew** |
| Deleted rows come back | **Repair not run within `gc_grace_seconds`** |
| A write returned an error but the data is there | **Partial write below w is not rolled back** |
| Quorum reads still return stale data | Rebalancing, restored-from-old-replica, or concurrent read/write (§4.4) |
| Cannot tell how stale a leaderless cluster is | **No fixed write order ⇒ no lag metric** |
***
# 6.10 Self-test (/docs/ddia/replication/self-test)
Why doesn't replication remove the need for backups? Give the scenario where replication actively hurts.
Why is it impracticable for all followers to be synchronous? What is the semisynchronous compromise, and what exactly does it guarantee?
Describe the four steps of adding a new follower without downtime. What must the snapshot be associated with, and what is that called in Postgres and MySQL?
A follower has been offline for a week. State the leader's dilemma and both bad outcomes.
List the three steps of automatic failover, and the four things that can go wrong. Which one caused the GitHub incident, and how?
Why does a short failover timeout make an overloaded system *worse*?
Compare the three replication-log formats on: coupling to the storage engine, ability to run different versions on leader and follower, and external parseability.
Why does WAL shipping prevent zero-downtime upgrades? Why doesn't logical replication?
Define eventual consistency. Why is "eventually" deliberately vague?
For each of read-after-write, monotonic reads, and consistent prefix reads: state the anomaly in one sentence and give one implementation technique.
Why does cross-device read-after-write consistency break the "remember the client's last write timestamp" approach?
Why does the book treat *synchronous* multi-leader replication as equivalent to single-leader?
Give the precise reason multi-leader replication cannot enforce username uniqueness.
Compare circular, star, and all-to-all topologies on failure tolerance. What problem is unique to all-to-all, and why don't timestamps fix it?
Explain how an app with no offline mode can still be multi-leader.
State the "real meaning" of LWW. Under what single condition is LWW harmless?
Reproduce the Amazon shopping-cart anomaly and explain what a proper collection CRDT does differently.
Write the quorum condition. Why is it a probability adjustment rather than a guarantee? Give three of the six edge cases.
Why is staleness easy to monitor in leader-based replication and hard in leaderless?
Define happens-before. Why can two operations be concurrent even when the speed of light would have permitted causality?
Walk through the shopping-cart version-number trace and explain why no writes are lost despite the clients never being up to date.
Why is a single version number insufficient with multiple replicas? What replaces it, and what does it let the database distinguish?
you run a SaaS with users in the US, EU, and APAC. Requirements: (a) usernames globally unique, (b) users' own documents editable offline on mobile, (c) survive the loss of an entire region, (d) p99 write latency \< 100 ms in-region. Design the replication topology — you will need more than one model. State exactly which data uses which, and what you give up.
# 6.1 Single-Leader Replication (/docs/ddia/replication/single-leader-replication)
**Also called leader-based, primary-backup, or active/passive replication.**
**Three steps:**
1. One replica is the **leader**. Clients send writes **only** to the leader, which **first writes to its local storage.**
2. The other replicas are **followers**. Whenever the leader writes locally, it **also sends the change to all followers as part of a replication log / change stream.** Each follower **applies all writes in the same order as the leader.**
3. Reads may go to **the leader or any follower**; **writes are accepted only by the leader.**
> **If the database is sharded (Ch 7), EACH SHARD has one leader. Different shards may have leaders on different nodes — but each shard must have one leader node.**
**Where it's used:** PostgreSQL, MySQL, **Oracle Data Guard**, **SQL Server Always On availability groups**, MongoDB, DynamoDB, **Kafka**, replicated block devices (**DRBD**), some network filesystems. **Also many consensus algorithms — Raft (CockroachDB, TiDB, etcd, RabbitMQ quorum queues) — are based on a single leader and automatically elect a new one if the old one fails.**
*(Terminology: older documents say "master–slave." **The term should be avoided as it is widely considered offensive.**)*
#### 1.1 Synchronous vs asynchronous replication [#11-synchronous-vs-asynchronous-replication]
| | Advantage | Disadvantage |
| ---------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Synchronous** | **The follower is GUARANTEED to have an up-to-date copy consistent with the leader.** If the leader suddenly fails, **the data is still available on the follower** | **If the synchronous follower doesn't respond** (crash, network fault, any reason), **the write cannot be processed. The leader must BLOCK ALL WRITES and wait until the replica is available again** |
| **Asynchronous** | The leader **can continue processing writes even if all followers have fallen behind** | **If the leader fails unrecoverably, any writes not yet replicated are LOST.** A write is **not guaranteed durable even if it was confirmed to the client** |
> **It is impracticable for ALL followers to be synchronous — any one node outage would grind the whole system to a halt.**
**Semisynchronous** is the practical compromise: **one follower is synchronous, the others asynchronous. If the synchronous follower becomes unavailable or slow, one of the asynchronous followers is made synchronous.** This guarantees **an up-to-date copy on at least two nodes: the leader and one synchronous follower.**
**Quorum variant:** some systems update **a majority of replicas synchronously (e.g. 3 of 5 including the leader)** and the minority asynchronously. **Majority quorums are often used in eventually consistent systems or systems using a consensus protocol for automatic leader election** (Ch 10).
**Normal replication speed:** most systems apply changes to followers **in less than a second — but there is NO GUARANTEE.** Followers might fall behind **by several minutes or more** if: a follower is recovering from a failure, the system is near maximum capacity, or there are network problems.
> **Weakening durability may sound like a bad trade-off, but asynchronous replication is nevertheless widely used, especially with many followers or geographically distributed ones.**
#### 1.2 Setting up new followers [#12-setting-up-new-followers]
**Why you can't just copy the files:** clients are constantly writing, **so a standard file copy would see different parts of the database at different points in time — the result might not make any sense.** You *could* lock the database, **but that goes against high availability.**
**The four-step process (no downtime):**
**The practical steps vary significantly by database — some fully automated, others an arcane multistep workflow performed manually by an administrator.**
**Archiving to object storage:** archive the replication log to an object store along with periodic snapshots. **This is a good way of implementing backups and disaster recovery, and steps 1–2 become "download those files."** **WAL-G** does this for PostgreSQL/MySQL/SQL Server; **Litestream** for SQLite.
#### 1.3 Aside: databases backed by object storage [#13-aside-databases-backed-by-object-storage]
**Four benefits:**
1. **Inexpensive** — cloud databases can store rarely queried data on cheaper, higher-latency storage while serving the working set from memory/SSD/NVMe
2. **Multi-zone, dual-region, or multi-region replication with very high durability** — and it **lets databases bypass inter-zone network fees**
3. **Conditional write** — essentially a **compare-and-set (CAS)** operation — used to implement **transactions and leadership election**
4. **Data integration** — multiple databases in one object store, especially with open formats like **Parquet and Iceberg**
> **These benefits dramatically simplify the database architecture by shifting the responsibility of transactions, leadership election, and replication to object storage.**
**Five trade-offs:**
* **Much higher read and write latencies** than local disks or virtual block devices (EBS)
* **Per-API-call fees** force batching, **which further increases latency**
* **Objects are often immutable**, making **random writes in a large object extremely resource-intensive**
* **No standard filesystem interface.** **FUSE** lets you mount buckets as filesystems, **but many FUSE interfaces lack POSIX features like nonsequential writes or symlinks** that systems may depend on
**Three architectural responses:**
| Approach | Design |
| --------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Tiered storage** | Less frequently accessed data on object storage; new/hot data on SSD/NVMe or memory |
| **Object storage + separate WAL** | Object storage as primary tier, but a **separate low-latency system for the WAL** (Amazon EBS, Neon's Safekeepers) |
| **Zero-disk architecture (ZDA)** | **ALL data persisted to object storage; disks and memory used strictly for caching. Nodes have NO PERSISTENT STATE, which dramatically simplifies operations** |
**ZDA in the wild:** **WarpStream, Confluent Freight, Buf's Bufstream, Redpanda Serverless** (all Kafka-compatible), **nearly every modern cloud data warehouse**, **Turbopuffer** (vector search), **SlateDB** (cloud-native LSM).
#### 1.4 Handling node outages [#14-handling-node-outages]
**Follower failure → catch-up recovery.** Each follower keeps a **local log of changes received from the leader**. After a crash or a temporary network interruption, **it knows the last transaction processed before the fault**, connects to the leader, requests everything since, applies it, and catches up.
**The performance catch:** with high write throughput or a long offline period, **there may be a lot to catch up on — high load on BOTH the recovering follower AND the leader** (which must send the backlog).
**The log-retention dilemma:** the leader can delete its log once all followers confirm processing. **If a follower is unavailable for a long time, the leader must choose:**
**Leader failure → failover.** One follower is promoted, clients are reconfigured, other followers start consuming from the new leader.
**Automatic failover, three steps:**
1. **Determining that the leader has failed.** Crashes, power outages, network issues. **There is no foolproof way of detecting what has occurred, so most systems simply use a TIMEOUT** — nodes bounce messages back and forth; no response for, say, 30 seconds ⇒ assumed dead. *(Planned maintenance doesn't need this — the leader can trigger a safe handoff before shutting down.)*
2. **Choosing a new leader.** By **election** (chosen by a majority of remaining replicas) or **appointed by a previously established controller node.** **The best candidate is usually the replica with the most up-to-date data, to minimize data loss.** Getting all nodes to agree is **a consensus problem** (Ch 10).
3. **Reconfiguring the system.** Clients must send writes to the new leader. **If the old leader comes back, it might still believe it is the leader** — the system must ensure it becomes a follower and recognizes the new leader.
**Four things that go wrong — all worth memorizing:**
> **These problems have no easy solutions. For this reason, some operations teams prefer to perform failovers MANUALLY, even if the software supports automatic failover.**
**The one rule:** **pick an up-to-date follower.** With sync/semisync replication that's the follower the old leader waited for. With async, **pick the follower with the highest log sequence number.** **Losing a fraction of a second's worth of writes may be tolerable; picking a follower behind by several days could be catastrophic.**
#### 1.5 Implementation of replication logs — three methods [#15-implementation-of-replication-logs--three-methods]
##### (a) Statement-based replication [#a-statement-based-replication]
The leader **logs every write statement it executes** and sends the statement log to followers; **each follower parses and executes that SQL as if received from a client.**
**Three ways it breaks:**
1. **Nondeterministic functions** — `NOW()`, `RAND()` — **generate a different value on each replica**
2. **Autoincrementing columns, or statements depending on existing data** (`UPDATE … WHERE `) **must execute in exactly the same order on each replica**, or they have a different effect. **Limiting when there are multiple concurrently executing transactions**
3. **Side effects** — triggers, stored procedures, UDFs — **may differ on each replica unless absolutely deterministic**
**Workarounds:** the leader can **replace nondeterministic function calls with a fixed return value when logging.** *Executing deterministic statements in a fixed order is the same idea as **event sourcing** (Ch 3), and is known as **state machine replication*** (Ch 10).
**Status:** used by MySQL before 5.1; **still sometimes used because it's quite compact**, but **MySQL now switches to row-based replication by default if there is any nondeterminism.** **VoltDB uses it and makes it safe by requiring transactions to be deterministic** — but **determinism is hard to guarantee in practice.**
##### (b) Write-ahead log (WAL) shipping [#b-write-ahead-log-wal-shipping]
**The WAL already contains all information necessary to restore indexes and heap to a consistent state** (Ch 4 §3.3), **so we can use the exact same log to build a replica on another node.** The leader writes the log to disk **and** sends it across the network; the follower **builds a copy of the exact same files as on the leader.**
Used by **PostgreSQL and Oracle.**
> **The main disadvantage: the log describes the data at a VERY LOW LEVEL — which bytes were changed in which disk blocks. This makes replication TIGHTLY COUPLED TO THE STORAGE ENGINE. If the database changes its storage format between versions, it is typically not possible to run different versions on the leader and the followers.**
**And that has a big operational impact:**
##### (c) Logical (row-based) log replication [#c-logical-row-based-log-replication]
**Use a DIFFERENT log format for replication than for the storage engine**, decoupling replication from storage internals. Called a **logical log** to distinguish it from the storage engine's **physical** representation.
**Granularity: a row.**
| Operation | What's logged |
| ---------- | -------------------------------------------------------------------------------------------------------------------------------------------- |
| **Insert** | **The new values of all columns** |
| **Delete** | **Enough information to uniquely identify the deleted row** — typically the primary key; **if there's no PK, the old values of all columns** |
| **Update** | **Enough to identify the row, plus new values of all columns** (or at least all changed ones) |
A multi-row transaction generates several such records **followed by a record indicating the transaction was committed.**
**Implementations:** **MySQL keeps a separate logical log — the binlog — in addition to the WAL.** **PostgreSQL implements logical replication by DECODING THE PHYSICAL WAL into row insert/update/delete events.**
**Two advantages:**
1. **Easier to keep backward compatible → leader and follower can run different versions → minimal-downtime upgrades**
2. **Easier for external applications to parse** — useful to send DB contents to a data warehouse, or to a system building custom indexes and caches. **This technique is CHANGE DATA CAPTURE** (Ch 12).
**Summary table:**
| Method | Coupling | Cross-version replication | Parseable externally | Determinism required |
| ---------------- | ---------------------- | ------------------------- | -------------------- | ------------------------ |
| **Statement** | none | yes | yes | **YES — the fatal flaw** |
| **WAL shipping** | **tight (byte-level)** | **no** | no | no |
| **Logical/row** | loose | **yes** | **yes (→ CDC)** | no |
***
# 6.6 Technology deep dives (/docs/ddia/replication/technology-deep-dives)
***
#### 6.1 PostgreSQL streaming replication (single-leader, WAL shipping + logical) [#61-postgresql-streaming-replication-single-leader-wal-shipping--logical]
**Problem it solves.** Keep a byte-identical hot standby for failover and read scaling, with no application changes.
**Why not statement-based?** Nondeterminism (`now()`, `random()`, `nextval`) and ordering dependencies make it unsafe under concurrency.
**Why add logical replication later?** **WAL shipping is tightly coupled to the storage format**, so leader and follower must run the same major version — **which makes zero-downtime major upgrades impossible.** Logical decoding solves both that and CDC.
**How it works internally.** The leader streams WAL records over a replication connection; each standby has a **replication slot** recording the LSN it has consumed, so the leader retains WAL until it's consumed. `synchronous_commit` controls the durability level (`off` / `local` / `remote_write` / `on` / `remote_apply`); `synchronous_standby_names` picks which standbys are synchronous. **Logical replication** runs the WAL through an **output plugin** (`pgoutput`) that turns physical records into row-level insert/update/delete events.
**Deployment.** Primary + N standbys; a failover manager (Patroni + etcd/Consul, repmgr, or a cloud provider's). **Patroni's real job is fencing** — using a distributed lock so only one node believes it's primary.
**Monitoring.**
* **`pg_stat_replication`**: `sent_lsn` − `replay_lsn` per standby, in **bytes** and in **seconds** (`replay_lag`) — both matter; a standby can be byte-close but seconds-behind if replay is blocked
* **Replication slot retained WAL** — the #1 way to fill a primary's disk: an inactive slot pins WAL forever
* **`pg_stat_database_conflicts`** — queries on the standby cancelled by replay
* Sync-standby availability; failover events; timeline ID after promotion
**Scaling.** Read replicas for reads; cascading replication to reduce load on the primary; but **all writes still go through one node.**
**Backup.** Base backup + **WAL archiving** for PITR (pgBackRest, WAL-G). **Replication is not backup** — §0.
**What actually breaks.**
* **An abandoned replication slot** fills `pg_wal` and takes the primary down. This is the single most common Postgres replication outage.
* **`max_standby_streaming_delay` vs long-running read queries on the standby** — either queries get cancelled ("conflict with recovery") or replay stalls and lag grows. You must choose which.
* **Split brain after a network partition** if the failover manager lacks proper fencing — the old primary keeps accepting writes.
* **Lost writes on promotion** with `synchronous_commit = off` or an async standby promoted.
* **Timeline divergence** — the old primary rejoins on a diverged timeline and needs `pg_rewind` or a full rebuild.
* **Logical replication doesn't replicate DDL** — a schema change on the publisher breaks the subscriber.
***
#### 6.2 MySQL binlog replication + GTIDs [#62-mysql-binlog-replication--gtids]
**Problem it solves.** Same as above, but with a **logical log by default**, which makes cross-version replication and CDC first-class.
**Why row-based over statement-based?** MySQL shipped statement-based first and **switched to row-based by default whenever a statement is nondeterministic** — the practical concession that statement-based replication cannot be made safe in general.
**How it works internally.** The primary writes the **binlog** (row events by default, plus a commit marker per transaction). Replicas have an **I/O thread** copying the binlog to a local relay log, and one or more **SQL applier threads** replaying it. **GTIDs** give each transaction a globally unique `source_uuid:txn_id`, which makes failover safe: a replica knows exactly which transactions it has, rather than a file+offset that becomes meaningless after a primary change. Semi-sync (`rpl_semi_sync_source_wait_for_slave_count`) waits for N replicas to *receive* (not apply) before acking.
**Monitoring.** `Seconds_Behind_Source` (crude — it measures the *applier*, and reads 0 when the I/O thread is broken, so always pair it with `SHOW REPLICA STATUS` errors), applier thread errors, **GTID gaps** (`gtid_executed` vs the primary's), relay log disk usage, semi-sync ack timeouts (which silently downgrade to async).
**What actually breaks.**
* **The GitHub incident from §1.4** — a lagging replica promoted, its **autoincrement counter behind**, reusing primary keys that were already referenced by Redis, leaking private data to the wrong users. The lesson generalizes: **cross-system consistency breaks catastrophically when the promoted node's counters go backwards.**
* **Semi-sync silently degrading to async** after a timeout — you think you have durability and you don't.
* **Replication drift** with statement-based mode still enabled for some statements.
* **A single-threaded applier** falling behind a multi-threaded writer (fixed by parallel replication, which introduces its own ordering subtleties).
* **`sync_binlog=0` + `innodb_flush_log_at_trx_commit=2`** — fast, and loses committed transactions on a power failure.
***
#### 6.3 Cassandra / ScyllaDB (leaderless, Dynamo-style) [#63-cassandra--scylladb-leaderless-dynamo-style]
**Problem it solves.** Multi-region, always-writable storage that tolerates node and region failure without failover, with tunable staleness.
**Why not single-leader?** Failover time, leader sensitivity to gray failures, and cross-region write latency. **Leaderless doesn't distinguish the normal case from the failure case** (§4.6), so a slow node degrades nothing.
**Why not multi-leader?** Reads can be arbitrarily stale; quorums give a middle ground.
**How it works internally.** Consistent hashing ring; RF replicas per key; per-request **consistency level** (`ONE`, `QUORUM`, `LOCAL_QUORUM`, `EACH_QUORUM`, `ALL`, `ANY`) implementing the w/r knobs; **read repair**, **hinted handoff**, and **anti-entropy repair** (Merkle-tree based) as the three catch-up mechanisms; **LWW at the cell level using wall-clock timestamps**; storage is an LSM tree (Ch 4).
**Deployment.** Multi-DC keyspace with `NetworkTopologyStrategy`; clients pinned to a **local coordinator**; `LOCAL_QUORUM` as the default; scheduled repairs (weekly, within `gc_grace_seconds`).
**Monitoring.** Per-CL latency and timeout rate; **hint count and hint delivery rate** (§4.5 — the only real staleness proxy); **repair completion time vs `gc_grace_seconds`** (miss this and deleted data resurrects); pending compactions; **dropped mutations** (a write accepted but silently discarded under load); tombstone-scanned-per-read histograms; **clock skew across nodes** (§4.4 ⑤ makes this a data-correctness metric, not just an ops metric).
**Scaling.** Add nodes to the ring; **quorums seldom exceed 4-of-7 or 5-of-9** because bigger quorums raise the chance of hitting a slow replica.
**Backup.** Snapshots (hard links over immutable SSTables) + incremental backups shipped off-node.
**What actually breaks.**
* **Clock skew silently dropping writes.** LWW with wall-clock timestamps means a node with a fast clock can make later writes lose. NTP failure becomes data loss with no error anywhere.
* **Zombie data.** If a node misses a delete and repair doesn't run within `gc_grace_seconds`, the tombstone is purged elsewhere and the deleted row **comes back** on the next repair.
* **Failed writes that partly succeeded** aren't rolled back (§4.4 ④) — "the write returned an error" tells you nothing about whether the data is there.
* **Sloppy quorum / CL=ANY** accepting a write that no correct replica will ever serve.
* **Read repair storms** after a node returns from a long outage.
* **Using it for a workload needing uniqueness or invariants** — LWW cannot express "this username is taken."
***
#### 6.4 Sync engines / CRDT libraries (Automerge, Yjs, PouchDB/CouchDB, Firestore) [#64-sync-engines--crdt-libraries-automerge-yjs-pouchdbcouchdb-firestore]
**Problem it solves.** Sub-frame-latency local reads/writes, offline operation, and real-time multi-user collaboration on the same document.
**Why not request/response to a server?** Every interaction pays a round trip; every call needs error handling in the UI; offline requires a separate code path.
**Why CRDTs rather than LWW?** LWW on a document means one user's paragraph silently vanishes. Users notice.
**How it works internally.** Each character/element gets an **immutable unique ID**; operations reference the ID of the element they follow, not an index — so **no transformation is needed and replicas converge from any delivery order** (§3.4). Automerge/Yjs implement RGA/YATA-family sequence CRDTs plus maps, lists, counters, and text, all composable into a JSON-shaped document. Changes are exchanged as compact binary deltas; the full change history is retained to allow merging from any peer.
**Deployment.** A relay/persistence server (y-websocket, Automerge sync server, CouchDB) that is **just a dumb replica**, not an authority — which is what makes local-first possible.
**Monitoring.** Document size growth (**the metric** — CRDT metadata and tombstones accumulate), sync round-trip latency, peer count per document, merge conflict/sibling counts, client memory.
**Scaling.** Per-document, not per-dataset. **Sync engines assume all data the user may need is downloaded in advance** — fine for a user's own files, not for an ecommerce catalog.
**What actually breaks.**
* **Document bloat.** A long-lived collaborative document accumulates tombstones and per-character metadata until it's tens of MB and slow to load. Requires compaction/snapshotting and, sometimes, history truncation.
* **Invariants you cannot express.** "No more than 5 items" — concurrent adds will exceed it and **your only option is to drop some** (§3.4).
* **Merge results that are correct but semantically wrong.** Both users' edits preserved, producing a sentence neither wrote.
* **Access control.** If every replica has the whole document history, you cannot hide part of it.
* **Schema migration across offline clients** — a client that's been offline for six months syncs a document written by three schema versions later.
* **Deleted-data compliance** — the change history is, by design, permanent.
***
# 6.11 Terminology introduced here (/docs/ddia/replication/terminology-introduced-here)
# 6.9 Worked examples (/docs/ddia/replication/worked-examples)
**① Quorum arithmetic.** n=5. Which (w,r) satisfy `w+r>n`, and what does each tolerate?
| w | r | w+r | Valid? | Write availability | Read availability |
| - | - | --- | ------ | ------------------ | ------------------------------------------- |
| 3 | 3 | 6 | ✔ | tolerates 2 down | tolerates 2 down |
| 5 | 1 | 6 | ✔ | **0 down** | tolerates 4 down |
| 1 | 5 | 6 | ✔ | tolerates 4 down | **0 down** |
| 4 | 2 | 6 | ✔ | tolerates 1 down | tolerates 3 down |
| 2 | 2 | 4 | ✗ | tolerates 3 down | tolerates 3 down — **stale reads possible** |
**② Failover data loss.** Async replication, leader committing 5,000 writes/s, follower lag p99 = 800 ms. Leader dies. Expected loss on promotion ≈ **5,000 × 0.8 = 4,000 writes**. With semisync (one sync follower), loss on promotion of *that* follower = **0**; on promotion of an async follower, unchanged. **This calculation is the entire argument for semisync.**
**③ Read-after-write window.** Replication lag p99 = 300 ms; a user's page reload happens \~200 ms after submit. Without mitigation, **roughly the p99 fraction of users hitting a lagging replica see stale data** — and since users reload immediately, this is not a rare event, it's the common path. Reading from the leader for 1 second after a write eliminates it at the cost of routing \~1 leader-read per write.
**④ Tail latency vs quorum size.** If each replica has an independent 1% chance of being slow (>100 ms), then waiting for the fastest `r` of `n`:
* `r=2, n=3`: P(at least 2 of 3 fast) — slow response only if ≥2 replicas are slow ≈ 3×(0.01)² ≈ **0.03%**
* `r=5, n=9`: needs 5 fast of 9; slow if ≥5 slow — negligible, **but** you now wait for the 5th-fastest rather than the 2nd-fastest, so **median latency rises.**
This is why **quorums are seldom more than 4-of-7 or 5-of-9.**
**⑤ Version-vector siblings.** Three replicas, client reads at `{A:2, B:1, C:1}` and writes back. Meanwhile another client wrote on B, producing `{A:2, B:2, C:1}`. Comparing: neither vector dominates the other **only if** one has a strictly greater entry somewhere and strictly lesser elsewhere. Here `{A:2,B:2,C:1}` **dominates** `{A:2,B:1,C:1}` (≥ in all positions, > in one) ⇒ **not concurrent**; the B write happened after. Now add a third: `{A:3,B:1,C:1}` vs `{A:2,B:2,C:1}` — neither dominates ⇒ **concurrent ⇒ siblings.**
**⑥ Hinted-handoff load.** A node is offline for 4 hours in a 3-replica cluster taking 10,000 writes/s. Each write it missed is stored as a hint on another node: **\~144 million hints**, which must all be delivered when it returns — **on top of normal traffic, to a node that is already cold.** This is exactly the "additional load at a time when the system is already under strain" the book warns about, and it's why hint TTLs exist.
***
# 7.8 Decision cheat sheet (/docs/ddia/sharding/decision-cheat-sheet)
**Should I shard at all?**
Only if **data volume or WRITE throughput** exceeds one machine. **Read throughput → use replicas instead.** A single modern machine handles far more than folklore suggests, and **sharding is a heavyweight solution mostly relevant at large scale.**
**Key-range or hash?**
**Composite keys are usually the answer.** `(partition_key, clustering_key)` gives you hash-distributed shards **and** efficient ranges within a partition. Almost every well-designed sharded schema looks like this.
**How many shards per node?**
More than one, always — it makes rebalancing cheap and lets you weight powerful nodes. Cassandra defaults to 16 vnodes/node, ScyllaDB to 256.
**Automatic or manual rebalancing?**
Manual or semi-automatic if: you're near max write throughput, your failure detection is timeout-based, or a known traffic surge is coming. Fully automatic if the system is well below capacity and the vendor owns the operational risk.
**Which routing model?**
**Shard-aware client** for lowest latency (one hop) if you control the client. **Routing tier** if clients are heterogeneous or you want a wire-protocol-compatible façade. **Any-node forwarding** for simplicity at the cost of an extra hop. **In all three cases, put the authoritative map in a consensus-backed coordination service.**
**Local or global secondary index?**
**Local** when writes dominate, or when queries usually include the partition key. **Global** when **read throughput exceeds write throughput and postings lists are short** — and be explicit about whether you accept staleness or pay for a distributed transaction.
**Multitenant: shard per tenant?**
Yes if tenants are meaningfully sized and you want resource/permission isolation, per-tenant restore, per-tenant residency, and easy GDPR deletion. No if you have hundreds of thousands of tiny tenants (overhead), or one whale tenant that doesn't fit a node (you'll need to shard *within* it anyway).
***
# 7.12 Forward links (/docs/ddia/sharding/forward-links)
| Concept here | Where it's developed |
| ---------------------------------------------- | ------------------------------------------- |
| Writing to several shards atomically | **Ch 8** — Transactions (two-phase commit) |
| Race conditions corrupting hand-built indexes | **Ch 8** — Weak isolation |
| Why timeout-based failure detection misfires | **Ch 9** — Trouble with Distributed Systems |
| Consensus behind shard-assignment coordination | **Ch 10** — Consistency and Consensus |
| Parallel query execution across many shards | **Ch 11** — Batch Processing |
| Keeping derived indexes in sync via streams | **Ch 12** — Stream Processing |
| Replication of each shard | **Ch 6** — Replication |
# 7. Sharding (/docs/ddia/sharding)
> "Clearly, we must break away from the sequential and not limit the computers. We must state definitions and provide for priorities and descriptions of data. We must state relationships, not procedures." — Grace Murray Hopper
**A distributed database distributes data in two ways:**
> **Normally, shards are defined such that each piece of data (record, row, document) belongs to EXACTLY ONE shard. In effect, each shard is a small database of its own** — although some systems support operations touching multiple shards at once.
**Sharding is usually combined with replication**, so copies of each shard live on multiple nodes: **each record belongs to exactly one shard, but may still be stored on several nodes for fault tolerance.**
#### Terminology — the same thing under many names [#terminology--the-same-thing-under-many-names]
| System | Name |
| ------------------------------ | --------------- |
| Kafka | **partition** |
| CockroachDB | **range** |
| HBase, TiDB | **region** |
| Couchbase | **vBucket** |
| Riak | **vnode** |
| Cassandra | **token-range** |
| Bigtable, YugabyteDB, ScyllaDB | **tablet** |
> **PostgreSQL treats them as DISTINCT concepts: partitioning splits a large table into several files ON THE SAME MACHINE** (advantages: e.g. **very fast to delete an entire partition**), **whereas sharding splits a dataset ACROSS MULTIPLE MACHINES.** In many other systems, partitioning is just another word for sharding.
*Etymology, because it's a good story: one theory traces "shard" to the online RPG **Ultima Online**, in which **a magic crystal was shattered into pieces, and each shard refracted a copy of the game world** — so "shard" came to mean one of a set of parallel game servers, then carried over to databases. Another theory: an acronym for **System for Highly Available Replicated Data**, a 1980s database whose details are lost to history.*
> ⚠️ **Partitioning has NOTHING to do with NETWORK PARTITIONS (netsplits)** — a type of fault in the network between nodes (Ch 9).
***
# 7.7 Production failure catalog for this chapter (/docs/ddia/sharding/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ---------------------------------------------------- | -------------------------------------------------------------------------------------------- |
| 9 of 10 nodes idle, one at 100% | **Hot shard** — skewed partition key |
| All writes hit one shard, others idle | **Timestamp as the first element of a key-range partition key** |
| Adding a node moves nearly all the data | **`hash mod N`** |
| Can't add more nodes than shards | **Fixed shard count** chosen too low |
| Resharding requires downtime | System doesn't allow resharding with concurrent writes |
| Shard split makes an already-hot shard worse | **Splitting rewrites all the data**, and the shard needing a split is usually the loaded one |
| Cluster death spiral after one node slows | **Auto-rebalance + auto-failure-detection** cascading failure |
| Rebalancing can't keep up with writes | System near max write throughput while splitting |
| A single celebrity melts one node | **Hot key** — uniform hashing ≠ uniform load |
| Salting the hot key didn't help reads | **Salting splits writes only; reads must still gather all 100 keys** |
| One query silently costs 100× | **Scatter query** — partition key omitted |
| More shards, same query throughput | **Local secondary index** — every shard processes every query |
| Secondary index disagrees with the data | Hand-rolled index + race conditions / partial write failures |
| Global index returns stale results | **Asynchronous** global index maintenance (DynamoDB) |
| Two coordinators assign the same shard differently | **Split brain** in the shard-assignment coordinator |
| Requests lost during a shard move | **Cutover window** with in-flight requests to the old node |
| One partition grows to 8 GB and everything times out | **Unbounded partition** — missing bucketing in the clustering key |
| Cross-shard write half-succeeded | No distributed transaction (Ch 8) |
***
# 7.1 Pros and Cons of Sharding (/docs/ddia/sharding/pros-cons-sharding)
**The primary reason is SCALABILITY** — data volume or **write throughput** has become too great for a single node.
> **If READ throughput is the problem, you don't necessarily need sharding — you can use read scaling** (Ch 6 followers).
**Sharding is one of the main tools for HORIZONTAL SCALING (scale-out)** — growing capacity **not by moving to a bigger machine, but by adding more smaller machines.** If you can divide the workload so each shard handles a roughly equal share, **you can assign shards to different machines to process data and queries in parallel.**
> **While replication is useful at BOTH small and large scale — it enables fault tolerance and offline operation — SHARDING IS A HEAVYWEIGHT SOLUTION THAT IS MOSTLY RELEVANT AT LARGE SCALE. If a single machine can handle your data volume and write throughput (AND A SINGLE MACHINE CAN DO A LOT NOWADAYS!), it's often better to avoid sharding and stick with a single-shard database.**
#### The four costs [#the-four-costs]
| Cost | Detail |
| ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **The partition key choice is consequential and hard to change** | **All records with the same partition key go in the same shard.** Accessing a record is **fast if you know which shard it's in; if you don't, you have to do an INEFFICIENT SEARCH ACROSS ALL SHARDS.** And **the sharding scheme is difficult to change** |
| **Relational data is harder than key-value** | Sharding **works well for key-value data, where you can easily shard by key**, but harder with relational data where you may want to **search by a secondary index or JOIN records distributed across shards** (§6) |
| **Cross-shard writes need distributed transactions** | **A write may need to update related records in several shards.** Single-node transactions are common; **cross-shard consistency requires a DISTRIBUTED TRANSACTION — usually much slower, and may become a bottleneck for the system as a whole** (Ch 8) |
| **Operational complexity** | Rebalancing, routing, monitoring skew — all new work |
**A non-scalability use: sharding on ONE machine.** Some systems run **one single-threaded process per CPU core** to exploit CPU parallelism or a **NUMA (nonuniform memory access)** architecture where some memory banks are closer to one CPU than others. **Redis, VoltDB, and FoundationDB use one process per core and rely on sharding to spread load across cores in the same machine.**
***
# 7.4 Request Routing (/docs/ddia/sharding/request-routing)
**The question: to read or write a particular key, which node — which IP address and port — do you connect to?**
> **Very similar to SERVICE DISCOVERY (Ch 5). The biggest difference: with application services, each instance is usually STATELESS and a load balancer can send a request to any instance. With sharded databases, A REQUEST FOR A KEY CAN BE HANDLED ONLY BY A NODE THAT IS A REPLICA FOR THE SHARD CONTAINING THAT KEY.**
**Three approaches:**
**Three key problems, whichever you pick:**
1. **Who decides which shard lives on which node?** **Simplest is a single coordinator — but then how do you make it fault-tolerant if that node goes down? And if the coordinator role can fail over, how do you prevent SPLIT BRAIN where two coordinators make CONTRADICTORY shard assignments?**
2. **How does the routing component learn about changes** in the shard→node assignment?
3. **The cutover window:** **while a shard is being moved, the new node has taken over but REQUESTS TO THE OLD NODE MAY STILL BE IN FLIGHT. How do you handle those?**
**The standard answer — a coordination service:**
| System | Coordination mechanism |
| ------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **HBase, SolrCloud** | **ZooKeeper** |
| **Kubernetes** | **etcd** (tracks which service instance runs where) |
| **MongoDB** | **Own config server implementation + `mongos` daemons as the routing tier** |
| **Kafka, YugabyteDB, TiDB, ScyllaDB** | **Built-in implementations of the RAFT consensus protocol** |
| **Riak** | **A GOSSIP PROTOCOL among nodes** — **much weaker consistency than consensus; SPLIT BRAIN IS POSSIBLE, with different parts of the cluster having different node assignments for the same shard.** *Leaderless databases can tolerate this because they make weak consistency guarantees anyway* |
**And for finding IPs in the first place:** **shard→node assignment is fast-changing, but node IP addresses are not — so DNS is often sufficient for that layer.**
> **This discussion focused on finding the shard for an INDIVIDUAL KEY — most relevant for sharded OLTP. Analytical databases shard too, but with a very different query execution: rather than executing in a single shard, a query commonly needs to AGGREGATE AND JOIN DATA FROM MANY SHARDS IN PARALLEL** (Ch 11).
***
# 7.10 Self-test (/docs/ddia/sharding/self-test)
Distinguish sharding from replication in one sentence each. How do they combine under single-leader replication?
Why is sharding recommended *only* at large scale, when replication is useful at any scale?
Name four costs sharding imposes. Which one is hardest to reverse?
Give three reasons a SaaS product might shard per tenant that have nothing to do with scalability.
Define skew, hot shard, and hot key. Give an example where hashing eliminates the first but not the third.
Why must key-range shard boundaries "adapt to the data"? Give the encyclopedia illustration.
A time-series database keyed by timestamp has one shard at 100% and the rest idle. Diagnose it, fix it, and state precisely what the fix costs you.
Why is splitting a shard both necessary and dangerous?
Work out how many keys move when going from 4 nodes to 5 under `mod N` versus under a fixed-shard scheme.
In the fixed-shard scheme, three things could change during rebalancing — which one actually does?
Why choose a shard count that is "divisible by many factors"? Why can you not have more nodes than shards?
State the Goldilocks problem of fixed shard counts.
How does hash-range sharding get the benefits of both key-range and hash sharding? What does it still lose, and what partially rescues it?
What two properties define a consistent hashing algorithm? What does "consistent" *not* mean here?
Explain why salting a hot key helps writes but not reads, and what bookkeeping it forces on you.
Draw the cascading-failure loop caused by automatic rebalancing plus timeout-based failure detection.
Name the three request-routing architectures and the three problems common to all of them.
Why is ZooKeeper/etcd used for shard assignment rather than a single coordinator? What does Riak do instead, and what does it give up?
For local and global secondary indexes: which is cheap on write, which is cheap on single-condition read, and why does neither give you cheap multi-condition reads?
Why does adding shards not increase query throughput for a local secondary index?
you're building an event-tracking system: 500k events/s, 30-day retention, queried as (a) "all events for user X in the last hour" (high volume, low latency) and (b) "count of event type Y per hour across all users" (analyst, minutes acceptable). Choose the sharding scheme, the partition and clustering keys, the secondary-index strategy, and the routing architecture. Justify each against a specific trade-off from this chapter, and name the failure mode you're most worried about.
# 7.3 Sharding of Key-Value Data (/docs/ddia/sharding/sharding-key-value-data)
**The goal: spread data AND query load EVENLY.** With a fair share each, **10 nodes should handle 10× the data and 10× the throughput** (ignoring replication). Adding/removing a node should allow **rebalancing.**
**Vocabulary of unfairness:**
| Term | Meaning |
| ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Skewed** | Some shards have more data or queries than others — **makes sharding much less effective** |
| **Hot shard / hot spot** | A shard with disproportionately high load. **In the extreme, all load on one shard → 9 of 10 nodes idle and your bottleneck is the single busy node** |
| **Hot key** | One key with particularly high load — **e.g. a celebrity in a social network** |
**We need an algorithm taking a record's PARTITION KEY and returning its shard.** In a key-value store the partition key is **usually the key, or the first part of the key**; in a relational model it might be **a column of a table — not necessarily its primary key.** **The algorithm must be amenable to rebalancing in order to relieve hot spots.**
#### 3.1 Sharding by key range [#31-sharding-by-key-range]
**Assign a contiguous range of partition keys (min to max) to each shard — like the volumes of a paper encyclopedia.**
**Who chooses the boundaries:**
| Mode | Systems |
| ------------- | --------------------------------------------------------------------------------- |
| **Manual** | **Vitess** (a sharding layer for MySQL) |
| **Automatic** | **Bigtable, HBase, MongoDB (range option), CockroachDB, RethinkDB, FoundationDB** |
| **Both** | **YugabyteDB** (manual and automatic tablet splitting) |
**Within each shard, keys are stored in SORTED order** (B-tree or SSTables, Ch 4). Two advantages:
* **Range scans are easy**
* **You can treat the key as a CONCATENATED INDEX to fetch several related records in one query** — e.g. sensor data keyed by timestamp, where a range scan gets **all readings from a particular month**
**The downside — hot shards from nearby writes:**
**Rebalancing key-range shards:**
* **Pre-splitting** — HBase and MongoDB let you **configure an initial set of shards on an empty database.** **This requires that you already have some idea what the key distribution will look like**
* **Splitting** — as volume/throughput grow, **split an existing shard into two or more, each holding a contiguous subrange.** Similar to **what happens at the top level of a B-tree**
* **Merging** — if large amounts of data are deleted, **merge adjacent small shards**
* **Split triggers** — a **configured size** (HBase default: **10 GB**) or, in some systems, **write throughput persistently above a threshold.** **Thus a HOT shard may be split even if it isn't storing much data**, so its write load spreads
> ✔ **The number of shards adapts to the data volume** — small data → few shards → small overheads; huge data → each shard capped at a configurable maximum.
>
> ✗ **Splitting a shard is EXPENSIVE — it requires all its data to be rewritten into new files, similarly to a compaction. And a shard that needs splitting is often ALSO one under high load, so the cost of splitting can EXACERBATE that load, risking it becoming overloaded.**
#### 3.2 Sharding by hash of key [#32-sharding-by-hash-of-key]
**If you don't care whether partition keys are near each other** (e.g. tenant IDs), **hash the partition key first.**
> **A good hash function takes SKEWED data and makes it UNIFORMLY DISTRIBUTED.** A 32-bit hash returns a seemingly random number from 0 to 2³²−1; **even very similar input strings hash to evenly distributed values — but the same input always produces the same output.**
**For sharding, the hash function need NOT be cryptographically strong:** **MongoDB uses MD5; Cassandra and ScyllaDB use Murmur3.**
> ⚠️ **Language built-in hashes may be unsuitable: in Java's `Object.hashCode()` and Ruby's `Object#hash`, THE SAME KEY MAY HAVE A DIFFERENT HASH VALUE IN DIFFERENT PROCESSES.**
##### (a) Hash mod N — the naive approach, and why it fails [#a-hash-mod-n--the-naive-approach-and-why-it-fails]
> **`mod N` is easy to compute but leads to VERY INEFFICIENT REBALANCING because of a lot of unnecessary movement. We need an approach that moves AS LITTLE DATA AS POSSIBLE.**
##### (b) Fixed number of shards [#b-fixed-number-of-shards]
**Create many more shards than nodes and assign several shards to each node.**
**Note the cutover detail:** **reassignment is not immediate — it takes time to transfer a large amount of data over the network — so THE OLD ASSIGNMENT IS USED FOR ANY READS AND WRITES THAT HAPPEN WHILE THE TRANSFER IS IN PROGRESS.**
**Two practical tips:**
* **Choose a shard count divisible by many factors**, so the dataset can be evenly split across various numbers of nodes — **not requiring the node count to be a power of 2**
* **Account for mismatched hardware: assign MORE shards to more powerful nodes** so they take a greater share of load
**Used by Citus (sharding layer for PostgreSQL), Riak, Elasticsearch, Couchbase.**
**The limitations:**
> **It works well AS LONG AS YOU HAVE A GOOD ESTIMATE of how many shards you'll need when you first create the database.** You can then add/remove nodes easily — **subject to the limitation that YOU CAN'T HAVE MORE NODES THAN SHARDS.**
>
> **If the number turns out wrong, an EXPENSIVE RESHARDING is required: split each shard and write it out to new files, using a lot of additional disk space. Some systems DON'T ALLOW RESHARDING WHILE CONCURRENTLY WRITING, which makes it difficult to change the shard count WITHOUT DOWNTIME.**
**The Goldilocks problem:** since each shard holds a **fixed fraction** of total data, **shard size grows proportionally to total data.**
##### (c) Sharding by hash range — the best of both [#c-sharding-by-hash-range--the-best-of-both]
**Combine key-range sharding with a hash function so each shard contains a range of HASH VALUES rather than a range of KEYS.**
**The saving grace for range queries:**
> **If keys consist of two or more columns and the partition key is only the FIRST of them, you can still perform efficient range queries over the SECOND and later columns. As long as all records in the range query have the SAME partition key, they will be in the SAME shard.**
*(So `(user_id, timestamp)` hashed on `user_id` still gives you efficient "all events for user X between T1 and T2".)*
**Used by:** **YugabyteDB, DynamoDB**; an option in **MongoDB**. **Cassandra and ScyllaDB use a variant:**
*Aside — the warehouse equivalent:* **BigQuery**: the partition key determines the partition, **"cluster columns" determine sort order within it**. **Snowflake** assigns **micro-partitions** automatically but lets you define **cluster keys**. **Delta Lake** supports both manual and automatic partition assignment plus cluster keys. **Clustering improves range-scan performance AND compression AND filtering.**
##### (d) Consistent hashing [#d-consistent-hashing]
> **A consistent hashing algorithm maps keys to a specified number of shards satisfying two properties:**
>
> 1. **The number of keys mapped to each shard is roughly equal**
> 2. **When the number of shards changes, AS FEW KEYS AS POSSIBLE are moved**
> ⚠️ **"Consistent" here has NOTHING to do with replica consistency (Ch 6) or ACID consistency (Ch 8)** — it describes **the tendency of a key to stay in the same shard if possible.**
**Cassandra/ScyllaDB's algorithm is similar to the original definition.** Other algorithms: **highest random weight (rendezvous hashing)** and **jump consistent hashing.**
> **With these approaches, rather than a small number of existing shards being SPLIT INTO SUBRANGES to create new shards for a new node, THE NEW NODE IS INSTEAD ASSIGNED INDIVIDUAL KEYS that were previously scattered across all the other nodes. Which is preferable depends on the application.**
#### 3.3 Skewed workloads and relieving hot spots [#33-skewed-workloads-and-relieving-hot-spots]
> **Consistent hashing ensures keys are uniformly distributed across nodes — but that DOESN'T MEAN THE ACTUAL LOAD IS UNIFORMLY DISTRIBUTED.**
**Skew means:** much more data under some partition keys than others, **or** the request rate to some keys is much higher. **You can still end up with some servers overloaded while others sit almost idle.**
**The canonical case:** **a post by a celebrity with millions of followers causes a storm of activity — a large volume of reads and writes to THE SAME KEY** (the celebrity's user ID, or the ID of the action people are commenting on).
**Three mitigations:**
**① Dedicated shard.** **A system that defines shards by ranges of keys (or hashes) can put an individual hot key in a shard BY ITSELF — perhaps even assigning it a dedicated machine.**
**② Application-level key splitting (salting).**
**The bookkeeping cost:** it makes sense to salt **only the small number of hot keys** — for the vast majority of low-throughput keys **this would be unnecessary overhead.** **So you also need a way to TRACK WHICH KEYS ARE SPLIT, and a PROCESS FOR CONVERTING a regular key into a specially managed hot key.**
**And it's dynamic:** **a viral post may be hot for a couple of days and then calm down. Some keys may be hot for WRITES while others are hot for READS, necessitating DIFFERENT STRATEGIES.**
**③ Automated heat management.** Some cloud services do this automatically — **Amazon calls it heat management or adaptive capacity.**
#### 3.4 Automatic vs manual rebalancing [#34-automatic-vs-manual-rebalancing]
| Mode | Examples |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| **Fully automatic** | DynamoDB — promoted as able to **automatically add and remove shards to adapt to big load changes within a matter of minutes** |
| **Fully manual** | Explicitly configured by an administrator |
| **Middle ground** | **Couchbase and Riak generate a suggested shard assignment automatically but REQUIRE AN ADMINISTRATOR TO COMMIT IT** |
**Why automatic is attractive:** less operational work; **systems can even autoscale** to adapt to workload changes.
**Why automatic is dangerous:**
> **For that reason, it can be good to have A HUMAN IN THE LOOP for rebalancing. It's slower than a fully automatic process, but it can help prevent operational surprises.**
>
> **Manual rebalancing is also useful for PREEMPTIVELY rebalancing when a surge is expected from a known event — Cyber Monday sales, or ticket sales for the World Cup.**
*(This is the same argument as Ch 6's manual-failover preference and Ch 2's "autoscaling is cool, but predictable load may prefer manual" — a consistent theme: automation that reacts to failure signals can amplify failures.)*
***
# 7.2 Sharding for Multitenancy (/docs/ddia/sharding/sharding-multitenancy)
**SaaS products are often multitenant — each tenant is a customer.** Multiple users may log in to the same tenant, but **each tenant has a self-contained dataset separate from other tenants.** (An email marketing service: each business is a tenant; its newsletter sign-ups and delivery data are separate from other businesses'.)
**Either each tenant gets a separate shard, or multiple small tenants are grouped into a larger shard.** Shards might be **physically separate databases** (cf. the embedded-storage-engine-per-tenant idea in Ch 4) **or separately manageable portions of a larger logical database.**
**Seven advantages:**
| Advantage | Why it matters |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Resource isolation** | **One tenant's expensive operation is less likely to affect other tenants** on different shards |
| **Permission isolation** | **If there's a bug in your access control logic, you're less likely to accidentally give one tenant access to another's data** when the datasets are physically separate |
| **Cell-based architecture** | Apply sharding **not only at storage but to the SERVICES running your application code.** Services + storage for a set of tenants form a self-contained **cell**; cells run largely independently. **A fault in one cell remains limited to that cell** — fault isolation |
| **Per-tenant backup and restore** | **Restore one tenant's state without affecting others** — useful when a tenant accidentally deletes or overwrites important data |
| **Regulatory compliance** | **GDPR/CCPA rights to access and deletion become simple export and deletion operations on that person's shard** |
| **Data residence** | A **region-aware database can assign a tenant's shard to a particular region** to satisfy data residency laws |
| **Gradual schema rollout** | **Roll out migrations one tenant at a time** — reduces risk by detecting problems before they affect all tenants, **but can be difficult to do transactionally** |
**Three challenges:**
1. **It assumes each individual tenant fits on a single node.** If you have one tenant too big for a machine, **you need sharding WITHIN that tenant — back to sharding for scalability**
2. **Many small tenants** → a shard each may incur **too much overhead**. Group them — **but then you have the problem of MOVING TENANTS between shards as they grow**
3. **Cross-tenant features become harder** if they need joins across shards
***
# 7.5 Sharding and Secondary Indexes (/docs/ddia/sharding/sharding-secondary-indexes)
**Everything so far relies on the client knowing the PARTITION KEY.** That works in key-value models where the partition key is the first part of (or all of) the primary key.
**A secondary index usually doesn't identify a record uniquely — it's a way of searching for occurrences of a value:** *find all actions by user 123, find all articles containing "hogwash", find all cars whose color is red.*
> **The problem with secondary indexes is that THEY DON'T MAP NEATLY TO SHARDS.** Two approaches: **local** and **global**.
#### 5.1 Local secondary indexes (document-partitioned) [#51-local-secondary-indexes-document-partitioned]
**Each shard independently maintains its own secondary indexes, covering only the records in that shard.**
**Writes are easy:** you deal **only with the shard containing the record you're writing.** *(When a red car is added, that shard automatically adds its ID to the postings list for `color:red`.)*
**Reads:**
| Situation | Cost |
| --------------------------------------------------------- | ----------------------------------------------------------------------------- |
| **You know the partition key** | Search **the appropriate shard only** ✔ |
| **You want only SOME results, not all** | **Send to any shard** ✔ |
| **You want ALL results and don't know the partition key** | **Send the query to ALL shards and combine results** — a **scatter/gather** ✗ |
> **This makes read queries on secondary indexes QUITE EXPENSIVE. Even querying shards in parallel, it is prone to TAIL LATENCY AMPLIFICATION** (Ch 2 §2.6). **It also LIMITS SCALABILITY: adding more shards lets you store more data, but it DOESN'T INCREASE QUERY THROUGHPUT if every shard has to process every query anyway.**
**Nevertheless widely used:** MongoDB, Riak, Cassandra, Elasticsearch, SolrCloud, VoltDB.
> ⚠️ **Warning on rolling your own:** if your database supports only key-value, you may be tempted to implement a secondary index in application code. **Take great care to ensure the indexes remain consistent with the underlying data. RACE CONDITIONS AND INTERMITTENT WRITE FAILURES (where some changes were saved but others weren't) CAN VERY EASILY CAUSE THE DATA TO GO OUT OF SYNC** (Ch 8).
#### 5.2 Global secondary indexes (term-partitioned) [#52-global-secondary-indexes-term-partitioned]
**Construct a global index covering data in all shards. But you can't store it on one node — it would become a bottleneck and defeat the purpose. So THE GLOBAL INDEX MUST ALSO BE SHARDED — but it can be sharded DIFFERENTLY from the primary-key index.**
*(A "term" generalizes the full-text-search notion of a keyword to mean **any value you can search for in the secondary index.**)*
**The index shard can hold a contiguous range of terms, or terms can be assigned by HASH of the term.**
**Reads:**
| Query | Cost |
| --------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Single condition** (`color = red`) | **Read from ONE index shard to fetch the postings list** ✔ |
| **Fetch the actual records, not just IDs** | **Still have to read from all the data shards responsible for those IDs** ✗ |
| **Multiple conditions** (color AND make; or multiple words in a text) | **Those terms will likely be on DIFFERENT index shards.** To compute the logical AND, **find all IDs in BOTH postings lists.** Fine if the lists are short — **but if they're long, it can be slow to send them over the network to compute the intersection** ✗ |
**Writes are the real problem:**
> **Writing a single record might affect MULTIPLE SHARDS OF THE INDEX (every term in the document might be on a different shard). This makes it harder to keep the secondary index in sync with the underlying data. One option is a DISTRIBUTED TRANSACTION to atomically update the shards storing the primary record and its secondary indexes** (Ch 8).
**Used by CockroachDB, TiDB, YugabyteDB. DynamoDB supports BOTH local and global.**
> **In DynamoDB, writes are ASYNCHRONOUSLY reflected in global indexes, so READS FROM A GLOBAL INDEX MAY BE STALE** — similar to replication lag (Ch 6).
>
> **Nevertheless, global indexes are useful IF READ THROUGHPUT IS HIGHER THAN WRITE THROUGHPUT, and if the postings lists are not too long.**
#### 5.3 The comparison [#53-the-comparison]
| | Local (document-partitioned) | Global (term-partitioned) |
| ------------------------------ | ---------------------------------------------------------- | ------------------------------------------------------------------------------------------- |
| **Index sharded by** | the **record's** partition key | the **indexed value (term)** |
| **Write** | **ONE shard** ✔ | **Several index shards** ✗ (may need a distributed transaction, or async ⇒ **stale reads**) |
| **Read (single condition)** | **ALL shards — scatter/gather** ✗ | **ONE shard for the postings list** ✔ |
| **Read (fetch records)** | already local | **still multiple data shards** |
| **Read (multi-condition AND)** | each shard ANDs locally | **cross-shard postings-list intersection over the network** |
| **Scalability** | **more shards ≠ more query throughput** | scales with terms |
| **Consistency** | naturally consistent | **hard — needs distributed txn or accepts staleness** |
| **Used by** | MongoDB, Riak, Cassandra, Elasticsearch, SolrCloud, VoltDB | CockroachDB, TiDB, YugabyteDB; DynamoDB (both) |
***
# 7.6 Technology deep dives (/docs/ddia/sharding/technology-deep-dives)
***
#### 6.1 Vitess (manual key-range sharding for MySQL) [#61-vitess-manual-key-range-sharding-for-mysql]
**Problem it solves.** Let an existing, huge MySQL deployment scale horizontally without rewriting the application, and without giving up MySQL's transactional semantics inside a shard.
**Why not just a bigger MySQL?** Single-node write throughput and dataset size ceilings; and at YouTube/Slack scale the ceiling is real.
**Why not automatic sharding?** Vitess deliberately gives operators **manual control over key-range boundaries** — because resharding is expensive and the timing matters (§3.4's "human in the loop").
**How it works internally.** `vtgate` is the **routing tier** (§4 approach ②) speaking the MySQL wire protocol, so the application thinks it's talking to one MySQL. It parses the query, consults the **VSchema** (which declares the sharding key per table), and routes or scatters. `vttablet` sits in front of each MySQL instance. **Resharding** is done online: new shards are created, data is copied, changes are streamed via **VReplication** (built on the binlog — Ch 6 §1.5c), and traffic is cut over per-shard-per-direction with the ability to reverse.
**Deployment.** Topology service (**etcd/ZooKeeper/Consul**) holding the authoritative shard map; `vtgate` fleet behind a normal L4 LB; `vttablet` + MySQL pairs.
**Monitoring.** **Scatter-query rate** (queries hitting all shards — this is the metric that tells you your VSchema is wrong); per-shard QPS and row-count skew; VReplication lag during a reshard; `vtgate` query-plan cache hit rate; cross-shard transaction rate.
**Scaling.** Split a shard when it's too big or too hot; the key range makes the split a clean range operation.
**Backup.** Per-shard MySQL backups + binlog; **plus the topology service, which must be backed up separately** — losing the shard map is losing the database.
**What actually breaks.**
* **Scatter queries by accident.** A query that omits the sharding key hits every shard. One such query on a hot path silently converts your 100-shard cluster into 100× the load. This is §5.1's problem, made operational.
* **Choosing a sharding key that doesn't match the access pattern** — irreversible without a full reshard.
* **Reshard cutover with in-flight writes** (§4 problem 3).
* **Cross-shard transactions** — Vitess offers 2PC but it is slow and discouraged (Ch 8).
* **`AUTO_INCREMENT` across shards** — needs a sequence table or a different ID scheme.
***
#### 6.2 Cassandra / ScyllaDB token-range sharding [#62-cassandra--scylladb-token-range-sharding]
**Problem it solves.** Sharding that rebalances with minimal data movement, no coordinator, and no fixed shard count.
**Why not `mod N`?** §3.2a — most keys move on every topology change.
**Why not a fixed shard count?** §3.2b — you must guess the count correctly at creation, can't exceed it with nodes, and resharding is a downtime event.
**How it works internally.** Murmur3 hashes the partition key into a token; the token space is divided into **ranges with random boundaries**, with **16 vnodes/node in Cassandra, 256 in ScyllaDB** by default. Adding a node **takes slices from several existing nodes** rather than splitting one, which spreads the streaming load. Within a partition, rows are **clustered (sorted) by the clustering columns** — which is exactly the "partition key first, range query on later columns" trick of §3.2c. ScyllaDB additionally shards **per CPU core** within a node (the §1 NUMA/thread-per-core idea).
**Deployment.** Multi-DC with `NetworkTopologyStrategy`; token allocation algorithm configured to reduce imbalance; clients use a **token-aware driver** — §4 approach ③, so the client computes the token itself and connects directly to a replica.
**Monitoring.** **Per-node ownership percentage** (should be near-uniform; if not, your vnode count or token allocation is wrong); **partition size distribution — the p99 and max are what matter**, because one giant partition is the classic Cassandra outage; **tombstones per read**; hot-partition detection via per-table read/write latency outliers; streaming progress during topology change.
**Scaling.** Add nodes; ranges adjust automatically. **You cannot easily change the partition key** — it determines everything.
**What actually breaks.**
* **The unbounded partition.** Modeling `(user_id)` as the whole partition key for a table that grows forever produces a multi-GB partition on one node. Reads time out, compactions take hours, and the node is a permanent hot spot. **The fix is bucketing** — `(user_id, month)` — which is the §3.1 timestamp-prefix lesson in another dress.
* **Hot key from a celebrity** (§3.3) — token-uniformity does not imply load-uniformity.
* **Local secondary indexes used as if they were global.** Cassandra's `CREATE INDEX` is a local index; a query on it scatters to every node. It is the single most misused feature in Cassandra.
* **Adding many nodes at once**, saturating the network with streaming.
* **Gossip-based topology disagreement** during a partition — weaker than consensus, by design.
***
#### 6.3 DynamoDB (hash-range sharding + adaptive capacity) [#63-dynamodb-hash-range-sharding--adaptive-capacity]
**Problem it solves.** Fully managed sharding where the operator never thinks about partitions, with automatic splitting on both size and throughput.
**Why not manual?** The whole product promise is that shard management is invisible.
**How it works internally.** Partition key hashed to place the item; **sort key** orders items within the partition (again §3.2c). Partitions split when they exceed **10 GB** or their throughput ceiling. **Adaptive capacity / heat management** (§3.3) reallocates throughput toward hot partitions, and **isolates a single hot key onto its own partition** when needed. **Global secondary indexes are asynchronously maintained** — hence eventually consistent (§5.2).
**Monitoring.** `ThrottledRequests` and `ConsumedCapacity` per table **and per GSI** (a throttled GSI throttles the base-table write); `SuccessfulRequestLatency`; **hot-partition indicators via CloudWatch Contributor Insights** — the only way to actually see key-level skew.
**Scaling.** On-demand or provisioned with autoscaling. Splits are automatic and one-way — **partitions never merge**, so a table that once spiked keeps its partition count, which permanently divides your provisioned throughput. This is a genuinely surprising operational fact.
**What actually breaks.**
* **A low-cardinality partition key.** `status = "PENDING"` as a partition key puts everything on one partition, and no amount of provisioned capacity helps.
* **Throttling on a GSI silently throttling base-table writes.**
* **Stale GSI reads** treated as strongly consistent (§5.2).
* **Scan operations** in production — the anti-pattern equivalent of Vitess scatter queries.
* **Partition count inflation after a traffic spike**, permanently diluting per-partition capacity.
***
#### 6.4 Elasticsearch (fixed shard count, local secondary indexes) [#64-elasticsearch-fixed-shard-count-local-secondary-indexes]
**Problem it solves.** Distributed full-text search where each shard is a self-contained Lucene index.
**Why local indexes?** A search index *is* a secondary index; term-partitioning it globally would require cross-shard postings intersection on every multi-term query (§5.2). Elasticsearch chooses **document-partitioned**, accepting scatter/gather.
**How it works internally.** `shard = hash(routing) % number_of_primary_shards` — the **fixed shard count** of §3.2b, chosen at index creation and **immutable** without a reindex. A search **scatters to one copy of every shard**, each computes local top-K, and the coordinating node **gathers and merges**. Scoring requires a second phase (query-then-fetch) because term statistics are per-shard.
**Monitoring.** Shard count per node; **search latency broken into query vs fetch phase**; **the slowest shard per query** (tail latency amplification made visible); rejected search-thread-pool tasks; skew in doc count per shard.
**Scaling.** More nodes redistributes shards but **does not increase query throughput for a query that must touch every shard** — §5.1's exact warning. Increasing parallelism requires more *indices* (e.g. time-based indices) plus routing so queries hit fewer of them.
**What actually breaks.**
* **Oversharding** — the fixed-count guess made too high, so every query fans out to hundreds of shards and tail latency dominates.
* **Undersharding** — the guess made too low, so shards grow past \~50 GB and you need a full reindex to change it.
* **Custom routing forgotten on the query side**, so a routed write is searched with a full scatter.
* **Deep pagination** across shards — every shard must produce `from+size` hits.
* **One slow node making every query slow**, because every query touches every shard.
***
# 7.11 Terminology introduced here (/docs/ddia/sharding/terminology-introduced-here)
# 7.9 Worked examples (/docs/ddia/sharding/worked-examples)
**① Why `mod N` is so bad.** With N nodes and a key space of K keys, moving from N to N+1 nodes:
* `mod N`: a key stays put only if `h mod N == h mod (N+1)` — **roughly `K/(N+1)` keys stay, so \~`N/(N+1)` of ALL keys move.** Going 3→4 nodes moves **\~75%** of the data.
* **Fixed shards / consistent hashing:** only **`K/(N+1)`** keys move — going 3→4 moves **\~25%**, which is the theoretical minimum (the new node's fair share).
**② Choosing a fixed shard count.** You expect to grow from 3 nodes to at most 60 nodes.
* Minimum shards = 60 (you can't have more nodes than shards)
* Pick a **highly divisible** number: **720** = 2⁴·3²·5, divisible by 1,2,3,4,5,6,8,9,10,12,15,16,18,20,24,30,36,40,45,48,60… so it splits evenly across almost any node count.
* At 3 nodes: 240 shards/node. At 60 nodes: 12 shards/node.
* If total data reaches 7.2 TB, each shard is **10 GB** — in the "just right" band. At 72 TB each shard is 100 GB and **rebalancing/recovery become expensive** — that's your resharding trigger.
**③ Hot key salting arithmetic.** A celebrity key takes 50,000 writes/s and 200,000 reads/s; a shard handles 10,000 ops/s.
* Unsalted: **one shard sees 250,000 ops/s → 25× over capacity.**
* Salt with 2 digits (100 keys): writes become 500/s per salted key across up to 100 shards ✔. **But every read must query all 100 keys**, so total read work becomes 200,000 × 100 = **20,000,000 key-reads/s** — catastrophically worse.
* **Correct design:** salt writes, and serve reads from a **cache or a materialized aggregate**, not by gathering 100 keys. This is why the book says salting splits *only* the write load.
**④ Scatter-query cost.** 100 shards, each shard's p99 = 10 ms, p50 = 2 ms. A scatter query waits for the **slowest** shard.
* P(all 100 shards respond under their p99) = 0.99¹⁰⁰ ≈ **36%**
* So **\~64% of scatter queries take ≥10 ms**, i.e. **the cluster's p50 for scatter queries ≈ each shard's p99.** That is tail latency amplification, and it's why local secondary indexes limit scalability.
**⑤ Global index write amplification.** A document with 40 indexed terms, global index sharded over 20 shards. One document write may touch **up to min(40, 20) = 20 index shards** plus 1 data shard. Either you run a 21-participant distributed transaction (slow, Ch 8) or you accept asynchronous, stale indexes (DynamoDB's choice).
**⑥ Partition size and bucketing.** IoT: 10,000 sensors × 1 reading/s × 200 bytes.
* Partition key `sensor_id`: each partition grows **17.3 MB/day → 6.3 GB/year** — over the 10 GB "big partition" threshold in under two years, and unbounded thereafter.
* Partition key `(sensor_id, yyyy-mm)`: each partition caps at **\~520 MB** ✔, and `WHERE sensor_id = X AND ts BETWEEN …` still hits one partition per month.
* **Cost:** a query spanning a year now touches 12 partitions instead of 1 — the deliberate trade.
***
# 4.3 B-Trees (/docs/ddia/storage-retrieval/b-trees)
**Introduced 1970; called "ubiquitous" less than 10 years later.** They remain **the standard index implementation in almost all relational databases**, and many nonrelational ones.
**Similarity to SSTables:** keys sorted → efficient key-value lookups and range queries.
**Where it ends:** *very* different design philosophy.
| | Log-structured | B-tree |
| ---------- | --------------------------------------- | ------------------------------------------------------------ |
| Unit | **variable-size segments**, several MB+ | **fixed-size blocks or pages** |
| Mutability | **written once, then immutable** | **may overwrite a page in place** |
| Page size | — | traditionally 4 KiB; **PostgreSQL uses 8 KiB, MySQL 16 KiB** |
**Pages reference each other by page number** — like a pointer, but on disk. If all pages are in one file, **page number × page size = byte offset.**
#### 3.1 The lookup [#31-the-lookup]
**Branching factor** = number of child-page references in one page. **Typically several hundred** (it depends on the space needed for page references and range boundaries).
> **The tree stays balanced: a B-tree with n keys always has depth O(log n). Most databases fit in a B-tree three or four levels deep.**
>
> **Worked example from the book: a four-level tree of 4 KiB pages with branching factor 500 stores up to 500⁴ × 4 KiB ≈ 250 TB.**
*(This structure is technically a B+ tree, but the book doesn't distinguish it from other B-tree variants.)*
#### 3.2 Writes and page splits [#32-writes-and-page-splits]
* **Update an existing key:** find the leaf page containing it, **overwrite that page on disk** with a version containing the new value.
* **Add a new key:** find the page whose range encompasses it and add it. **If there isn't enough free space, split the page into two half-full pages and update the parent** to account for the new subdivision.
**Deleting keys (which may require nodes to be merged) is more complex.**
#### 3.3 Making B-trees reliable [#33-making-b-trees-reliable]
> **The basic write operation is to OVERWRITE A PAGE ON DISK.** It's assumed the overwrite doesn't change the page's location, so all references to it remain intact. **This is in stark contrast to LSM-trees, which only append and eventually delete, never modifying files in place.**
**Two ways this goes wrong:**
| Failure | Result |
| ------------------------------------------------------- | ------------------------------------------------------------------------------ |
| Crash mid-split (several pages must be written at once) | **Corrupted tree** — e.g. **an orphan page that is not a child of any parent** |
| Hardware can't atomically write an entire page | **Torn page** — partially written |
**The fix: a write-ahead log (WAL).** An **append-only file to which every B-tree modification must be written BEFORE it is applied to the tree pages.** On restart after a crash, the log restores the B-tree to a consistent state. *(In filesystems the equivalent is called journaling.)*
**Performance interaction:** implementations **don't immediately write every modified page to disk — they buffer pages in memory first.** The WAL is what makes that safe. **As long as data has been written to the WAL and flushed with `fsync`, it will be durable.**
#### 3.4 B-tree variants worth knowing [#34-b-tree-variants-worth-knowing]
| Variant | Idea |
| -------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Copy-on-write** (LMDB) | Instead of overwriting + WAL, write the modified page **to a different location** and create **a new version of the parent pages pointing at it.** Also useful for **concurrency control** (→ snapshot isolation, Ch 8) |
| **Key abbreviation** | Don't store the entire key. **Interior keys need only enough information to act as boundaries between ranges.** Packing more keys per page → **higher branching factor → fewer levels** |
| **Sequential leaf layout** | Lay out leaf pages **in sequential order on disk** to speed up range scans and reduce seeks. **Difficult to maintain as the tree grows** |
| **Sibling pointers** | Each leaf page references its **left and right siblings**, so you can scan keys in order **without jumping back to parent pages** |
***
# 4.4 Comparing B-Trees and LSM-Trees (/docs/ddia/storage-retrieval/comparing-b-trees-lsm)
> **Rule of thumb: LSM-trees are better for write-heavy applications; B-trees are faster for reads.**
>
> **But benchmarks are often sensitive to workload details — you need to test with YOUR workload.** And it's **not a strict either/or**: storage engines sometimes blend both (e.g. multiple B-trees merged LSM-style).
#### 4.1 Read performance [#41-read-performance]
| | B-tree | LSM |
| ---------------- | -------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Point lookup** | Read one page per level; **few levels ⇒ fast and predictable** | Often must check **several SSTables at different compaction stages**; **Bloom filters reduce the disk I/O** |
| **Range query** | **Simple and fast** — uses the sorted tree directly | Can use SSTable sorting, but must **scan all segments in parallel and combine results.** **Bloom filters DON'T help** (you'd need the hash of every possible key in the range — impractical) ⇒ **range queries are more expensive than point queries in LSM** |
**LSM write-throughput hazard:** **high write throughput can cause latency spikes if the memtable fills up** — data can't be written to disk fast enough, perhaps because **compaction can't keep up with incoming writes.** Many engines including **RocksDB apply backpressure: they suspend all reads and writes until the memtable has been written out.**
**Modern SSDs**, especially **NVMe** (PCIe bus rather than SATA), **perform many independent read requests in parallel.** Both structures can deliver high read throughput, but **the engine must be carefully designed to exploit that parallelism.**
#### 4.2 Sequential vs random writes [#42-sequential-vs-random-writes]
**Disks have higher sequential write throughput than random**, so **a log-structured engine can generally handle higher write throughput on the same hardware.** The difference is **particularly big on spinning disks; on SSDs it's smaller but still noticeable.**
**Why SSDs still prefer sequential writes** (this is the sub-section people most often skip and most often need):
* Flash can be **read or written one page at a time (typically 4 KiB)** but **erased only one block at a time (typically 512 KiB).**
* Before erasing a block, the controller **must move pages containing valid data into other blocks** — **garbage collection (GC)**.
* **Sequential workload:** large chunks written at once → **a whole 512 KiB block likely belongs to a single file** → when deleted, **the whole block can be erased with no GC.**
* **Random workload:** a block **more likely contains a mixture of valid and invalid pages**, so **GC must do more work before erasing.**
* Two consequences: **GC's write bandwidth is not available to the application**, and **the extra writes wear the flash** — **random writes wear out the drive faster than sequential writes.**
#### 4.3 Write amplification [#43-write-amplification]
> **Write amplification = (total bytes written to disk in a workload) ÷ (bytes you'd write with a plain append-only log and no index).** *(Sometimes defined in I/O operations rather than bytes.)*
**Where the extra writes come from:**
| Engine | Writes per logical write |
| ---------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **LSM** | (1) the WAL for durability, (2) the memtable flush to disk, (3) **again every time the pair is part of a compaction** |
| **B-tree** | (1) the WAL, (2) the tree page itself — **plus sometimes writing out an entire page even if only a few bytes changed**, to guarantee correct recovery after a crash or power failure |
**Optimization for LSM:** if values are much larger than keys, **store values separately from keys and compact only the SSTables containing keys + value references** (this is WiscKey-style key-value separation, used by RocksDB's BlobDB and TiKV's Titan).
**Why it matters:** in write-heavy applications the bottleneck may be the rate at which the database can write to disk. **The higher the write amplification, the fewer writes/second within the available disk bandwidth.** It also **determines SSD wear.**
**Which is better depends on:** key/value lengths, and **how often you overwrite existing keys vs insert new ones.** **For typical workloads, LSM-trees tend to have LOWER write amplification** because they **don't have to write entire pages and can compress chunks of the SSTable.**
> ⚠️ **Benchmarking warning:** run the experiment **long enough that write amplification becomes visible.** **When writing to an empty LSM-tree there are no compactions yet, so all disk bandwidth is available for new writes.** As the database grows, **new writes must share bandwidth with compaction.** Short benchmarks systematically flatter LSM engines.
#### 4.4 Disk space usage [#44-disk-space-usage]
| | B-tree | LSM |
| --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Fragmentation** | **Yes.** Deleting many keys leaves unused pages in the middle of the file; they can be reused but **can't easily be returned to the OS** → needs a background process to move pages around, **e.g. PostgreSQL's vacuum** | **Less of a problem** — compaction periodically rewrites the files anyway, and **SSTables have no pages with unused space** |
| **Compression** | — | **Blocks of key-value pairs compress better in SSTables, often resulting in smaller files than B-trees** |
| **Space overhead** | — | Overwritten/deleted keys **consume space until removed by compaction** — **low with leveled compaction; size-tiered uses more, especially temporarily during compaction** |
| **Verified deletion** | — | **A deleted record may still exist in higher levels until the tombstone has propagated through all compaction levels, which might take a long time** — a genuine problem for data-protection compliance. **Specialist designs can propagate deletions faster** |
| **Snapshots** | **Harder** — pages are overwritten | **Easy**: write out the memtable, **record which segment files existed at that time. As long as you don't delete those files, you don't need to copy anything** |
***
# 4.7 Data Storage for Analytics (/docs/ddia/storage-retrieval/data-storage-analytics)
Data warehouses are usually **relational**, because SQL fits analytical queries well, and many graphical tools generate SQL and support **drill-down** and **slicing and dicing**.
> **On the surface a warehouse and an OLTP database look similar — both have a SQL interface. But the internals look quite different, because they're optimized for very different query patterns.** Many vendors now focus on one or the other.
HTAP systems (SQL Server, SAP HANA, SingleStore) support both — **but they are increasingly becoming two separate storage and query engines that happen to be accessible through a common SQL interface.**
#### 7.1 The unbundled cloud warehouse [#71-the-unbundled-cloud-warehouse]
Established vendors (Teradata, Vertica, SAP HANA) offer on-prem and cloud. Cloud-only warehouses (**BigQuery, Redshift, Snowflake**) exploit **object storage and serverless computation**, integrate with cloud services (automatic log ingestion; Dataflow, Kinesis), and are **more elastic because they decouple query computation from the storage layer** — data in object storage rather than local disks, so **storage capacity and query compute are adjusted independently.**
**Open source warehouses have broken apart into four separable components** — this is one of the most useful mental models in the chapter:
#### 7.2 Column-oriented storage [#72-column-oriented-storage]
**The motivating query** (Example 4-1): are people more inclined to buy fresh fruit or candy depending on the day of the week?
```sql
SELECT dim_date.weekday, dim_product.category,
SUM(fact_sales.quantity) AS quantity_sold
FROM fact_sales
JOIN dim_date ON fact_sales.date_key = dim_date.date_key
JOIN dim_product ON fact_sales.product_sk = dim_product.product_sk
WHERE dim_date.year = 2024
AND dim_product.category IN ('Fresh fruit', 'Candy')
GROUP BY dim_date.weekday, dim_product.category;
```
**It touches a huge number of rows but only THREE columns** of `fact_sales`: `date_key`, `product_sk`, `quantity`. Fact tables are **often over one hundred columns wide**, and **`SELECT *` is rarely needed for analytics.**
**Why row storage fails here:** even with indexes on `date_key`/`product_sk`, a row-oriented engine **still loads all those rows (each with 100+ attributes) from disk into memory, parses them, and filters.**
> **The layout relies on each column storing the rows in THE SAME ORDER.** To reassemble a row, take the 23rd entry from each column.
**In practice engines don't store a whole column (trillions of rows) in one go.** They **break the table into blocks of thousands or millions of rows, and within each block store each column separately.** Since many queries are restricted to a date range, **it's common to make each block contain rows for a particular timestamp range** — then the query loads only the needed columns from the blocks overlapping the required range. *(This is zone-map / block-skipping territory.)*
**Where it's used:** almost all analytical databases — **Snowflake**, **DuckDB** (single-node embedded), **Pinot** and **Druid** (product analytics); formats **Parquet, ORC, Lance, Nimble**; in-memory formats **Apache Arrow, Pandas/NumPy**; time-series **InfluxDB IOx, TimescaleDB**.
*Note: columnar applies to non-relational data too — **Parquet supports a document data model** (based on Google's **Dremel**) using **shredding/striping**.*
> ⚠️ **Do NOT confuse column-oriented databases with the WIDE-COLUMN (column-family) data model**, where a row can have thousands of columns and rows needn't share columns. **Despite the name, wide-column databases are ROW-oriented — they store all values from a row together.** Examples: **Bigtable, Accumulo, HBase.**
#### 7.3 Column compression [#73-column-compression]
**Why this is so effective:** **the number of distinct values in a column is often small compared to the number of rows** (a retailer has billions of sales but only 100,000 distinct products).
**Roaring bitmaps** switch between the plain-bitmap and run-length representations, **using whichever is more compact.**
**Why bitmaps are perfect for warehouse queries:**
**Bitmaps also answer graph queries** — e.g. find all users of a social network **followed by user X who also follow user Y.**
#### 7.4 Sort order in column storage [#74-sort-order-in-column-storage]
**Rows can be stored in insertion order** (then inserting = appending to each column). **But you can impose an order and use it as an indexing mechanism**, as with SSTables.
> ⚠️ **Sorting each column INDEPENDENTLY would make no sense** — you'd no longer know which items belong to the same row. **The data must be sorted an entire row at a time, even though it's stored by column.**
**Choosing sort keys:**
* **First sort key**: pick from knowledge of common queries. If queries often target date ranges (last month), make **`date_key` first** — the query then scans only last month's rows.
* **Second sort key** determines order among rows tying on the first. If `date_key` is first, **`product_sk` second** groups all sales of the same product on the same day together, helping queries that group or filter by product within a date range.
**Sorting also boosts compression — but unevenly:**
> **The compression effect is strongest on the first sort key.** Still, **having the first few columns sorted is a win overall.**
#### 7.5 Writing to column-oriented storage [#75-writing-to-column-oriented-storage]
**Writes in a warehouse tend to be bulk imports, often via ETL.**
**Why individual row inserts are terrible here:** writing a row in the middle of a sorted table means **rewriting all the compressed columns from the insertion position onward.** But **a bulk write of many rows at once amortizes that cost.**
**The solution is the same log-structured approach as LSM:**
Done by **Snowflake, Vertica, Apache Pinot, Apache Druid**, and many others.
#### 7.6 Query execution: compilation vs vectorization [#76-query-execution-compilation-vs-vectorization]
A complex analytical SQL query becomes a **query plan of stages called operators**, possibly **distributed across machines for parallel execution.** The planner optimizes **which operators, in what order, and where each runs.**
**The problem:** for queries scanning millions of rows, it's not just disk bytes — it's **CPU time.** The simplest operator is **an interpreter**: while iterating over each row, it consults a data structure representing the query to find out which comparisons/calculations to perform on which columns. **Too slow for analytics.**
**Two solutions, both used in practice:**
| Approach | Mechanism |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Query compilation** | The engine **generates code** for executing the query — iterate rows, look at the columns of interest, perform comparisons, copy values to an output buffer if conditions hold. **Compile that generated code to machine code (often via LLVM) and run it on column-encoded data in memory.** Similar to **JIT compilation in the JVM.** |
| **Vectorized processing** | The query is **interpreted, not compiled**, but made fast by **processing many values from a column in a batch** rather than row by row. A **fixed set of predefined operators is built into the database**; pass arguments, get back a batch of results. |
**Vectorization example:**
**Both exploit the same four CPU characteristics** (this list is the real payoff of the section):
1. **Preferring sequential memory access over random access, to reduce cache misses**
2. **Doing most work in tight inner loops** — few instructions, **no function calls** — to keep the instruction pipeline busy and **avoid branch mispredictions**
3. **Parallelism: multiple threads and SIMD instructions**
4. **Operating directly on COMPRESSED data without decoding it into a separate in-memory representation**, saving memory allocation and copying
#### 7.7 Materialized views and data cubes [#77-materialized-views-and-data-cubes]
| | Virtual view | Materialized view |
| ------------------------- | ----------------------------------------------------------------------------------------------- | -------------------------------------------------------- |
| What it is | **A shortcut for writing queries** | **An actual copy of the query results, written to disk** |
| On read | SQL engine **expands it into the underlying query on the fly** and processes the expanded query | Read the stored copy |
| On underlying data change | nothing to do | **Must be updated** — more work on writes |
Some databases update them automatically; **Materialize** is a system specializing in materialized view maintenance.
**Materialized aggregates / data cubes (OLAP cubes).** Warehouse queries often use `COUNT`, `SUM`, `AVG`, `MIN`, `MAX`. If many queries use the same aggregates, **crunching raw data each time is wasteful.** A data cube **creates a grid of aggregates grouped by different dimensions.**
| | Data cube |
| ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| ✔ **Advantage** | **Certain queries become very fast because they're effectively precomputed.** Total sales per store yesterday = read a total along one dimension — **no need to scan millions of rows** |
| ✗ **Disadvantage** | **No flexibility.** You **cannot** compute "what proportion of sales came from items costing more than $100," **because price isn't one of the dimensions** |
> **Most data warehouses therefore keep as much raw data as possible and use aggregates like data cubes only as a performance boost for certain queries.**
***
# 4.11 Decision cheat sheet (/docs/ddia/storage-retrieval/decision-cheat-sheet)
**LSM or B-tree?**
Write-heavy, high ingest, sequential-write-friendly, want cheap snapshots and lower write amplification → **LSM**. Read-heavy, latency-predictability matters, lots of range scans, mature transactional needs → **B-tree**. Test with your workload — and **run the test long enough for compaction to reach steady state.**
**Do I add this index?**
Only if a real query pattern needs it. Every index costs disk and slows every write. Prefer **covering indexes** when a hot query can be answered from the index alone; accept the extra space knowingly.
**Row or column store?**
Point lookups and updates of individual records → row. **Aggregates over many rows touching few columns → column.** If a query would touch >5–10% of rows, columnar wins; if it touches a handful of rows by key, row wins.
**How do I pick the sort/order key for a columnar table?**
Put the **most commonly filtered range predicate first** (usually time). Second key = the next most common grouping dimension. Remember: **compression benefit is concentrated in the first key**, and this choice is effectively irreversible without a full rewrite.
**Concatenated index or multidimensional?**
Concatenated works when queries always constrain a **prefix** of the columns. If you need two independent ranges simultaneously (lat *and* lon, date *and* temperature), you need a **real multidimensional index** (R-tree/BKD/Z-order).
**Full-text or vector search?**
Exact terms, names, codes, filters, and explainability → **inverted index**. Meaning-level matching, paraphrase, cross-lingual → **vectors**. In practice, **hybrid** (BM25 + vector, fused) beats either alone for most product search.
**In-memory database?**
When the dataset fits in RAM *and* you need either the lowest possible latency or **data structures that are awkward on disk** (Redis's sorted sets, priority queues). Remember the speedup comes from **avoiding disk-encoding overhead**, not from avoiding disk reads.
***
# 4.15 Forward links (/docs/ddia/storage-retrieval/forward-links)
| Concept here | Where it's developed |
| ------------------------------------------------- | ------------------------------------------- |
| Durability, `fsync`, crash recovery | **Ch 8** — Transactions |
| Copy-on-write B-trees for MVCC | **Ch 8** — Snapshot Isolation |
| Spreading a storage engine across machines | **Ch 6** (Replication), **Ch 7** (Sharding) |
| Maintaining materialized views incrementally | **Ch 12** — Stream Processing |
| Object storage as the segment-file substrate | **Ch 1**, **Ch 11** |
| Star/snowflake schemas the columnar layout serves | **Ch 3** |
| Parquet/Avro as encoding formats | **Ch 5** — Encoding and Evolution |
# 4. Storage and Retrieval (/docs/ddia/storage-retrieval)
> "A computer does not primarily compute in the sense of doing arithmetic. \[…] They primarily are filing systems." — Richard Feynman
**Ch 3 was the user's view** (what format you give the database, what interface you query it through). **Ch 4 is the database's view**: how it stores what you give it, and how it finds it again.
**Why you should care even though you'll never write a storage engine:** you must **select** the right one from many available, and **to configure it to perform well on your workload you need a rough idea of what it's doing under the hood.**
The chapter's map:
***
# 4.6 Keeping everything in memory (/docs/ddia/storage-retrieval/keeping-everything-memory)
Everything so far is **an answer to the limitations of disks.** We tolerate the awkwardness because disks are **durable** and have **lower cost per gigabyte than RAM.**
**As RAM gets cheaper, the cost-per-GB argument erodes. Many datasets are simply not that big.** Hence in-memory databases.
| System | Durability approach |
| ---------------------------------------- | ------------------------------------------------------------------------------------------------ |
| **Memcached** | **Caching only** — acceptable to lose data on restart |
| **VoltDB, SingleStore, Oracle TimesTen** | In-memory **relational**; vendors claim big gains from removing on-disk data structure overheads |
| **RAMCloud** | Open source in-memory KV store **with durability**, log-structured for both memory and disk |
| **Redis, Couchbase** | **Weak durability** — asynchronous disk writes |
Durability is achieved by **special hardware (battery-powered RAM)** or, more commonly, **writing a change log to disk, writing periodic snapshots, or replicating the in-memory state to other machines.**
> **These are still "in-memory" databases because the disk is merely an append-only log for durability, and reads are served entirely from memory.** Writing to disk also brings operational advantages: **files can be backed up, inspected, and analyzed by external utilities.**
**⚠️ The counterintuitive point most people get wrong:**
> **The performance advantage of in-memory databases is NOT because they avoid reading from disk.** A disk-based engine may never read from disk either, if you have enough memory, **because the OS caches recently used disk blocks anyway.**
>
> **They are faster because they avoid the overheads of ENCODING in-memory data structures into a form that can be written to disk.**
**A second, under-appreciated use case:** in-memory databases can offer **data models that are difficult to implement with disk-based indexes.** **Redis offers a database-like interface to priority queues and sets — its implementation is comparatively simple because everything is in memory.**
***
# 4.2 Log-Structured Storage (/docs/ddia/storage-retrieval/log-structured-storage)
#### 2.1 Step 1: hash index in memory [#21-step-1-hash-index-in-memory]
Keep an in-memory hash map: **key → byte offset of its most recent value in the log.**
Write = append to log **and** update the map. Read = look up offset, seek, read. **If that part of the file is already in the filesystem cache, a read requires no disk I/O at all.**
**Four problems, and each one motivates the next design step:**
| Problem | Consequence |
| ------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------- |
| Old log entries are never freed | **You might run out of disk space** |
| The hash map isn't persisted | **Rebuild on restart** by scanning the whole log → slow restarts with a lot of data |
| **The hash table must fit in memory** | On-disk hash maps perform badly: **lots of random-access I/O, expensive to grow when full, and hash collisions require fiddly logic** |
| **Range queries are inefficient** | Can't scan keys 10000–19999; you must look up each key individually |
> **In practice, hash tables are not used very often for database indexes.** It is much more common to keep data **sorted by key**.
#### 2.2 SSTables — Sorted String Tables [#22-sstables--sorted-string-tables]
Same key-value pairs, but **sorted by key**, and **each key appears only once in the file.**
**How the sparse index works:** group pairs into **blocks of a few kilobytes** and store **only the first key of each block** in the index. Looking for `handiwork`, which isn't in the sparse index: **because of the sorting you know it must lie between `handbag` and `handsome`**, so seek to `handbag`'s offset and **scan forward.** *A block of a few kilobytes can be scanned very quickly.*
The sparse index itself lives in a separate part of the SSTable — implemented as an **immutable B-tree, a trie, or similar.**
**Consequence: you no longer need all keys in memory.** That kills problem #3 above.
#### 2.3 Constructing and merging SSTables — the LSM algorithm [#23-constructing-and-merging-sstables--the-lsm-algorithm]
**The problem SSTables create:** you can't just append, or the file stops being sorted. Rewriting the whole SSTable per insert would be far too expensive.
**The solution — a hybrid of an append-only log and a sorted file:**
**The four steps:**
1. **Write → in-memory ordered map (the memtable)**: red-black tree, skip list, or trie — structures that let you insert in any order, look up efficiently, and **read back in sorted order.**
2. **When the memtable exceeds a threshold (typically a few MB), write it out to disk in sorted order as an SSTable** — the **most recent segment**, a separate file alongside older ones, each with its own index. **While it's being written, the database keeps writing to a new memtable instance**; the old memtable's memory is freed on completion.
3. **Read**: try the memtable, then the most recent on-disk segment, then next-older, etc. **If the key is in no segment, it doesn't exist in the database.**
4. **Background merge & compaction** combines segments and discards overwritten/deleted values.
**Merging works like mergesort:** read the input files side by side, look at the first key in each, **copy the lowest key to the output, repeat. If the same key appears in more than one input file, keep only the more recent value.** Output is sorted, one value per key, **and it uses minimal memory because you iterate one key at a time.**
**Crash safety of the memtable:** a **separate log on disk to which every write is immediately appended.** It's **not sorted by key — that doesn't matter, because its only purpose is to restore the memtable after a crash.** Every time the memtable is written out, the corresponding part of the log can be discarded.
**Deletion = a tombstone.** Append a special deletion record. **When segments are merged, the tombstone tells the merge process to discard any previous values for that key. Once the tombstone is merged into the OLDEST segment, it can be dropped.**
**Provenance:** this is essentially what **RocksDB, Cassandra, ScyllaDB, and HBase** do, all inspired by **Google's Bigtable paper** (which introduced the terms *SSTable* and *memtable*). Published in **1996 as the Log-Structured Merge-tree (LSM-tree)**, building on earlier work on **log-structured filesystems**.
**Key properties that fall out of immutability:**
* A segment file is **written in one pass and thereafter immutable.**
* **Merging happens in a background thread**, and **reads continue to be served from the input segments during the merge.** When it completes, **switch reads to the merged segment, then delete the inputs.**
* **Segment files needn't be on local disk — they're well suited to object storage.** SlateDB and Delta Lake take this approach.
* **Crash recovery is simple:** crash during a memtable flush or a merge → **just delete the unfinished SSTable and start afresh.** The WAL may contain incomplete records (crash mid-record, or disk full) — **detected via checksums and discarded.**
#### 2.4 Bloom filters [#24-bloom-filters]
**The problem:** reading a key last updated long ago, or **a key that doesn't exist**, is slow — the engine must check several segment files.
**The fix:** a **Bloom filter per segment** — a fast, *approximate* check of whether a key appears in an SSTable.
* **If at least one bit is 0 → the key definitely does NOT appear.** Safe to skip the SSTable.
* **If all bits are 1 → the key is likely present** — but possibly all those bits were set by *other* keys. That's a **false positive**.
**Cost of a false positive is small:** consult the sparse index, decode the block, discover it's not there, **continue the search with the next-oldest segment.** A bit of unnecessary work; no harm done.
**Sizing rule of thumb (worth memorizing):**
> **\~10 bits of Bloom filter per key ⇒ \~1% false-positive probability, and the probability drops TENFOLD for every 5 additional bits per key.**
So: 10 bits/key → 1%; 15 bits/key → 0.1%; 20 bits/key → 0.01%.
Checks use **bitwise operations all CPUs support** — extremely fast. The filter is **small compared to the rest of the SSTable.**
#### 2.5 Compaction strategies [#25-compaction-strategies]
**When to compact, and which SSTables to include** — usually configurable.
| Strategy | Best for |
| --------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| **Size-tiered** | **Mostly writes, few reads** |
| **Leveled** | **Read-dominated workloads**; also good when you **write a small number of keys frequently and a large number of keys rarely** |
**Most LSM implementations provide a variety of strategies for different workloads.**
> **The basic idea — keeping a cascade of SSTables merged in the background — is simple and effective.**
#### 2.6 Aside: embedded storage engines [#26-aside-embedded-storage-engines]
Not all databases are network services. **Embedded databases are libraries running in the same process as your application**, reading/writing local files, invoked by normal function calls.
Examples: **RocksDB, SQLite, LMDB, DuckDB, KùzuDB.**
* Very common in **mobile apps** for local user data
* On the backend: appropriate **if the data fits on a single machine and there aren't many concurrent transactions**
* **Multitenant pattern:** if each tenant is small and completely separate (**you never run queries combining data from multiple tenants**), you can use **a separate embedded database instance per tenant**
***
# 4.8 Multidimensional and Full-Text Indexes (/docs/ddia/storage-retrieval/multidimensional-full-text-indexes)
#### 8.1 Concatenated vs multidimensional [#81-concatenated-vs-multidimensional]
**Concatenated index** = combine several fields into one key by appending one column to another (order specified in the index definition). **Like an old-fashioned paper phone book: an index from `(lastname, firstname)` to phone number.**
**Why it fails for geospatial:**
```sql
SELECT * FROM restaurants WHERE latitude > 51.4946 AND latitude < 51.5079
AND longitude > -0.1162 AND longitude < -0.1004;
```
**Solutions:**
1. **Translate 2-D into a single number via a space-filling curve** (Z-order/Morton, Hilbert), then use a regular B-tree
2. **More commonly, specialized spatial indexes: R-trees or Bkd-trees**, which **divide the space so nearby points tend to be grouped in the same subtree.** **PostGIS implements geospatial indexes as R-trees using PostgreSQL's GiST facility**
3. **Regularly spaced grids of triangles, squares, or hexagons** (H3, S2)
**Multidimensional indexes are not just geographic:**
* Ecommerce: a **3-D index on (red, green, blue)** to search products in a color range
* Weather: a **2-D index on (date, temperature)** to find observations in a given year with temperature 25–30°C. With a 1-D index you'd have to **scan all records from that year and then filter by temperature, or vice versa.**
#### 8.2 Full-text search and the inverted index [#82-full-text-search-and-the-inverted-index]
**The reframing that makes it click:**
> **Full-text search is another kind of multidimensional query. Each term that might appear in a text is a DIMENSION.** A document containing term *x* has value 1 in dimension *x*; otherwise 0. Searching "red apples" = **a 1 in the `red` dimension AND simultaneously a 1 in the `apples` dimension.** **The number of dimensions may be very large.**
**Inverted index** = key-value structure where **key = a term, value = the list of IDs of all documents containing it (the postings list).**
**If document IDs are sequential numbers, the postings list can be a sparse bitmap** — the *n*th bit for term *x* is 1 if document *n* contains *x*.
**Real implementations:**
* **Lucene** (the engine behind **Elasticsearch and Solr**) **stores term → postings list in SSTable-like sorted files, merged in the background using the same log-structured approach** as §2. *(So Lucene is an LSM engine — that connection is worth internalizing.)*
* **PostgreSQL's GIN index** uses postings lists for full-text search **and for indexing inside JSON documents.**
**Two extensions:**
| Technique | How | Trade-off |
| ------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------- |
| **n-grams / trigrams** | Instead of breaking into words, index all substrings of length *n*. Trigrams of `hello` = `hel`, `ell`, `llo`. **Lets you search arbitrary substrings ≥3 chars, and even supports regular expressions in queries** | **The indexes are quite large** |
| **Edit distance / fuzzy** | Lucene stores the term set as **a finite state automaton over the characters in the keys (like a trie)** and transforms it into a **Levenshtein automaton**, supporting efficient search within a given edit distance (distance 1 = one letter added, removed, or replaced) | more complex |
**Note the honest scoping:** information retrieval is a big specialist topic involving **language-specific processing** — several **Asian languages are written without spaces or punctuation between words, so splitting text into words requires a model** that says which character sequences constitute a word — plus **synonyms and grammatical forms.** Beyond the book's scope.
#### 8.3 Vector embeddings and semantic search [#83-vector-embeddings-and-semantic-search]
**The problem:** a help page titled *"canceling your subscription"* should be findable by *"how to close my account"* or *"terminate contract"* — **close in meaning, completely different words.** Important for **retrieval-augmented generation (RAG)**, which incorporates search results into an LLM's output.
**How:** an **embedding model** (often an LLM) translates a document into a **vector embedding** — a vector of floating-point values. **The vector is a point in multidimensional space; each value is the document's location along one dimension's axis. Semantically similar inputs get vectors near each other.**
> ⚠️ **Terminology collision:** in **vectorized processing** (§7.6), a *vector* is **a batch of values processed with specially optimized code.** In **embedding models**, a *vector* is **an array of floats representing a location in multidimensional space.** Same word, different meanings.
**Model history:** early text models **Word2Vec, BERT, GPT**, usually implemented as **neural networks**. Then models for **video, audio, images**. More recently **multimodal**: one model generating embeddings for multiple modalities.
**Query flow:** user query + related context (e.g. the user's location) → embedding model → query vector → **vector index** returns documents whose vectors are closest.
**Three index types** (R-trees don't work well for high-dimensional vectors):
Implemented in **Facebook's Faiss** (several variants of each) and **PostgreSQL's pgvector** (both).
***
# 4.10 Production failure catalog for this chapter (/docs/ddia/storage-retrieval/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| --------------------------------------------------------- | ------------------------------------------------------------------------------------ |
| Write latency jumps 1 ms → 30 s, no code change | LSM **compaction can't keep up**; engine applied **backpressure** |
| Reads slow down over days, then recover after a restart | L0 file accumulation / read amplification |
| Query times out scanning a "small" table | **Tombstones** not yet compacted away |
| Deleted data still readable weeks later | Tombstone hasn't propagated through all compaction levels |
| Disk full during routine maintenance | **Size-tiered compaction** temp space |
| Benchmark 5× better than production | Benchmarked an **empty LSM** — no compaction yet |
| Range queries far slower than point queries | Bloom filters **don't help range queries** |
| Insert throughput collapses after switching to UUIDv4 PK | **Random writes** + constant **page splits** in a clustered B-tree |
| Table is 4× its logical size after a mass delete | **B-tree fragmentation**; needs vacuum/rebuild |
| Periodic latency spikes every few minutes | **Checkpoint** flushing buffered dirty pages |
| Everything fine until data exceeded RAM, then 100× slower | Working set no longer fits the **buffer pool / page cache** |
| Warehouse query costs $4,000 | No sort key / no partition pruning → **full scan** |
| Query planning takes longer than the query | **Small-file problem** — a footer read per file |
| `SELECT *` is 50× slower than naming 3 columns | Defeats **columnar** storage |
| Search cluster falls over with thousands of shards | **Oversharding** — every shard is a full Lucene index |
| Semantic search quality quietly degrades | **Recall collapse** from low `ef_search`/`nprobe`, or mixed embedding model versions |
| Vector search returns nothing when filtered | **Filtered ANN** — pre/post-filtering interacting badly with the graph |
***
# 4.5 Secondary indexes and where values live (/docs/ddia/storage-retrieval/secondary-indexes-where-values)
**Primary key** uniquely identifies one row / document / vertex; other records refer to it by that key, and the index resolves such references.
**Secondary indexes** (`CREATE INDEX`) let you search by other columns. **The key difference: indexed values are not necessarily unique** — many rows may share one index entry. Two solutions:
1. **Make each value in the index a list of matching row identifiers** — like a **postings list** in a full-text index
2. **Make each entry unique by appending a row identifier to it**
Both B-trees and log-structured storage can implement an index.
#### 5.1 Three ways to store the value [#51-three-ways-to-store-the-value]
**The heap-file update gotcha:** updating a value without changing the key **can overwrite in place, provided the new value is not larger.** If it *is* larger, the record **probably must move to a new heap location with enough space** — and then **either all indexes must be updated to point at the new location, or a forwarding pointer is left behind in the old location.**
*(This is exactly why PostgreSQL's HOT updates and MySQL's clustered-PK design exist, and why "update a TEXT column from 20 bytes to 2000 bytes" can be surprisingly expensive.)*
***
# 4.13 Self-test (/docs/ddia/storage-retrieval/self-test)
Why is `db_get` O(n) but `db_set` fast, and what single data structure fixes reads without changing the write path?
Name the four problems with an in-memory hash index over a log. Which one does an SSTable's **sparse index** solve, and how?
Why can't you simply append to an SSTable? Describe the memtable→segment→compaction cycle that resolves this.
What exactly does a Bloom filter guarantee, and what does it *not* guarantee? Why is a false positive harmless in an LSM engine?
You have a read-heavy workload with a large key space. Size-tiered or leveled compaction, and why?
A B-tree page split crashes halfway. What two forms of corruption can result, and what mechanism prevents them?
Why do SSDs — with no moving parts — still prefer sequential writes? Trace the answer through page size, block size, and GC.
Define write amplification. Give one reason LSM is usually lower and one reason a B-tree can be very high for a small update.
Why can taking a consistent snapshot be nearly free in an LSM engine and expensive in a B-tree?
What's the difference between a clustered index, a heap file, and a covering index? When does updating a heap-file row force every index to be rewritten?
Why are in-memory databases fast? (The answer is *not* "they avoid disk reads.")
Why does sorting rows improve compression most for the **first** sort key and barely at all for the fourth?
Why must a column store sort *whole rows*, never individual columns?
Explain how `WHERE product_sk = 30 AND store_sk = 3` is answered with two bitmaps. What property of columnar storage makes the bitwise AND valid?
Contrast query compilation with vectorized processing. Name two CPU characteristics both exploit.
Give one query a data cube answers instantly and one it fundamentally cannot answer. Why?
Why can a concatenated `(latitude, longitude)` index not answer a map-viewport query? Name two index types that can.
Reframe full-text search as a multidimensional query. What is a postings list, and why can it be a bitmap?
Why don't R-trees work for 1536-dimensional embeddings? Compare flat, IVF, and HNSW on accuracy and speed.
you're storing 10 TB of IoT sensor readings, written 500k rows/sec, queried as "average value per sensor per hour for the last 7 days." Choose a storage engine, a sort key, and a compaction/partition strategy — and justify each against the trade-offs in this chapter.
# 4.9 Technology deep dives (/docs/ddia/storage-retrieval/technology-deep-dives)
***
#### 9.1 LSM-tree engines (RocksDB, Cassandra, ScyllaDB, HBase, LevelDB) [#91-lsm-tree-engines-rocksdb-cassandra-scylladb-hbase-leveldb]
**Problem it solves.** Sustain very high write throughput on disks that strongly prefer sequential writes, while still supporting sorted key access and range scans.
**Why wasn't a B-tree enough?** B-trees convert scattered key writes into **random page overwrites**, plus they write **every value at least twice** (WAL + page), sometimes writing a whole page for a few changed bytes. On write-heavy workloads that caps throughput and burns SSD endurance.
**Why wasn't the hash-index-over-a-log enough?** The hash table must fit in memory, restarts require a full log scan, and **range queries are impossible.**
**How it works internally.** §2 in full: memtable → immutable sorted segments → background merge/compaction; WAL for memtable durability; tombstones for deletes; Bloom filter + sparse index per segment; size-tiered or leveled compaction.
**Deployment.** Embedded (RocksDB inside another system — Kafka Streams, TiKV, CockroachDB, MyRocks all embed it) or as a distributed database (Cassandra/Scylla: peer-to-peer ring, replication factor + tunable consistency; HBase: on HDFS with a master).
**Monitoring.**
* **Compaction debt / pending compaction bytes** — the single most predictive metric. If it climbs monotonically, you are losing and will eventually stall.
* **Write stalls / backpressure events** — RocksDB explicitly suspends writes when the memtable can't flush fast enough; count these.
* **Number of SSTables per level** and **L0 file count** (L0 overload is the classic read-amplification cliff)
* **Read amplification** — SSTables touched per lookup; **Bloom filter false-positive rate**
* **Space amplification** — bytes on disk ÷ logical bytes; especially during size-tiered compaction
* **Tombstone counts and droppable-tombstone ratio** (Cassandra) — this is where "deleted data still returns" and "queries suddenly time out" come from
**Scaling.** Horizontal via partitioning (Ch 7). Vertically: more memory for memtables + block cache, faster NVMe, and **more compaction threads** — but compaction competes with foreground I/O, so this is a genuine dial, not a free win.
**Backup.** Trivially good: **snapshot = flush the memtable, record which segment files exist, and hard-link them.** Files are immutable, so nothing needs copying until compaction wants to delete them. This is why RocksDB checkpoints and Cassandra snapshots are near-instant.
**What actually breaks in production.**
* **Compaction can't keep up** → L0 files pile up → read amplification explodes → the engine applies backpressure → **write latency goes from 1 ms to 30 s with no code change.** Almost always caused by write rate exceeding sustained compaction throughput, or too few compaction threads, or a saturated disk.
* **Tombstone accumulation.** Cassandra's canonical failure: a queue-like table where rows are written and deleted. A read scanning past 100,000 tombstones times out. Deletes are *writes*, and they're not free until they've propagated through every compaction level.
* **Range queries far slower than point queries**, because Bloom filters don't help and every segment must be scanned. People benchmark point lookups, deploy, then discover their scan workload.
* **Space amplification during size-tiered compaction** — merging four large tables needs temporary disk equal to their combined size. Running at 70% disk usage is how you get a full disk during a routine compaction.
* **Short benchmarks lie.** An empty LSM does no compaction, so early numbers are 3–5× the sustained figure. Benchmark past the point where the working set exceeds memory *and* compaction is in steady state.
* **"Deleted" data still on disk** for compliance purposes, until tombstones reach the oldest level.
***
#### 9.2 B-tree engines (PostgreSQL, MySQL/InnoDB, SQL Server, Oracle) [#92-b-tree-engines-postgresql-mysqlinnodb-sql-server-oracle]
**Problem it solves.** Predictable, low-latency point lookups and range scans by key, with in-place updates and mature transactional semantics.
**Why wasn't an LSM enough?** Reads must consult multiple segments; range queries can't use Bloom filters; compaction introduces background load and latency spikes; and the mature transaction/isolation machinery of relational databases grew up around page-based storage.
**How it works internally.** §3 in full: fixed-size pages, page numbers as on-disk pointers, branching factor in the hundreds, depth 3–4, splits propagating to the root, WAL for crash recovery, buffered dirty pages flushed later.
**Deployment.** Primary + replicas via WAL shipping (Ch 6). Buffer pool sized to a large fraction of RAM (Postgres conventionally 25% `shared_buffers` plus reliance on the OS page cache; InnoDB conventionally 70–80% in `innodb_buffer_pool_size` because it bypasses the OS cache).
**Monitoring.**
* **Buffer pool / cache hit ratio** — the cliff-edge metric; when the working set stops fitting, latency changes by orders of magnitude
* **WAL generation rate** and **checkpoint frequency** — a checkpoint storm is a classic latency spike
* **Index bloat and table bloat**; **vacuum progress** (Postgres) — §4.4's fragmentation problem, made operational
* **Page splits per second** (InnoDB) — high rates mean a bad insert pattern
* **`fsync` latency** — durability is bounded by it; a degraded disk shows up here first
* Lock waits, deadlocks, long-running transactions
**Scaling.** Read replicas; partitioning; bigger buffer pool. The dominant lever is **keeping the working set in memory**.
**Backup.** Base backup + **WAL archiving** → point-in-time recovery. **The WAL is what makes both replication and PITR possible** — same mechanism, two uses.
**What actually breaks.**
* **Random-insert index hotspots vs. monotonic keys.** A UUIDv4 primary key on InnoDB scatters inserts across the whole clustered index → random writes, constant page splits, poor fill factor. A monotonic key (UUIDv7, auto-increment) keeps inserts at the right edge — but then *that* page becomes a contention hotspot under high concurrency. Both directions have a failure mode.
* **Index bloat** after mass deletes: pages remain allocated in the middle of the file and **can't be returned to the OS** without a rebuild.
* **Torn pages** on hardware without atomic page writes — which is why Postgres has `full_page_writes` (and why it inflates WAL volume so much after each checkpoint).
* **Checkpoint spikes** — buffered dirty pages all flushed at once, saturating the disk and stalling foreground queries.
* **Too many indexes.** Every index multiplies write cost; the "just add an index" reflex silently halves write throughput.
* **Long-running transaction blocks vacuum** → bloat grows unbounded → eventually the table is mostly dead tuples and even indexed queries slow down.
***
#### 9.3 Columnar formats and engines (Parquet/ORC + Snowflake, BigQuery, ClickHouse, DuckDB) [#93-columnar-formats-and-engines-parquetorc--snowflake-bigquery-clickhouse-duckdb]
**Problem it solves.** Aggregate over billions of rows while reading only the few columns the query touches, and spend as little CPU per row as possible.
**Why wasn't a row store enough?** A 100+-column fact table forces a row store to read \~100× more bytes than the query needs, then parse and discard them.
**Why weren't indexes enough?** Indexes help you find *rows*; they don't help when you must touch a billion rows regardless. Analytical queries are scan-bound, not lookup-bound.
**How it works internally.** §7 in full: column chunks with per-column codecs; bitmap + RLE encodings; sorting to create long runs; block/row-group granularity with min/max statistics for skipping; LSM-style batched writes; JIT compilation or vectorized batch operators; SIMD; **operating directly on compressed data.**
**Parquet file anatomy** (worth knowing concretely):
**Deployment.** Files in object storage, a table format (Iceberg/Delta) for atomicity and time travel, a catalog for discovery, a query engine on top — the four-layer split of §7.1.
**Monitoring.** Bytes scanned per query (the cost metric); **row groups pruned vs read** (if pruning is 0%, your sort/partition key is wrong); file size distribution; spill-to-disk volume; per-operator time in the query profile.
**Scaling.** Partition + sort so that predicates prune; size row groups around 128 MB; scale compute independently of storage.
**Backup.** Files are immutable and versioned by the table format; the catalog is the thing that must be backed up.
**What actually breaks.**
* **The small-file problem.** A streaming writer producing a file per minute yields millions of tiny Parquet files. **Query planning time exceeds query time**, because the engine must read a footer per file. Requires scheduled compaction.
* **Wrong sort key**, chosen once and effectively permanent, so no query ever prunes and every query is a full scan.
* **`SELECT *` in a columnar store** — defeats the entire design and is often 50× slower than naming columns.
* **High-cardinality `GROUP BY` OOM** — the hash table doesn't fit; the engine spills or dies.
* **Row-at-a-time inserts** into ClickHouse/Druid → "too many parts", the columnar analogue of L0 overload.
* **Schema evolution mismatches** — a column added in 2025 doesn't exist in 2023 files; engines differ in whether that's a null or an error.
* **Timezone/type coercion** between writer and reader (INT96 timestamps, anyone) producing silently wrong dates.
***
#### 9.4 Full-text search engines (Lucene / Elasticsearch / OpenSearch) [#94-full-text-search-engines-lucene--elasticsearch--opensearch]
**Problem it solves.** Find documents by keywords appearing anywhere in the text, ranked by relevance, with tolerance for typos and grammatical variation.
**Why wasn't `LIKE '%term%'` enough?** It's a full scan with no index and no ranking. Why wasn't a B-tree on the text column enough? It indexes whole values, not the terms inside them.
**How it works internally.** Analysis chain (tokenize → lowercase → stopwords → stemming → synonyms) produces terms; **inverted index maps term → postings list**; postings stored in **SSTable-like sorted files merged in the background** — Lucene *is* an LSM engine, and its "segments" and "merges" are exactly §2's. Ranking via **BM25**. Fuzzy matching via a **Levenshtein automaton** over a trie of terms. Numeric/geo fields use BKD-trees.
**Deployment.** Index split into **shards** (each an independent Lucene index) with replicas; a coordinating node scatters the query to shards and gathers/merges results.
**Monitoring.** Segment count and merge queue; **JVM heap pressure and GC pauses** (the classic Elasticsearch failure); indexing rate vs refresh interval; search latency by phase (query vs fetch); **field data / fielddata cache size**; shard count per node.
**Scaling.** **Shard count is fixed at index creation** and is the single most consequential decision — too few and you can't scale; too many and per-shard overhead dominates (each shard is a full Lucene index with its own segments and merges). Reindex to change it.
**Backup.** Snapshot API to object storage; incremental because segments are immutable — the same property as §4.4.
**What actually breaks.**
* **Oversharding.** Thousands of tiny shards, each with segments and merge threads, consuming heap for metadata until the cluster falls over.
* **Mapping explosion** — dynamic mapping on user-supplied JSON keys creates thousands of fields, and the cluster state grows until updates stall.
* **Deep pagination** (`from: 100000`) — every shard must produce and sort 100,000+ hits. Use `search_after`.
* **Refresh interval vs indexing throughput** — a 1 s refresh during a bulk load creates enormous segment churn.
* **Split brain / red cluster** after a network blip with badly configured quorum settings (Ch 9–10 territory).
* **Using it as a system of record.** It's a derived data system; it has no transactions worth the name, and its answer to corruption is "reindex from the source."
***
#### 9.5 Vector databases (pgvector, Faiss, and dedicated stores) [#95-vector-databases-pgvector-faiss-and-dedicated-stores]
**Problem it solves.** Retrieve semantically similar items — the retrieval half of RAG — where lexical matching fails because the query and the document share no words.
**Why wasn't full-text search enough?** "How to close my account" and "canceling your subscription" have zero terms in common. Synonym lists don't generalize.
**Why wasn't an R-tree enough?** **R-trees don't work well for vectors with many dimensions** — above roughly 10–20 dimensions, space-partitioning indexes degenerate to full scans (the curse of dimensionality). Embeddings have 384–4096 dimensions.
**How it works internally.** §8.3: flat (exact, slow), IVF (partition by centroid, `nprobe` controls the accuracy/speed dial), HNSW (multi-layer proximity graph, greedy descent). Plus **quantization** — PQ/SQ compress vectors to a fraction of their size, trading recall for memory, which is what makes billion-scale indexes affordable.
**Deployment.** Either an extension on your existing database (pgvector — strongly preferred when the vector count is in the millions, because you keep transactions and joins with metadata) or a dedicated store when you need billions of vectors or heavy filtered search.
**Monitoring.** **Recall\@k measured against a flat/exact baseline** — this is the metric everyone forgets, and without it you cannot tell a broken index from a bad embedding model. Also: index build time and memory, query latency vs `nprobe`/`ef_search`, and index staleness after writes.
**Scaling.** HNSW memory is roughly `(dim × 4 bytes + M × 8 bytes) × N` — a million 1536-dim vectors is \~6 GB before graph overhead. Quantize, or shard.
**Backup.** The **embeddings are derived data** — regenerable from source documents, but regeneration costs real money in embedding-API calls, so back up the vectors. **Critically: back up which model version produced them.**
**What actually breaks.**
* **Silent recall collapse.** `ef_search`/`nprobe` set too low, so the index returns plausible-but-wrong neighbours. Nothing errors. Search quality quietly degrades and users just stop trusting it.
* **Mixing embeddings from two model versions in one index.** Vectors from different models are not comparable, and the results are noise. This happens every time someone upgrades an embedding model without a full reindex.
* **Filtered search performance.** "Nearest neighbours *where tenant\_id = X*" is the hard case: pre-filtering breaks the graph's connectivity, post-filtering may return nothing. This is the #1 real-world vector-search problem and most benchmarks ignore it.
* **HNSW deletes** — most implementations only tombstone, so a heavily updated index degrades until rebuilt.
* **Index build time** measured in hours, blocking a deploy nobody planned for.
* **Normalization mismatch** — cosine similarity on unnormalized vectors gives wrong rankings.
***
# 4.14 Terminology introduced here (/docs/ddia/storage-retrieval/terminology-introduced-here)
# 4.12 Worked examples (/docs/ddia/storage-retrieval/worked-examples)
**① B-tree capacity.** Page 4 KiB, branching factor 500, depth 4.
Leaf pages = 500³ = 125,000,000. Capacity = 125,000,000 × 4 KiB = **500 GB of leaf pages**; the book's figure of **\~250 TB** assumes 500⁴ leaf entries' worth of addressable data (4 levels of *references*). The point to retain: **depth 3–4 covers essentially any real database, so a lookup is 3–4 page reads.**
**② Bloom filter sizing.** You have 10 million keys per SSTable and want a 0.1% false-positive rate.
Rule: 10 bits/key → 1%; **each +5 bits/key divides FPP by 10.** So 0.1% needs **15 bits/key** = 150,000,000 bits = **\~18.8 MB** per SSTable. For 0.01%: 20 bits/key = **\~25 MB**.
**③ Write amplification.** An LSM with leveled compaction, 7 levels, fan-out 10.
A key is written: once to the WAL, once on memtable flush, and **once per level it is promoted through** — roughly 1 + 1 + 7 ≈ **10× write amplification** (real-world leveled RocksDB is typically 10–30×). Size-tiered is lower (\~4–10×) but has higher **space** amplification. A B-tree writing a full 8 KiB page for a 100-byte row change has **\~80× amplification for that write** — which is why `full_page_writes` after a checkpoint is so expensive.
**④ Columnar I/O saving.** Fact table: 1 billion rows × 100 columns × 8 bytes = 800 GB.
Query touches 3 columns → row store reads **800 GB**; column store reads 1e9 × 3 × 8 = **24 GB**, and with 4:1 compression **\~6 GB**. That's the **>100× difference** that makes interactive analytics possible.
**⑤ Bitmap compression.** `product_sk` with 100,000 distinct values over 1 billion rows.
Raw: 1e9 × 4 bytes = 4 GB. Bitmaps: 100,000 bitmaps × 1e9 bits = 12.5 TB uncompressed — **worse!** But each bitmap is \~99.999% zeros, so run-length encoding collapses it. **The lesson: bitmap indexes only pay off when combined with RLE/roaring, and the win grows as cardinality *falls*.** For a column with 5 distinct values sorted as the primary key, RLE takes it to a few kilobytes over a billion rows.
**⑥ HNSW memory.** 5 million documents, 1536-dim float32 embeddings, M=16.
Vectors: 5e6 × 1536 × 4 B = **30.7 GB**. Graph: 5e6 × 16 × 8 B ≈ **0.64 GB** per layer, \~1.3 GB total. **\~32 GB RAM** — before you've served a single query. With product quantization to 8× compression, \~4 GB. This is why "just add vector search" is a capacity-planning conversation.
***
# 4.1 The world's simplest database (/docs/ddia/storage-retrieval/world-s-simplest-database)
```bash
#!/bin/bash
db_set () { echo "$1,$2" >> database; }
db_get () { grep "^$1," database | sed -e "s/^$1,//" | tail -n 1; }
```
```bash
$ db_set 12 '{"name":"London","attractions":["Big Ben","London Eye"]}'
$ db_set 42 '{"name":"San Francisco","attractions":["Golden Gate Bridge"]}'
$ db_set 42 '{"name":"San Francisco","attractions":["Exploratorium"]}'
$ db_get 42
{"name":"San Francisco","attractions":["Exploratorium"]}
$ cat database
12,{"name":"London",...}
42,{"name":"San Francisco","attractions":["Golden Gate Bridge"]} ← old version kept
42,{"name":"San Francisco","attractions":["Exploratorium"]} ← tail -n 1 wins
```
**Every `db_set` appends.** Updates don't overwrite; **you find the latest value by looking at the LAST occurrence of a key** (hence `tail -n 1`).
| | Performance | Why |
| -------- | ---------------------------- | ------------------------------------------------------------------------------------------- |
| `db_set` | **Very good** | **Appending to a file is generally very efficient** — the simplest possible write operation |
| `db_get` | **Terrible** — **O(n)** | Scans the entire file every lookup. Double the records → double the time |
**Terminology note (important, used all book):** *log* here does **not** mean application logs. It means **an append-only sequence of records on disk.** Not necessarily human-readable; may be binary and purely internal.
Real databases add: **handling concurrent writes**, **reclaiming disk space so the log doesn't grow forever**, and **handling partially written records when recovering from a crash.** The principle is the same.
#### The index trade-off — the sentence that governs the whole chapter [#the-index-trade-off--the-sentence-that-governs-the-whole-chapter]
> **An index is an additional structure derived from the primary data.** Adding/removing an index **doesn't affect the contents of the database — only the performance of queries.**
>
> **Well-chosen indexes speed up read queries, but every index consumes additional disk space and slows down writes, sometimes substantially.**
That's why databases **don't index everything by default** — they require *you* to choose indexes using knowledge of the application's typical query patterns.
***
# 12.3 Databases and Streams (/docs/ddia/stream-processing/databases-streams)
> **Every write to a database is an event that can be captured, stored, and processed. The connection between databases and streams runs deeper than just the physical storage of logs on disk — IT IS QUITE FUNDAMENTAL.**
>
> * **A REPLICATION LOG is a stream of database write events produced by the leader.** Followers apply that stream and end up with an accurate copy.
> * **STATE MACHINE REPLICATION (Ch 10): if every event represents a write, and every replica processes the same events in the same order, all replicas end in the same state.** *(Processing is assumed DETERMINISTIC.)* **IT'S JUST ANOTHER CASE OF EVENT STREAMS.**
#### 3.1 The dual-write problem [#31-the-dual-write-problem]
**The setup:** an OLTP database for user requests, a cache, a full-text index, a warehouse — **each with its own copy of the data, in its own representation optimized for its own purposes.** They must be kept in sync.
**Dual writes — application code explicitly writes to each system — has two serious problems:**
> **The root cause: "In Figure 12-4 THERE ISN'T A SINGLE LEADER. The database may have a leader and the search index may have a leader, BUT NEITHER FOLLOWS THE OTHER, so conflicts can occur"** — it's accidental multi-leader replication (Ch 6).
>
> **The fix: "The situation would be better if there REALLY WAS ONLY ONE LEADER — the database — and if we could MAKE THE SEARCH INDEX A FOLLOWER OF THE DATABASE."**
#### 3.2 Change Data Capture (CDC) [#32-change-data-capture-cdc]
> **The problem with most databases' replication logs is that they have long been considered AN INTERNAL IMPLEMENTATION DETAIL, NOT A PUBLIC API. For decades, many databases simply DID NOT HAVE A DOCUMENTED WAY of getting the log of changes.**
>
> **CDC is the process of OBSERVING ALL DATA CHANGES written to a database and EXTRACTING THEM IN A FORM IN WHICH THEY CAN BE REPLICATED to other systems.**
**Implementations:** **Debezium** (source connectors for **MySQL, PostgreSQL, Oracle, SQL Server, Db2, Cassandra**, and more — attaching to replication logs and surfacing changes in **a standard event schema**), **Kafka Connect**, **Maxwell** (parses the MySQL binlog), **GoldenGate** (Oracle), **pgcapture** (PostgreSQL).
> **Like message brokers, CDC is usually ASYNCHRONOUS: the source database DOES NOT WAIT for a change to be applied to consumers before committing. This has the operational advantage that ADDING A SLOW CONSUMER DOES NOT AFFECT THE SYSTEM OF RECORD TOO MUCH, but the downside that ALL THE ISSUES OF REPLICATION LAG APPLY.**
##### Initial snapshot [#initial-snapshot]
> **If you have the log of ALL changes ever made, you can reconstruct the entire state by replaying it. However, keeping all changes forever would require TOO MUCH DISK SPACE, and replaying would TAKE TOO LONG — so the log needs to be truncated.**
>
> **Building a new full-text index requires A FULL COPY of the database. Applying only recent changes would MISS ITEMS THAT WERE NOT RECENTLY UPDATED.**
>
> ### **The snapshot MUST CORRESPOND TO A KNOWN POSITION OR OFFSET IN THE CHANGE LOG, so you know where to start applying changes afterward.** *(Debezium uses Netflix's **DBLog watermarking algorithm** for INCREMENTAL SNAPSHOTS.)* [#the-snapshot-must-correspond-to-a-known-position-or-offset-in-the-change-log-so-you-know-where-to-start-applying-changes-afterward-debezium-uses-netflixs-dblog-watermarking-algorithm-for-incremental-snapshots]
##### Log compaction — the elegant alternative [#log-compaction--the-elegant-alternative]
> ### **The payoff: "whenever you want to rebuild a derived data system such as a search index, START A NEW CONSUMER FROM OFFSET 0 of the log-compacted topic and sequentially scan all messages. THE LOG IS GUARANTEED TO CONTAIN THE MOST RECENT VALUE FOR EVERY KEY. You can use it to obtain A FULL COPY OF THE DATABASE CONTENTS WITHOUT HAVING TO TAKE ANOTHER SNAPSHOT."** [#the-payoff-whenever-you-want-to-rebuild-a-derived-data-system-such-as-a-search-index-start-a-new-consumer-from-offset-0-of-the-log-compacted-topic-and-sequentially-scan-all-messages-the-log-is-guaranteed-to-contain-the-most-recent-value-for-every-key-you-can-use-it-to-obtain-a-full-copy-of-the-database-contents-without-having-to-take-another-snapshot]
>
> **This allows the message broker to be used FOR DURABLE STORAGE, NOT JUST FOR TRANSIENT MESSAGING.**
##### API support today [#api-support-today]
> **Most popular databases now expose change streams AS A FIRST-CLASS INTERFACE, rather than the RETROFITTED AND REVERSE-ENGINEERED CDC efforts of the past.** MySQL and PostgreSQL send changes through **the same replication log they use for their own replicas.**
**The quorum-database challenge, and Cassandra's answer:**
> **CDC support for QUORUM WRITES is challenging because THERE'S NO SINGLE SOURCE OF TRUTH TO SUBSCRIBE TO. Whether data is visible DEPENDS ON EACH READER'S CONSISTENCY PREFERENCES.**
>
> **Cassandra SIDESTEPS this by exposing RAW LOG SEGMENTS FOR EACH NODE rather than a single stream of mutations. Systems that wish to consume must READ THE RAW SEGMENTS FOR EACH NODE AND DECIDE HOW BEST TO MERGE THEM — MUCH AS A QUORUM READER DOES.**
#### 3.3 CDC vs event sourcing [#33-cdc-vs-event-sourcing]
| | **CDC** | **Event sourcing** |
| ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Application's view** | **Uses the database in a MUTABLE way, updating and deleting at will** | **Application logic is EXPLICITLY BUILT on immutable events written to an event log; updates/deletes are DISCOURAGED OR PROHIBITED** |
| **Level of abstraction** | **Extracted at a LOW LEVEL (parsing the replication log)** — which is what **ensures the order of writes matches the order they were actually written**, avoiding the dual-write race | **Events reflect things that happened AT THE APPLICATION LEVEL rather than low-level state changes** |
| **Adoption cost** | **Can be added to an existing database WITH MINIMAL CHANGES; the application might NOT EVEN KNOW CDC IS OCCURRING** | **A BIG CHANGE for an application not already doing it** |
| **Log compaction** | ✔ **A CDC update event contains THE ENTIRE NEW VERSION of the record, so the current value is entirely determined by the most recent event ⇒ previous events CAN be discarded** | ✗ **Events express the INTENT of a user action, not the mechanics of the state update. LATER EVENTS TYPICALLY DO NOT OVERRIDE PRIOR ONES, so YOU NEED THE FULL HISTORY. LOG COMPACTION IS NOT POSSIBLE IN THE SAME WAY** |
*(Event-sourced apps typically store **snapshots** of current state — **but this is ONLY A PERFORMANCE OPTIMIZATION; the intention is that the system can store all raw events forever and reprocess the full log whenever required.**)*
##### The schema-as-public-API problem — and the outbox pattern [#the-schema-as-public-api-problem--and-the-outbox-pattern]
> **In a microservices architecture, a database is typically accessed from ONLY ONE SERVICE, making it AN INTERNAL IMPLEMENTATION DETAIL that developers can change freely.**
>
> ### **HOWEVER, CDC SYSTEMS TYPICALLY USE THE UPSTREAM DATABASE'S SCHEMA WHEN REPLICATING, WHICH TURNS THESE SCHEMAS INTO PUBLIC APIs THAT MUST BE MANAGED like the public API of the service. REMOVING A COLUMN WILL BREAK DOWNSTREAM CONSUMERS.** [#however-cdc-systems-typically-use-the-upstream-databases-schema-when-replicating-which-turns-these-schemas-into-public-apis-that-must-be-managed-like-the-public-api-of-the-service-removing-a-column-will-break-downstream-consumers]
>
> **Such challenges always existed with data pipelines, but they typically impacted ONLY DATA WAREHOUSE ETL. SINCE CDC IS OFTEN A DATA STREAM, OTHER PRODUCTION SERVICES MIGHT BE CONSUMERS. BREAKING THEM CAN CAUSE A CUSTOMER-FACING OUTAGE.** *(Data contracts are used to prevent this.)*
**The outbox pattern:**
#### 3.4 State, Streams, and Immutability — the philosophical core [#34-state-streams-and-immutability--the-philosophical-core]
> **Whenever you have state that changes, THAT STATE IS THE RESULT OF THE EVENTS THAT MUTATED IT OVER TIME.** Your list of available seats is the result of reservations processed; the current balance is the result of credits and debits; the response-time graph is an aggregation of individual response times.
>
> **No matter how the state changes, THERE WAS ALWAYS A SEQUENCE OF EVENTS THAT CAUSED THOSE CHANGES. EVEN AS THINGS ARE DONE AND UNDONE, THE FACT REMAINS TRUE THAT THOSE EVENTS OCCURRED.**
>
> ### **MUTABLE STATE AND AN APPEND-ONLY LOG OF IMMUTABLE EVENTS DO NOT CONTRADICT EACH OTHER; THEY ARE TWO SIDES OF THE SAME COIN.** [#mutable-state-and-an-append-only-log-of-immutable-events-do-not-contradict-each-other-they-are-two-sides-of-the-same-coin]
> **Jim Gray and Andreas Reuter, 1992: "THERE IS NO FUNDAMENTAL NEED TO KEEP A DATABASE AT ALL; THE LOG CONTAINS ALL THE INFORMATION THERE IS. THE ONLY REASON FOR STORING THE DATABASE (i.e., the current end-of-the-log) IS PERFORMANCE OF RETRIEVAL OPERATIONS."**
**Advantages of immutable events:**
**① The accounting precedent — centuries old.**
> **When a transaction occurs, it is recorded in an APPEND-ONLY LEDGER. The accounts — profit and loss, the balance sheet — are DERIVED from the transactions by adding them up.**
>
> **IF A MISTAKE IS MADE, ACCOUNTANTS DON'T ERASE OR CHANGE THE INCORRECT TRANSACTION. Instead they ADD ANOTHER TRANSACTION THAT COMPENSATES for the mistake — e.g. refunding an incorrect charge. THE INCORRECT TRANSACTION REMAINS IN THE LEDGER FOREVER, because it might be important for auditing. If incorrect figures were already published, the NEXT ACCOUNTING PERIOD INCLUDES A CORRECTION. THIS PROCESS IS ENTIRELY NORMAL IN ACCOUNTING.**
**② Recovery from buggy code.** *"If you accidentally deploy buggy code that writes bad data, RECOVERY IS MUCH HARDER IF THE CODE IS ABLE TO DESTRUCTIVELY OVERWRITE DATA."* Also: **customer service can use an audit log to diagnose requests and complaints.**
**③ More information than the current state.**
> **On a shopping site, a customer may add an item to their cart and then remove it. Although the second event cancels the first FROM THE POINT OF VIEW OF ORDER FULFILLMENT, IT MAY BE USEFUL FOR ANALYTICS to know THE CUSTOMER WAS CONSIDERING A PARTICULAR ITEM BUT THEN DECIDED AGAINST IT. Perhaps they will buy it in the future, or found a substitute. THIS INFORMATION WOULD BE LOST IN A DATABASE THAT DELETES ITEMS.**
**④ Deriving several views from the same log.**
> **Having an explicit translation step makes it easier to EVOLVE YOUR APPLICATION. To introduce a new feature presenting existing data in a new way, USE THE EVENT LOG TO BUILD A SEPARATE READ-OPTIMIZED VIEW and RUN IT ALONGSIDE THE EXISTING SYSTEMS WITHOUT MODIFYING THEM. RUNNING OLD AND NEW SIDE BY SIDE IS OFTEN EASIER THAN PERFORMING A COMPLICATED SCHEMA MIGRATION. Once readers have switched, SHUT THE OLD ONE DOWN AND RECLAIM ITS RESOURCES.**
>
> ### **"The traditional approach to database and schema design is based on the FALLACY THAT DATA MUST BE WRITTEN IN THE SAME FORM AS IT WILL BE QUERIED. DEBATES ABOUT NORMALIZATION AND DENORMALIZATION BECOME LARGELY IRRELEVANT if you can translate data from a write-optimized event log to read-optimized application state. IT IS ENTIRELY REASONABLE TO DENORMALIZE in the read-optimized views, as the translation process gives you a mechanism for KEEPING IT CONSISTENT WITH THE EVENT LOG."** [#the-traditional-approach-to-database-and-schema-design-is-based-on-the-fallacy-that-data-must-be-written-in-the-same-form-as-it-will-be-queried-debates-about-normalization-and-denormalization-become-largely-irrelevant-if-you-can-translate-data-from-a-write-optimized-event-log-to-read-optimized-application-state-it-is-entirely-reasonable-to-denormalize-in-the-read-optimized-views-as-the-translation-process-gives-you-a-mechanism-for-keeping-it-consistent-with-the-event-log]
*(The Ch 2 home timeline is exactly this: **highly denormalized read-optimized state, kept in sync by the fan-out service.**)*
**Concurrency control — the two-sided effect:**
**Limitations of immutability — three of them:**
1. **Churn.** *"Some workloads MOSTLY ADD data and rarely update or delete; they are EASY to make immutable. Other workloads have a HIGH RATE OF UPDATES AND DELETES ON A COMPARATIVELY SMALL DATASET; in these cases THE IMMUTABLE HISTORY MAY GROW PROHIBITIVELY LARGE, FRAGMENTATION MAY BECOME AN ISSUE, and THE PERFORMANCE OF COMPACTION AND GARBAGE COLLECTION BECOMES CRUCIAL."*
2. **Legal deletion.** GDPR erasure, or containing an accidental leak. **"It's NOT SUFFICIENT to just append another event indicating the data should be considered deleted — YOU ACTUALLY WANT TO REWRITE HISTORY AND PRETEND THE DATA WAS NEVER WRITTEN."** *(Datomic calls this **excision**; Fossil calls it **shunning**.)*
3. **Truly deleting is surprisingly hard.** **"Copies can live in many places. Storage engines, filesystems, and SSDs often WRITE TO A NEW LOCATION RATHER THAN OVERWRITING IN PLACE, and BACKUPS ARE OFTEN DELIBERATELY IMMUTABLE to prevent accidental deletion."**
**Crypto-shredding, and its honest limits:**
> **Store data you may want to delete ENCRYPTED; when you want to get rid of it, FORGET THE ENCRYPTION KEY. The encrypted data is still there, but nobody can use it.**
>
> **In a sense, THIS ONLY MOVES THE PROBLEM AROUND; the actual data is still immutable, BUT YOUR KEY STORAGE IS MUTABLE. Moreover, YOU HAVE TO DECIDE UP FRONT which data is encrypted with the same key — an important decision, since YOU CAN LATER CRYPTO-SHRED EITHER ALL OR NONE OF THE DATA ENCRYPTED WITH A PARTICULAR KEY, BUT NOT SOME OF IT. Storing a separate key for every data item would get TOO UNWIELDY, as the key storage would get AS BIG AS THE PRIMARY DATA STORAGE.** *(Puncturable encryption allows selective revocation but is not yet widely used.)*
>
> ### **"Overall, deletion is more a matter of MAKING IT HARDER TO RETRIEVE THE DATA than actually MAKING IT IMPOSSIBLE. Nevertheless, you sometimes have to try."** [#overall-deletion-is-more-a-matter-of-making-it-harder-to-retrieve-the-data-than-actually-making-it-impossible-nevertheless-you-sometimes-have-to-try]
***
# 12.8 Decision cheat sheet (/docs/ddia/stream-processing/decision-cheat-sheet)
**Which broker style?**
**How do I keep N systems in sync?**
**Never dual-write.** Pick one system of record and make everything else a follower via **CDC** (existing mutable app) or **event sourcing** (new app, intent matters, auditability required). Use the **outbox pattern** if you don't want your internal schema to become a public contract.
**Do I need log compaction?**
Yes if consumers must be able to **bootstrap from the log** without a separate snapshot — which is the whole point of "rebuild a derived system from scratch." No if events are **intent-level** (event sourcing), because later events don't supersede earlier ones.
**Event time or processing time?**
**Event time**, essentially always, because it's the only choice that is **deterministic under reprocessing.** Use processing time only when the delay is negligibly short and you don't care about replay. Budget for **watermarks, allowed lateness, and a dropped-event metric.**
**Which window?**
Fixed reporting intervals → **tumbling**. Smoothed trends → **hopping**. "Within N minutes of each other" → **sliding** (costly: buffers events). Per-user activity bursts → **session**.
**Which fault-tolerance mechanism?**
| Situation | Use |
| ------------------------------------------------------------------ | ------------------------------------------------------------------------- |
| All effects stay inside the framework | **Checkpointing / microbatching** — free exactly-once |
| Writing to an external store that supports conditional writes | **Idempotence with the offset as a dedup key** — cheapest |
| Effects span the processor and one specific system that cooperates | **Internal atomic commit** (Kafka transactions, Dataflow) |
| Effects are truly external and irreversible (emails, payments) | **Idempotency keys at the external boundary** — no framework can help you |
**Local state or remote?**
**Local + periodic replication** by default (remote lookup per message is slow). Remote only if state is enormous or shared. And **know your restore time** — that's your recovery objective.
***
# 12.5 Fault Tolerance (/docs/ddia/stream-processing/fault-tolerance)
**Why batch's approach doesn't transfer:**
> **Batch fault tolerance works because INPUT FILES ARE IMMUTABLE, EACH TASK WRITES TO A SEPARATE FILE, AND OUTPUT IS MADE VISIBLE ONLY WHEN A TASK COMPLETES SUCCESSFULLY. "IT APPEARS AS THOUGH EVERY INPUT RECORD WAS PROCESSED EXACTLY ONCE. Although restarting tasks means records MAY BE PROCESSED MULTIPLE TIMES, THE VISIBLE EFFECT IN THE OUTPUT IS AS IF THEY HAD BEEN PROCESSED ONLY ONCE."**
>
> **This is EXACTLY-ONCE SEMANTICS — "although EFFECTIVELY-ONCE WOULD BE A MORE DESCRIPTIVE TERM."**
>
> ### **"WAITING UNTIL A TASK IS FINISHED BEFORE MAKING ITS OUTPUT VISIBLE IS NOT AN OPTION, BECAUSE A STREAM IS INFINITE, SO YOU CAN NEVER FINISH PROCESSING IT."** [#waiting-until-a-task-is-finished-before-making-its-output-visible-is-not-an-option-because-a-stream-is-infinite-so-you-can-never-finish-processing-it]
#### 5.1 Microbatching and checkpointing [#51-microbatching-and-checkpointing]
| Approach | System | Mechanism |
| ----------------- | ------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Microbatching** | **Spark Streaming** | **Break the stream into small blocks and treat each like a miniature batch process.** Batch size typically **\~1 second — "a performance compromise: SMALLER batches incur greater SCHEDULING AND COORDINATION OVERHEAD, while LARGER batches mean A LONGER DELAY before results become visible."** ⚠️ **"Microbatching IMPLICITLY PROVIDES A TUMBLING WINDOW EQUAL TO THE BATCH SIZE (windowed by PROCESSING time, not event timestamps); jobs requiring larger windows must EXPLICITLY CARRY OVER STATE from one microbatch to the next"** |
| **Checkpointing** | **Apache Flink** | **Periodically generate ROLLING CHECKPOINTS of state and write them to durable storage. If an operator crashes, RESTART FROM THE MOST RECENT CHECKPOINT and DISCARD ANY OUTPUT generated between the checkpoint and the crash. Checkpoints are triggered by BARRIERS in the message stream — similar to microbatch boundaries, BUT WITHOUT FORCING A PARTICULAR WINDOW SIZE** |
> ### ⚠️ **THE CRITICAL LIMIT: "Within the confines of the framework, these provide the same exactly-once semantics as batch processing. HOWEVER, AS SOON AS OUTPUT LEAVES THE STREAM PROCESSOR — when it writes to a database, publishes to an external broker, or TRIGGERS THE SENDING OF EMAILS — THE FRAMEWORK IS NO LONGER ABLE TO DISCARD THE OUTPUT OF A FAILED MICROBATCH. Restarting causes the EXTERNAL SIDE EFFECT TO HAPPEN TWICE."** [#️-the-critical-limit-within-the-confines-of-the-framework-these-provide-the-same-exactly-once-semantics-as-batch-processing-however-as-soon-as-output-leaves-the-stream-processor--when-it-writes-to-a-database-publishes-to-an-external-broker-or-triggers-the-sending-of-emails--the-framework-is-no-longer-able-to-discard-the-output-of-a-failed-microbatch-restarting-causes-the-external-side-effect-to-happen-twice]
#### 5.2 Atomic commit revisited [#52-atomic-commit-revisited]
> **To give the appearance of exactly-once processing, ALL OUTPUTS AND SIDE EFFECTS MUST PERSIST IF AND ONLY IF THE PROCESSING IS SUCCESSFUL. That includes:**
>
> * **messages sent to downstream operators or external messaging systems (including EMAIL OR PUSH NOTIFICATIONS)**
> * **database writes**
> * **changes to operator state**
> * **acknowledgments of input messages (INCLUDING MOVING THE CONSUMER OFFSET FORWARD)**
>
> **These must all happen ATOMICALLY.**
**Why this works where XA didn't (Ch 8 §5.4):**
> **"In more RESTRICTED environments it is possible to implement such an atomic commit facility EFFICIENTLY. Used in Google Cloud Dataflow, VoltDB, and Apache Kafka. UNLIKE XA, THESE IMPLEMENTATIONS DO NOT ATTEMPT TO PROVIDE TRANSACTIONS ACROSS HETEROGENEOUS TECHNOLOGIES, but instead KEEP THE TRANSACTIONS INTERNAL by managing BOTH STATE CHANGES AND MESSAGING WITHIN THE FRAMEWORK. THE OVERHEAD OF THE TRANSACTION PROTOCOL CAN BE AMORTIZED BY PROCESSING SEVERAL INPUT MESSAGES WITHIN A SINGLE TRANSACTION."**
#### 5.3 Idempotence — the cheaper route [#53-idempotence--the-cheaper-route]
> **"An IDEMPOTENT operation is one you can perform MULTIPLE TIMES and it has THE SAME EFFECT as if you performed it ONCE. Deleting a key is idempotent; INCREMENTING A COUNTER IS NOT."**
>
> **"Even if an operation is NOT NATURALLY idempotent, IT CAN OFTEN BE MADE IDEMPOTENT WITH A BIT OF EXTRA METADATA. When consuming from Kafka, every message has a persistent, monotonically increasing OFFSET. WHEN WRITING TO AN EXTERNAL DATABASE, INCLUDE THE OFFSET OF THE MESSAGE THAT TRIGGERED THE LAST WRITE WITH THE VALUE. Thus you can tell whether an update HAS ALREADY BEEN APPLIED and avoid performing it again."**
**Four assumptions this relies on — all of them load-bearing:**
#### 5.4 Rebuilding state after a failure [#54-rebuilding-state-after-a-failure]
| Option | Detail |
| ------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Remote datastore, replicated** | **"Having to QUERY A REMOTE DATABASE FOR EACH INDIVIDUAL MESSAGE CAN BE SLOW"** |
| **Local state, replicated periodically** ✔ | On recovery, **"the new task can READ THE REPLICATED STATE AND RESUME PROCESSING WITHOUT DATA LOSS"** |
| **Flink** | **Periodic SNAPSHOTS of operator state written to durable storage (a DFS)** |
| **Kafka Streams** | **Replicates state changes by sending them to A DEDICATED KAFKA TOPIC WITH LOG COMPACTION — similar to CDC** |
| **VoltDB** | **Replicates state by REDUNDANTLY PROCESSING EACH INPUT MESSAGE ON SEVERAL NODES** (Ch 8's serial execution) |
| **Rebuild from the input** | **"Replicating the state may not even be NECESSARY. If the state is aggregations over a FAIRLY SHORT WINDOW, it may be fast enough to simply REPLAY THE INPUT EVENTS. If the state is a local replica of a database maintained by CDC, the database can be REBUILT FROM THE LOG-COMPACTED CHANGE STREAM"** |
> **"All this depends on the performance characteristics of the underlying infrastructure. In some systems, NETWORK DELAY MAY BE LOWER THAN DISK ACCESS LATENCY, and network bandwidth may be COMPARABLE TO DISK BANDWIDTH. NO SOLUTION IS UNIVERSALLY IDEAL, and the merits of LOCAL VERSUS REMOTE state MAY ALSO SHIFT AS STORAGE AND NETWORKING TECHNOLOGIES EVOLVE."**
***
# 12.12 Forward links (/docs/ddia/stream-processing/forward-links)
| Concept here | Where it's developed |
| ------------------------------------------------------------------------ | --------------------------------------------- |
| Composing batch and stream deliberately; unbundling the database | **Ch 13** — A Philosophy of Streaming Systems |
| End-to-end correctness and enforcing constraints without linearizability | **Ch 13** |
| Timeliness vs integrity | **Ch 13** |
| The log as consensus / total order broadcast | **Ch 10** §3 |
| Replication logs and their formats | **Ch 6** §1.5 |
| LSM compaction — the same algorithm as log compaction | **Ch 4** §2 |
| Two-phase commit and why XA fails | **Ch 8** §5 |
| Event sourcing and CQRS | **Ch 3** §3 |
| Clock untrustworthiness on user devices | **Ch 9** §3 |
| Batch fault tolerance for contrast | **Ch 11** §2.3 |
| GDPR, deletion, and data minimization | **Ch 1**, **Ch 14** |
# 12. Stream Processing (/docs/ddia/stream-processing)
> "A complex system that works is invariably found to have evolved from a simple system that works. The inverse proposition also appears to be true: A complex system designed from scratch never works and cannot be made to work." — John Gall
**The assumption Ch 11 quietly made, now removed:**
> **Batch processing assumed the input is BOUNDED — of a known and finite size — so the process knows when it has finished reading. (MapReduce's central sort MUST read its entire input before producing output, because THE VERY LAST INPUT RECORD COULD BE THE ONE WITH THE LOWEST KEY that needs to be the very first output record.)**
>
> **In reality, a lot of data is UNBOUNDED because it arrives gradually over time. Users produced data yesterday and today, and will produce more tomorrow. UNLESS YOU GO OUT OF BUSINESS, THIS PROCESS NEVER ENDS, SO THE DATASET IS NEVER "COMPLETE" IN ANY MEANINGFUL WAY.**
**A *stream* = data incrementally made available over time.** The concept appears in **Unix stdin/stdout, lazy lists, `FileInputStream`, TCP connections, audio and video over the internet.**
***
# 12.2 Log-Based Message Brokers (/docs/ddia/stream-processing/log-based-message-brokers)
**The mindset difference this fixes:**
#### 2.1 How it works [#21-how-it-works]
**Kafka, Amazon Kinesis Streams** work this way; **Google Cloud Pub/Sub is architecturally similar but exposes a JMS-style API rather than a log abstraction.**
> **Even though these brokers write ALL messages to disk, they achieve throughput of MILLIONS OF MESSAGES PER SECOND by sharding across machines, and fault tolerance by replicating.**
#### 2.2 Log vs traditional messaging — the honest comparison [#22-log-vs-traditional-messaging--the-honest-comparison]
**Fan-out is trivial:** *"several consumers can independently read the log without affecting one another; READING A MESSAGE DOES NOT DELETE IT."*
**Load balancing is coarse:** *"the broker assigns ENTIRE SHARDS to nodes in the consumer group instead of assigning individual messages."*
**Two downsides of coarse-grained balancing:**
1. **"The number of nodes sharing the work can be AT MOST THE NUMBER OF LOG SHARDS in that topic."** *(You could have two consumers split even/odd offsets, or use a thread pool — **but that complicates consumer offset management. In general, SINGLE-THREADED PROCESSING OF A SHARD IS PREFERABLE, AND PARALLELISM CAN BE INCREASED BY USING MORE SHARDS.**)*
2. **"If a SINGLE MESSAGE IS SLOW TO PROCESS, IT HOLDS UP THE PROCESSING OF SUBSEQUENT MESSAGES IN THAT SHARD"** — head-of-line blocking (Ch 2).
> ### **THE DECISION RULE: when messages may be EXPENSIVE to process, you want to parallelize PER MESSAGE, and ORDERING IS NOT SO IMPORTANT → JMS/AMQP style. When throughput is HIGH, each message is FAST to process, and ORDERING IS IMPORTANT → LOG-BASED.** [#the-decision-rule-when-messages-may-be-expensive-to-process-you-want-to-parallelize-per-message-and-ordering-is-not-so-important--jmsamqp-style-when-throughput-is-high-each-message-is-fast-to-process-and-ordering-is-important--log-based]
>
> *(The distinction is blurring — Kafka now supports JMS/AMQP-style consumer groups allowing multiple consumers to receive messages from the same partition.)*
**Routing for ordering:** *"Since sharded logs preserve ordering only WITHIN a shard, ALL MESSAGES THAT NEED TO BE CONSISTENTLY ORDERED MUST BE ROUTED TO THE SAME SHARD. E.g. events relating to one particular user appear in a fixed order — achieved by making THE USER ID THE PARTITION KEY."*
#### 2.3 Consumer offsets [#23-consumer-offsets]
> **Consuming a shard sequentially makes it easy to tell which messages have been processed: all messages with an offset LESS than the current offset are done. THE BROKER DOES NOT NEED TO TRACK ACKNOWLEDGMENTS FOR EVERY MESSAGE — ONLY TO PERIODICALLY RECORD THE CONSUMER OFFSETS. The reduced bookkeeping and the opportunities for BATCHING AND PIPELINING help increase throughput.**
>
> ⚠️ **"If a consumer fails, it will RESUME FROM THE LAST RECORDED OFFSET rather than the more recent last offset it saw. THIS CAN CAUSE THE CONSUMER TO SEE SOME MESSAGES TWICE."**
> **The offset is in fact very similar to the LOG SEQUENCE NUMBER in single-leader replication. EXACTLY THE SAME PRINCIPLE: THE MESSAGE BROKER BEHAVES LIKE A LEADER DATABASE AND THE CONSUMER LIKE A FOLLOWER.**
#### 2.4 Disk space, and the ring buffer [#24-disk-space-and-the-ring-buffer]
**Tiered and object storage:** **Kafka and Redpanda serve older messages from object storage as TIERED STORAGE. WarpStream, Confluent Freight, and Bufstream store ALL data in the object store.**
> **In addition to COST EFFICIENCY, this architecture MAKES DATA INTEGRATION EASIER: messages in object storage are stored AS ICEBERG TABLES, which enable BATCH AND DATA WAREHOUSE JOB EXECUTION DIRECTLY ON THE DATA WITHOUT HAVING TO COPY IT INTO ANOTHER SYSTEM.**
**When consumers can't keep up — the operational advantages:**
#### 2.5 Replaying old messages — the batch-like property [#25-replaying-old-messages--the-batch-like-property]
> **In a log-based broker, consuming messages is MORE LIKE READING FROM A FILE: a READ-ONLY operation that does not change the log. The only side effect is that THE CONSUMER OFFSET MOVES FORWARD — and THE OFFSET IS UNDER THE CONSUMER'S CONTROL.**
>
> **You can start a copy of a consumer WITH YESTERDAY'S OFFSET and write output to a DIFFERENT LOCATION to reprocess the last day's messages. YOU CAN REPEAT THIS ANY NUMBER OF TIMES, VARYING THE PROCESSING CODE.**
>
> ### **This makes log-based messaging MORE LIKE THE BATCH PROCESSES OF CH 11, where derived data is clearly separated from input data through a REPEATABLE TRANSFORMATION PROCESS. It allows more experimentation and easier recovery from errors and bugs — A GOOD TOOL FOR INTEGRATING DATAFLOWS WITHIN AN ORGANIZATION.** [#this-makes-log-based-messaging-more-like-the-batch-processes-of-ch-11-where-derived-data-is-clearly-separated-from-input-data-through-a-repeatable-transformation-process-it-allows-more-experimentation-and-easier-recovery-from-errors-and-bugs--a-good-tool-for-integrating-dataflows-within-an-organization]
***
# 12.4 Processing Streams (/docs/ddia/stream-processing/processing-streams)
**Three things you can do with a stream:**
1. **Write it to a database, cache, search index** — *"the streaming equivalent of Ch 11's batch use cases"*
2. **Push events to users** — email alerts, push notifications, a real-time dashboard. **"In this case, a human is the ultimate consumer"**
3. **Process input streams to produce output streams** — a pipeline of stages ← **this section**
> **A piece of code that processes streams is an OPERATOR or a JOB. It's closely related to Unix processes and MapReduce jobs, and the dataflow pattern is similar: a stream processor CONSUMES INPUT STREAMS READ-ONLY and WRITES ITS OUTPUT APPEND-ONLY. Sharding and parallelization patterns are also VERY SIMILAR.**
>
> ### **THE ONE CRUCIAL DIFFERENCE: A STREAM NEVER ENDS.** [#the-one-crucial-difference-a-stream-never-ends]
>
> **⇒ Sorting doesn't make sense on unbounded data, so SORT-MERGE JOINS CANNOT BE USED.**
> **⇒ Fault tolerance must change: "with a batch job running for a few minutes, a failed task can simply be RESTARTED FROM THE BEGINNING; with a stream job THAT HAS BEEN RUNNING FOR SEVERAL YEARS, restarting from the beginning after a crash MAY NOT BE A VIABLE OPTION."**
#### 4.1 Five uses of stream processing [#41-five-uses-of-stream-processing]
**The classic four monitoring applications:** **fraud detection** (unexpected credit card usage patterns → block the card) · **trading systems** (examine price changes, execute per rules) · **manufacturing** (monitor machine status, quickly identify malfunctions) · **military and intelligence** (track a potential aggressor, raise the alarm at signs of attack).
##### (a) Complex event processing (CEP) [#a-complex-event-processing-cep]
> **Developed in the 1990s. "Similarly to the way a REGULAR EXPRESSION lets you search for patterns of characters in a string, CEP lets you specify rules to search for CERTAIN PATTERNS OF EVENTS in a stream."**
>
> **CEP engines maintain A STATE MACHINE that performs the matching; on a match, they EMIT A COMPLEX EVENT (hence the name).**
>
> ### **THE RELATIONSHIP BETWEEN QUERIES AND DATA IS REVERSED. "Usually a database STORES DATA PERSISTENTLY AND TREATS QUERIES AS TRANSIENT — it searches for data matching the query and FORGETS ABOUT THE QUERY when finished. CEP ENGINES REVERSE THESE ROLES: QUERIES ARE STORED LONG-TERM; AS EACH EVENT ARRIVES, THE ENGINE CHECKS WHETHER IT HAS NOW SEEN A PATTERN MATCHING ANY OF ITS STANDING QUERIES."** [#the-relationship-between-queries-and-data-is-reversed-usually-a-database-stores-data-persistently-and-treats-queries-as-transient--it-searches-for-data-matching-the-query-and-forgets-about-the-query-when-finished-cep-engines-reverse-these-roles-queries-are-stored-long-term-as-each-event-arrives-the-engine-checks-whether-it-has-now-seen-a-pattern-matching-any-of-its-standing-queries]
*(Implementations: **Esper, Apama, TIBCO StreamBase**; Flink and Spark Streaming also have SQL support for declarative stream queries.)*
##### (b) Stream analytics [#b-stream-analytics]
> **"The boundary between CEP and stream analytics is blurry, but as a general rule, stream analytics is LESS focused on detecting specific event SEQUENCES and MORE oriented toward AGGREGATIONS AND STATISTICAL METRICS over large volumes of events."**
Examples: **rate of a certain event type** · **rolling average over a time period** · **comparing current statistics to previous intervals** (to detect trends or alert on metrics unusually high or low **compared to the same time last week**).
> **Averaging over a few minutes SMOOTHS OUT IRRELEVANT FLUCTUATIONS from one second to the next, while still giving a TIMELY picture of changes.**
**Probabilistic algorithms:** **Bloom filters** (set membership), **HyperLogLog** (cardinality estimation), **percentile estimation** (Ch 2).
> ⚠️ **"This use of approximation algorithms SOMETIMES LEADS PEOPLE TO BELIEVE THAT STREAM PROCESSING SYSTEMS ARE ALWAYS LOSSY AND INEXACT, BUT THAT IS WRONG. THERE IS NOTHING INHERENTLY APPROXIMATE ABOUT STREAM PROCESSING, AND USING PROBABILISTIC ALGORITHMS IS MERELY AN OPTIMIZATION."**
##### (c) Maintaining materialized views [#c-maintaining-materialized-views]
> **Unlike stream analytics, "considering only events within a certain TIME WINDOW is usually NOT SUFFICIENT. Building the materialized view potentially requires ALL EVENTS OVER AN ARBITRARY TIME PERIOD, apart from any obsolete events discarded by log compaction. IN EFFECT, YOU NEED A WINDOW THAT STRETCHES ALL THE WAY BACK TO THE BEGINNING OF TIME."**
>
> **"In principle any stream processor could be used, ALTHOUGH THE NEED TO MAINTAIN EVENTS FOREVER RUNS COUNTER TO THE ASSUMPTIONS OF SOME ANALYTICS-ORIENTED FRAMEWORKS that mostly operate on windows of limited duration."** *(Kafka Streams and ksqlDB support it, building on log compaction.)*
**Incremental view maintenance (IVM) — why databases aren't enough:**
##### (d) Search on streams [#d-search-on-streams]
> **"Conventional search engines FIRST INDEX THE DOCUMENTS AND THEN RUN QUERIES OVER THE INDEX. By contrast, searching a stream TURNS THE PROCESSING ON ITS HEAD: THE QUERIES ARE STORED, AND THE DOCUMENTS ARE EVALUATED AGAINST THEM, as in CEP."**
>
> **In the simplest case, test every document against every query — "although this can be SLOW if you have a LARGE NUMBER OF QUERIES. To optimize, IT IS POSSIBLE TO INDEX THE QUERIES AS WELL AS THE DOCUMENTS and thus narrow the set of queries that may match."** *(Elasticsearch's **percolator**.)*
Examples: **media monitoring services searching news feeds for mentions of companies/products/topics**; **real estate sites notifying users when a matching property appears.**
##### (e) Why actor frameworks are NOT stream processors [#e-why-actor-frameworks-are-not-stream-processors]
| | Actors | Stream processors |
| ------------- | ----------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------- |
| Purpose | **Primarily a mechanism for MANAGING CONCURRENCY and distributed execution of communicating modules** | **Primarily a DATA MANAGEMENT technique** |
| Communication | **Often EPHEMERAL and ONE-TO-ONE** | **Event logs are DURABLE and MULTI-SUBSCRIBER** |
| Topology | **Arbitrary, INCLUDING CYCLIC request/response patterns** | **Usually ACYCLIC PIPELINES where every stream is the output of one particular job and derived from a well-defined set of inputs** |
*(Some crossover: **Storm's distributed RPC** farms user queries out to nodes that also process event streams, interleaving queries with events. And **you can process streams with actor frameworks — but many don't guarantee message delivery on crashes, so THE PROCESSING IS NOT FAULT-TOLERANT unless you add retry logic.**)*
#### 4.2 Reasoning About Time — the hardest part [#42-reasoning-about-time--the-hardest-part]
**Why batch is easy and streaming is not:**
> **A batch process must look at THE TIMESTAMP EMBEDDED IN EACH EVENT. "THERE IS NO POINT IN LOOKING AT THE SYSTEM CLOCK OF THE MACHINE RUNNING THE PROCESS, because the time at which it is run HAS NOTHING TO DO with the time at which the events actually occurred. A batch process may read A YEAR'S WORTH of historical events WITHIN A FEW MINUTES."**
>
> **Moreover, using event timestamps makes the processing DETERMINISTIC: RUNNING THE SAME PROCESS AGAIN ON THE SAME INPUT YIELDS THE SAME RESULT.**
>
> **"Many stream processing frameworks use THE LOCAL SYSTEM CLOCK (the PROCESSING TIME) to determine windowing. This is SIMPLE, and reasonable IF THE DELAY BETWEEN EVENT CREATION AND PROCESSING IS NEGLIGIBLY SHORT. HOWEVER, IT BREAKS DOWN WITH ANY SIGNIFICANT PROCESSING LAG."**
**Six causes of processing lag:** **queueing · network faults · a performance issue causing contention in the broker or processor · a restart of the stream consumer · reprocessing past events while recovering from a fault · or after fixing a bug.**
**Out-of-order arrival:** *"A user makes request 1 (handled by server A), then request 2 (handled by server B). B's event REACHES THE BROKER BEFORE A's. Stream processors see B then A, EVEN THOUGH THEY OCCURRED IN THE OPPOSITE ORDER."*
**The Star Wars analogy:**
> **Episode IV was released in 1977, V in 1980, VI in 1983, then I, II, III in 1999, 2002, 2005, and VII, VIII, IX in 2015, 2017, 2019. "If you watched them in the order they came out, THE ORDER IN WHICH YOU PROCESSED THEM IS INCONSISTENT WITH THE ORDER OF THEIR NARRATIVE. (THE EPISODE NUMBER IS LIKE THE EVENT TIMESTAMP, AND THE DATE YOU WATCHED IS THE PROCESSING TIME.) As humans we cope with such discontinuities, BUT STREAM PROCESSING ALGORITHMS NEED TO BE SPECIFICALLY WRITTEN TO ACCOMMODATE THEM."**
**The concrete damage:**
##### Straggler events [#straggler-events]
> **"You can never be sure whether you have received ALL the events for a particular window or SOME ARE STILL TO COME."** You've counted events in minute 37; time has moved on to 38 and 39. **When do you declare minute 37 finished?**
>
> **You can TIME OUT after seeing no new events for a while. "However, SOME EVENTS COULD BE BUFFERED ON ANOTHER MACHINE SOMEWHERE, DELAYED BY A NETWORK INTERRUPTION."**
**Two options:**
1. **IGNORE the stragglers** — *"they are probably a small percentage in normal circumstances. TRACK THE NUMBER OF DROPPED EVENTS AS A METRIC AND ALERT IF YOU START DROPPING A SIGNIFICANT AMOUNT OF DATA"*
2. **PUBLISH A CORRECTION** — an updated value for the window with stragglers included. **"You may also need to RETRACT the previous output"**
*(A third mechanism: a **special message meaning "from now on there will be no more messages with a timestamp earlier than t"** — a watermark. **"However, if several producers on different machines are generating events, each with its own minimum threshold, THE CONSUMERS NEED TO TRACK EACH PRODUCER INDIVIDUALLY. ADDING AND REMOVING PRODUCERS IS TRICKIER IN THIS CASE."**)*
##### Whose clock are you using? [#whose-clock-are-you-using]
> **A mobile app reports usage metrics. It may be used OFFLINE, buffering events locally and sending them WHEN A CONNECTION IS NEXT AVAILABLE — WHICH MAY BE HOURS OR EVEN DAYS LATER. "To any consumers of this stream, THE EVENTS WILL APPEAR AS EXTREMELY DELAYED STRAGGLERS."**
>
> **The timestamp SHOULD really be the time the user interaction occurred, per the DEVICE's local clock. HOWEVER, THE CLOCK ON A USER-CONTROLLED DEVICE OFTEN CANNOT BE TRUSTED — it may be accidentally or DELIBERATELY set wrong. The time the server received it is more likely ACCURATE (the server is under your control) but LESS MEANINGFUL in terms of describing the user interaction.**
**The three-timestamp trick:**
> **"This problem is NOT UNIQUE TO STREAM PROCESSING; BATCH PROCESSING SUFFERS FROM EXACTLY THE SAME ISSUES. It is just MORE NOTICEABLE in a streaming context, where we are more aware of the passage of time."**
##### The four window types [#the-four-window-types]
> **State cost varies enormously: "a COUNTING operation will have ONLY ONE COUNTER regardless of window size or event count. On the other hand, SLIDING WINDOWS OR STREAM JOINS REQUIRE THAT EVENTS BE BUFFERED UNTIL THE WINDOW FINISHES. Therefore, LARGE WINDOW SIZES OR HIGH-THROUGHPUT STREAMS CAN CAUSE STREAM PROCESSORS TO KEEP A LOT OF TEMPORARY STATE."**
#### 4.3 Stream Joins — three kinds [#43-stream-joins--three-kinds]
**Why it's harder than batch:** **"the fact that NEW EVENTS CAN APPEAR AT ANY TIME makes joins on streams more challenging."**
##### (a) Stream–stream join (window join) [#a-streamstream-join-window-join]
**The example: search click-through rate.**
##### (b) Stream–table join (stream enrichment) [#b-streamtable-join-stream-enrichment]
##### (c) Table–table join (materialized view maintenance) [#c-tabletable-join-materialized-view-maintenance]
**The social network timeline, as a join:**
```sql
SELECT follows.follower_id AS timeline_id,
array_agg(posts.* ORDER BY posts.timestamp DESC)
FROM posts
JOIN follows ON follows.followee_id = posts.sender_id
GROUP BY follows.follower_id
```
**The four events the stream processor must handle:**
* **User u sends a post → add it to the timeline of every user following u**
* **User deletes a post, or their entire account → remove it from all timelines**
* **u1 starts following u2 → add u2's recent posts to u1's timeline**
* **u1 unfollows u2 → remove u2's posts from u1's timeline**
> **"The join of the streams corresponds DIRECTLY to the join of the tables in this query. THE TIMELINES ARE EFFECTIVELY A CACHE OF THE RESULT OF THE QUERY, UPDATED EVERY TIME THE UNDERLYING TABLES CHANGE."**
>
> **The calculus aside, which is genuinely illuminating: "If you regard a stream as THE DERIVATIVE OF A TABLE, and regard a join as A PRODUCT of two tables u·v, something interesting happens: THE STREAM OF CHANGES TO THE MATERIALIZED JOIN FOLLOWS THE PRODUCT RULE (u·v)′ = u′v + uv′. ANY CHANGE OF POSTS IS JOINED WITH THE CURRENT FOLLOWERS, AND ANY CHANGE OF FOLLOWS IS JOINED WITH THE CURRENT POSTS."**
##### Time dependence of joins — the subtle killer [#time-dependence-of-joins--the-subtle-killer]
> **All three require the processor to MAINTAIN STATE derived from one input and QUERY THAT STATE when processing the other. THE ORDER OF THE EVENTS THAT MAINTAIN THE STATE IS IMPORTANT — it matters whether you first FOLLOW and then UNFOLLOW, or the other way round.**
>
> **In a sharded log, ordering within a single partition is preserved, BUT THERE IS TYPICALLY NO ORDERING GUARANTEE ACROSS DIFFERENT STREAMS OR SHARDS.**
>
> ### **"If state changes over time, and you join with a state, WHAT POINT IN TIME DO YOU USE FOR THE JOIN?"** [#if-state-changes-over-time-and-you-join-with-a-state-what-point-in-time-do-you-use-for-the-join]
**The tax rate example:** *"If you sell things, you need to apply the right tax rate, which depends on country/state, product type, and DATE OF SALE (since tax rates change). WHEN JOINING SALES TO A TABLE OF TAX RATES, YOU PROBABLY WANT THE TAX RATE AT THE TIME OF THE SALE — WHICH MAY DIFFER FROM THE CURRENT RATE IF YOU ARE REPROCESSING HISTORICAL DATA."*
> **"If the ordering across streams is undetermined, THE JOIN BECOMES NONDETERMINISTIC — YOU CANNOT RERUN THE SAME JOB ON THE SAME INPUT AND NECESSARILY GET THE SAME RESULT."**
**The warehouse solution — slowly changing dimensions (SCD):**
> **Use A UNIQUE IDENTIFIER FOR A PARTICULAR VERSION of the joined record — every time the tax rate changes it gets a new identifier, and the invoice includes the identifier for the rate at the time of sale.**
>
> **This makes the join DETERMINISTIC, "but it has the consequence that LOG COMPACTION IS NOT POSSIBLE, since ALL VERSIONS of the records need to be retained. Alternatively, you can DENORMALIZE the data and include the applicable tax rate DIRECTLY IN EVERY SALE EVENT."**
***
# 12.7 Production failure catalog for this chapter (/docs/ddia/stream-processing/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ---------------------------------------------------- | --------------------------------------------------------------------------------- |
| Metrics quietly wrong; nobody noticed | **UDP/at-most-once delivery**; dropped messages aren't visible |
| Messages processed out of order | **Load balancing + redelivery** (§1.4), or **cross-partition ordering assumed** |
| One bad message loops forever, blocking the queue | **Poison message with no DLQ** |
| A new consumer can't see historical data | **AMQP/JMS destructive consumption** |
| Consumer falls behind, then silently skips data | **Retention expiry** past the consumer offset |
| Search index and database permanently disagree | **Dual-write race** (§3.1) |
| Half a write succeeded | **Dual write partial failure** — the atomic commit problem |
| Source database's disk fills up | **Inactive CDC replication slot** pinning WAL |
| Dropping a column caused a customer-facing outage | **CDC turned the schema into a public API** |
| Cannot rebuild a derived system without a snapshot | **No log compaction** on the change topic |
| Event-sourced log can't be compacted | Events express **intent**, not final state |
| GDPR erasure impossible | Immutability + copies everywhere; **crypto-shredding not designed in** |
| A traffic "spike" that never happened | **Windowing by processing time** after a redeploy backlog |
| Window results wrong after a network blip | **Straggler events** arriving after the window closed |
| Mobile events arrive days late with wrong timestamps | **Untrusted device clock**; no three-timestamp correction |
| Job emits nothing; looks healthy | **Watermark stalled** by an idle partition |
| Checkpoints grow until the job can't progress | **Unbounded state** — no TTL, too-wide join window |
| Stream job OOMs at high throughput | **Sliding windows / joins buffering events** |
| Reprocessing produces different results | **Nondeterministic join** across streams (time dependence) |
| Emails sent twice after a crash | **Side effect outside the framework** — microbatching/checkpointing can't undo it |
| Counter double-incremented on retry | **Non-idempotent operation** + at-least-once delivery |
| Zombie processor writes stale state | **No fencing** on failover (Ch 9) |
| Rebalance takes 10 minutes | **Large local state restore** from the changelog |
***
# 12.10 Self-test (/docs/ddia/stream-processing/self-test)
Why can't MapReduce's sort start producing output before reading all input? What does that imply about unbounded data?
Define an event. What are the streaming counterparts of "file" and "filename"?
Why does polling get *relatively* more expensive the more frequently you do it?
State the two questions that differentiate messaging systems, and the three options for the first.
When is message loss acceptable, and what is the trap even then?
Give three direct-messaging approaches and the shared assumption that limits all of them.
Give four differences between a message broker and a database.
Distinguish load balancing from fan-out. How do Kafka consumer groups provide both?
Draw the sequence in which load balancing + redelivery reorders messages. How do you avoid it, and what do you give up?
Describe the poison-message loop. Why does strong ordering make it worse, and what fixes it?
Contrast the "transient messaging" mindset with the database mindset. What two capabilities does the log-based hybrid recover?
What is a partition offset, and what ordering guarantee does it give — and *not* give?
Give the two downsides of assigning whole shards to consumers. State the rule for choosing between log-based and JMS-style brokers.
Why does a log-based broker only need to record periodic offsets rather than per-message acks? What does that cost you on failure?
Do the 20 TB / 250 MB/s calculation. What does it tell you about how long you have to fix a slow consumer?
Give three operational advantages of a slow consumer in a log-based broker versus a traditional one.
Why is consuming a log "more like reading a file"? What does that enable that AMQP cannot?
Draw the dual-write race. Why won't you notice it? Name the second, independent problem with dual writes.
What does CDC fundamentally do to the topology of your systems?
Why is an initial snapshot needed, and what must it be tied to?
Explain log compaction. What determines the disk space of a compacted log, and what does it let you do that a snapshot otherwise would?
Give three differences between CDC and event sourcing. Which one makes log compaction impossible, and why?
Explain the CDC-schema-as-public-API problem and how the outbox pattern addresses it. What two costs does the outbox add?
State the integrate/differentiate relationship between state and event streams. Quote the Gray–Reuter line about databases.
Give the accounting analogy for handling mistakes in an immutable log.
Give an example of information present in an event log but absent from current state.
Why does "normalized vs denormalized" become "largely irrelevant" with an event log?
Give one way event sourcing *worsens* and one way it *simplifies* concurrency control.
Give three limitations of immutability. Explain crypto-shredding and its two structural limits.
What is the one crucial difference between batch and stream processing, and give two consequences.
Explain the query/data role reversal in CEP. Where else in the chapter does the same reversal appear?
Why is it wrong to say stream processing is inherently approximate?
Why do analytics-oriented frameworks fit materialized-view maintenance badly?
Give the two drawbacks of `REFRESH MATERIALIZED VIEW`, and explain what IVM does instead.
Contrast event time with processing time. Use the Star Wars analogy. Draw what a redeploy does to a processing-time rate metric.
What is a straggler? Give the two handling options and the problem with watermark-style "no more messages before t" signals.
Give the three-timestamp scheme for untrusted device clocks and the two assumptions it makes.
Define tumbling, hopping, sliding, and session windows. Which two have unbounded state cost, and why?
For the search/click join: why isn't embedding search details in the click event equivalent? What does the join emit when no click arrives?
Why is a stream–table join really a stream–stream join? What is the window on the table side?
Write the timeline query as a table–table join and explain the product rule (u·v)′ = u′v + uv′ in that context.
What is the time-dependence problem in joins? Give the tax-rate example, the SCD fix, and what the fix costs.
Why can't batch's fault-tolerance approach be used directly for streams?
Compare microbatching and checkpointing. What implicit window does microbatching impose?
State precisely where exactly-once semantics stop working, and list the four things that must commit atomically.
Why do stream-internal atomic commits work where XA failed?
Give the four assumptions that idempotence-based exactly-once relies on. Which one requires a log-based broker?
Give four ways to recover operator state after a failure, and one case where no replication is needed at all.
you run an e-commerce platform. Requirements: (a) a search index, a cache, and a warehouse must all reflect order changes within seconds; (b) fraud detection must flag a card used in 3 countries within 10 minutes; (c) a "customers who viewed this also viewed" view must be maintained continuously; (d) an order-confirmation email must be sent exactly once; (e) you must honour GDPR erasure. For each requirement: name the mechanism, the window type (if any), the ordering guarantee you depend on, the fault-tolerance strategy, and what breaks when a consumer is down for 12 hours. Identify the one requirement where the framework cannot give you exactly-once and say what you do instead.
# 12.6 Technology deep dives (/docs/ddia/stream-processing/technology-deep-dives)
***
#### 6.1 Apache Kafka [#61-apache-kafka]
**Problem it solves.** Be simultaneously a **durable log** (rereadable, replayable, retained) and a **low-latency notification system** — the hybrid §2 opens with.
**Why wasn't RabbitMQ enough?** Consumption is **destructive**; a new consumer can't read the past; per-message acking limits throughput; and load balancing + redelivery **inevitably reorders** (§1.4).
**Why wasn't a database enough?** Polling cost grows as the hit rate falls (§1); triggers are an afterthought.
**How it works internally.** Topics → **partitions**, each an append-only segmented log on disk. Producers pick a partition by **key hash** (which is how you get per-key ordering). Each partition is replicated; one replica is **leader**, and the **ISR (in-sync replicas)** set determines what counts as committed — `acks=all` + `min.insync.replicas=2` is the durability contract that actually matters. Consumers in a **group** are assigned whole partitions; progress is a committed **offset**. **Log compaction** keeps the latest value per key forever (§3.2). Reads use **zero-copy `sendfile`**, which is why disk-backed throughput is so high. Transactions + idempotent producer give the internal atomic commit of §5.2.
**Deployment.** 3+ brokers across AZs; RF=3, `min.insync.replicas=2`, `acks=all`; **`unclean.leader.election.enable=false`** (Ch 10 §3.5 — otherwise you trade correctness for availability silently); KRaft (Raft) instead of ZooKeeper in modern versions; tiered storage to object store for long retention.
**Monitoring.**
* **Consumer lag per partition** — *the* metric. §2.4: the buffer is large enough that a human can fix a slow consumer *before* it starts missing messages, **but only if you're watching**
* **Under-replicated partitions** and **ISR shrink/expand rate** — a shrinking ISR is a durability degradation happening quietly
* **Offline partitions** (no leader available)
* **Request latency p99 by request type**; **fetch-purgatory size**
* **Rebalance frequency** per consumer group — frequent rebalances mean processing time exceeds `max.poll.interval.ms`
* **Partition skew** — bytes/messages per partition; a hot key means one consumer does all the work
**Scaling.** **Partition count bounds consumer parallelism** (§2.2) and **increasing it changes key→partition mapping**, breaking per-key ordering for existing keys. Choose generously up front.
**Backup.** Mirroring to a second cluster (MirrorMaker 2 / Cluster Linking) and/or tiered storage. **A compacted topic used for event sourcing is a system of record and needs real backups** (Ch 3 §3.4).
**What actually breaks in production.**
* **Retention expiry before a stalled consumer catches up** — silent, permanent data loss with no error, just an offset jump. §2.4's ring buffer, unmonitored.
* **Unclean leader election** enabled → an out-of-sync replica becomes leader → **acknowledged writes vanish.**
* **Hot partition from a low-cardinality key** — 90% of traffic on one partition, one consumer saturated, the rest idle.
* **Rebalance storms** when a consumer's processing time exceeds the poll interval, so the group never stabilizes and throughput goes to zero.
* **Assuming global ordering.** There is none across partitions (§2.1). Every "events arrived out of order" incident traces back to this.
* **Increasing partitions on a keyed topic**, silently breaking ordering for keys that move.
* **Consumer commits offsets before processing** ("at-most-once" by accident) → messages silently dropped on crash.
***
#### 6.2 Debezium / CDC pipelines [#62-debezium--cdc-pipelines]
**Problem it solves.** Make one database the leader and every derived system a follower (§3.2), eliminating dual-write races.
**Why wasn't dual writing enough?** §3.1's race condition and partial-failure problem — and crucially, **you don't even notice it happened.**
**Why not query the database periodically?** You miss intermediate states, you miss deletes, and the polling cost/freshness trade-off is bad.
**How it works internally.** A connector reads the database's **replication log** (MySQL binlog, Postgres logical decoding slot / `pgoutput`, Oracle LogMiner) and emits row-level change events with `before`/`after` images plus source metadata, into Kafka. **Initial snapshot** is bound to a **specific log position** (§3.2), and Debezium uses **DBLog watermarking** for incremental, non-blocking snapshots.
**Deployment.** Kafka Connect cluster; **one connector per source database**; schema registry for event schemas (Ch 5); an **outbox table** if you don't want the internal schema to become a public API (§3.3).
**Monitoring.**
* **Replication slot lag / retained WAL on the source** — *the* Postgres CDC failure: **an inactive slot pins WAL forever and fills the primary's disk** (Ch 6 §6.1). This is a *source database outage* caused by your pipeline.
* **Connector status** and **snapshot progress**
* **Event lag** (source commit time → Kafka append time)
* **Schema-change events** — every DDL is a potential downstream break (§3.3)
* **Tombstone/delete event rate** — an unexpected spike means someone ran a mass delete
**Scaling.** Bounded by the source's log generation and the connector's single-threaded decode. Filter tables aggressively.
**What actually breaks.**
* **The replication slot filling the source disk** when the connector is down or slow. This takes down the *production database*, not just the pipeline.
* **A dropped or renamed column breaking downstream production consumers** — §3.3's "schema became a public API," which "can cause a customer-facing outage."
* **DDL that logical decoding can't represent**, silently skipping data.
* **Snapshot + stream boundary bugs** — duplicates or gaps at the handover if the snapshot isn't tied to an exact log position.
* **CDC on a quorum database** (§3.2): per-node raw log segments that must be merged, with no single source of truth.
* **Confusing CDC with event sourcing** — expecting business intent from what is actually a row diff.
***
#### 6.3 Apache Flink [#63-apache-flink]
**Problem it solves.** Stateful, event-time-correct stream processing with exactly-once state semantics and large windows/joins.
**Why wasn't microbatching enough?** §5.1: microbatching imposes a **tumbling window by processing time equal to the batch size**, and forces a latency/overhead trade-off. Flink's **barrier-triggered checkpoints** decouple fault tolerance from window size.
**Why wasn't a simple consumer loop enough?** Because windows, joins, and aggregations need **state that must survive failure** and **event-time semantics** that a naive loop doesn't provide.
**How it works internally.** A dataflow graph of operators; **Chandy–Lamport-style asynchronous barrier snapshotting** — barriers flow through the stream, each operator snapshots its state when barriers from all inputs align, and the aligned snapshot is written to durable storage. **Watermarks** carry event-time progress and trigger windows; **allowed lateness** handles stragglers (§4.2). State backends: heap or **RocksDB** (for state larger than memory), with **incremental checkpoints**. **Two-phase commit sinks** extend exactly-once past the framework boundary for sinks that support it (§5.2).
**Deployment.** JobManager + TaskManagers on K8s/YARN; **checkpoint storage on a DFS/object store**; **savepoints** for planned upgrades (a savepoint is how you change code without losing state).
**Monitoring.**
* **Checkpoint duration, size, and failure rate** — growing duration is the leading indicator of every Flink problem
* **Backpressure** per operator (Flink exposes this directly — it tells you *which* operator is the bottleneck)
* **Watermark lag** per source — how far behind event time you are, and therefore when windows will fire
* **State size** per operator; RocksDB compaction and memory
* **Records dropped as late** (§4.2's "track the number of dropped events as a metric and alert")
* Restart count and time-to-recover
**Scaling.** Parallelism per operator; **keyed state is partitioned by key**, so a hot key means a hot subtask. Rescaling requires a savepoint.
**What actually breaks.**
* **Checkpoints taking longer than the checkpoint interval**, so they overlap and eventually the job can't make progress. Usually caused by backpressure or by state that grew unbounded.
* **Unbounded state** — a keyed state entry per user with no TTL, a stream–stream join with too wide a window, or session windows on a key space that never stops growing.
* **Watermark stalls.** One idle source partition holds the watermark back and **no windows ever fire** — the job looks healthy and emits nothing.
* **Late data silently dropped** because allowed lateness is zero and nobody instrumented the drop counter.
* **Nondeterministic joins** across streams (§4.3) making reprocessing produce different results.
* **Losing state on redeploy** because someone restarted without a savepoint.
***
#### 6.4 Kafka Streams / ksqlDB (and IVM engines: Materialize, RisingWave) [#64-kafka-streams--ksqldb-and-ivm-engines-materialize-risingwave]
**Problem it solves.** Maintain **materialized views** — the "window that stretches back to the beginning of time" (§4.1c) — continuously and incrementally.
**Why wasn't `REFRESH MATERIALIZED VIEW` enough?** §4.1: **poor efficiency** (all data reprocessed) and **poor freshness** (stale until the next scheduled run).
**Why weren't triggers enough?** They work only when **"the data is easily partitioned and the computation is naturally incremental"** — and **"many SQL queries can't be easily or efficiently converted to incremental computation."**
**How it works internally.** Kafka Streams: a library, not a cluster; **KTable** = a compacted changelog interpreted as a table; **KStream** = an event stream; joins follow §4.3's three types directly. Local state in RocksDB, **backed by a compacted Kafka changelog topic** so it can be rebuilt anywhere (§5.4). **IVM engines** (Materialize, RisingWave, Feldera) go further: they compile SQL into **incremental dataflow operators** (differential dataflow lineage), buffering recent events in memory and periodically updating on-disk views, with **reads combining both** for a real-time answer.
**Monitoring.** View staleness / end-to-end lag; state-store size and restore time (**how long to rebuild from the changelog — that's your RTO**); changelog topic size; rebalance-induced state migration; for IVM engines, memory per materialized view and arrangement size.
**What actually breaks.**
* **State restore time** after a rebalance — a multi-GB RocksDB store rebuilt from a changelog topic can take many minutes, during which that partition is unavailable. Standby replicas exist for exactly this.
* **Unbounded KTable growth** where keys are never deleted.
* **A materialized view that's expensive to maintain incrementally** — some joins produce huge intermediate arrangements.
* **Read-your-writes violations** (§3.4) — the user writes, then reads the view before the update lands.
* **Assuming the view is transactionally consistent with the source** — it isn't; CDC is asynchronous.
***
# 12.11 Terminology introduced here (/docs/ddia/stream-processing/terminology-introduced-here)
# 12.1 Transmitting Event Streams (/docs/ddia/stream-processing/transmitting-event-streams)
**An EVENT is the streaming counterpart of a batch record:**
> **A small, SELF-CONTAINED, IMMUTABLE object containing the details of SOMETHING THAT HAPPENED AT A POINT IN TIME. It usually contains a TIMESTAMP indicating when it happened according to a TIME-OF-DAY clock.**
**Vocabulary mapping:**
| Batch | Stream |
| -------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| file written once, read by multiple jobs | **event generated once by a PRODUCER (publisher, sender), processed by multiple CONSUMERS (subscribers, recipients)** |
| filename identifies a set of related records | **TOPIC or STREAM groups related events** |
**Why not just poll a database?**
> **In principle a file or database is sufficient: the producer writes every event, and each consumer PERIODICALLY POLLS for events since it last ran. THIS IS ESSENTIALLY WHAT A BATCH PROCESS DOES.**
>
> **But when moving toward continual processing with low delays, POLLING BECOMES EXPENSIVE if the datastore isn't designed for it. THE MORE OFTEN YOU POLL, THE LOWER THE PERCENTAGE OF REQUESTS THAT RETURN NEW EVENTS, AND THUS THE HIGHER THE OVERHEADS.**
>
> **Databases have traditionally not supported notification well. Relational databases have TRIGGERS, but they are VERY LIMITED and have been SOMEWHAT OF AN AFTERTHOUGHT in database design.**
#### 1.1 The two questions that differentiate every messaging system [#11-the-two-questions-that-differentiate-every-messaging-system]
**Whether loss is acceptable is application-specific:**
> **With periodic sensor readings and metrics, an occasional missing data point is perhaps not important — an updated value arrives shortly. HOWEVER, BEWARE THAT IF A LARGE NUMBER OF MESSAGES ARE DROPPED, IT MAY NOT BE IMMEDIATELY APPARENT THAT THE METRICS ARE INCORRECT.**
>
> **If you are COUNTING events, reliable delivery matters more, since EVERY LOST MESSAGE MEANS INCORRECT COUNTERS.**
#### 1.2 Direct messaging — and why it's limited [#12-direct-messaging--and-why-its-limited]
| Approach | Where used |
| ------------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **UDP multicast** | **Widely used in FINANCE for stock market feeds, where low latency is important.** UDP is unreliable, **but application-level protocols can recover lost packets — the producer must REMEMBER PACKETS IT HAS SENT so it can retransmit on demand** |
| **Brokerless libraries (ZeroMQ, nanomsg)** | Publish/subscribe over TCP or IP multicast |
| **StatsD** (metrics agents) | **Unreliable UDP. "Counter metrics are correct ONLY IF ALL MESSAGES ARE RECEIVED; using UDP makes the metrics AT BEST APPROXIMATE"** |
| **Webhooks** | **A callback URL of one service registered with another, which requests that URL whenever an event occurs** |
> **These generally require THE APPLICATION CODE TO BE AWARE OF THE POSSIBILITY OF MESSAGE LOSS. The faults they tolerate are QUITE LIMITED — they generally assume PRODUCERS AND CONSUMERS ARE CONSTANTLY ONLINE.**
>
> **If a consumer is OFFLINE, it may miss messages sent while unreachable. Some protocols let the producer retry — BUT THIS BREAKS DOWN IF THE PRODUCER CRASHES, LOSING THE BUFFER OF MESSAGES IT WAS SUPPOSED TO RETRY.**
#### 1.3 Message brokers [#13-message-brokers]
> **A kind of DATABASE OPTIMIZED FOR HANDLING MESSAGE STREAMS. By CENTRALIZING the data, these systems more easily tolerate clients that come and go, and THE QUESTION OF DURABILITY IS MOVED TO THE BROKER.**
>
> **Faced with slow consumers, they generally allow UNBOUNDED QUEUEING (as opposed to dropping or backpressure).**
>
> **A consequence of queueing is that CONSUMERS ARE GENERALLY ASYNCHRONOUS: when a producer sends, it normally waits ONLY for the broker to confirm it has BUFFERED the message — NOT for consumers to process it.**
**Brokers vs databases — four differences:**
| | Database | Message broker |
| ----------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Retention** | **Keeps data until EXPLICITLY DELETED** | **Some AUTOMATICALLY DELETE a message once successfully delivered. NOT SUITABLE FOR LONG-TERM DATA STORAGE** |
| **Working set** | Large | **Assumes queues are SHORT. If it must buffer a lot because consumers are slow (spilling to disk), EACH MESSAGE TAKES LONGER TO PROCESS AND OVERALL THROUGHPUT MAY DEGRADE** |
| **Selection** | **Secondary indexes, a query language** | **Subscribing to a subset of topics matching a pattern** — *"both are ways for a client to select the portion of the data it wants, but databases offer MUCH MORE ADVANCED query functionality"* |
| **Change notification** | **Result is a POINT-IN-TIME SNAPSHOT; the client is NOT told when its result becomes outdated** unless it repeats the query or polls | **No arbitrary queries and no updates after send — BUT THEY NOTIFY CLIENTS WHEN DATA CHANGES** |
*(Standards: **JMS, AMQP.** Implementations: **RabbitMQ, ActiveMQ, HornetQ, Qpid, TIBCO EMS, IBM MQ, Azure Service Bus, Google Cloud Pub/Sub.** And: **"although it is possible to use databases as queues, TUNING THEM TO GET GOOD PERFORMANCE IS NOT STRAIGHTFORWARD."**)*
**Two consumer patterns — and their combination:**
#### 1.4 Acknowledgments, redelivery, and the reordering trap [#14-acknowledgments-redelivery-and-the-reordering-trap]
> **Brokers use ACKNOWLEDGMENTS: a client must explicitly tell the broker when it has FINISHED PROCESSING so the broker can remove the message. If the connection closes or times out without an ack, the broker ASSUMES THE MESSAGE WAS NOT PROCESSED and DELIVERS IT AGAIN to another consumer.**
>
> ⚠️ **"It could happen that the message ACTUALLY WAS FULLY PROCESSED, BUT THE ACKNOWLEDGMENT WAS LOST IN THE NETWORK. Handling this requires an ATOMIC COMMIT PROTOCOL — unless the operation was IDEMPOTENT or exactly-once semantics are not required."**
**Load balancing + redelivery ⇒ inevitable reordering:**
**The poison-message loop and dead letter queues:**
> **A common scenario: a producer IMPROPERLY SERIALIZES a message — e.g. leaving out a required key in JSON. The message causes a consumer to CRASH AND RESTART, so it won't acknowledge, so the broker RESENDS IT, which causes ANOTHER CONSUMER TO FAIL. THIS LOOP REPEATS ITSELF INDEFINITELY.**
>
> **If the broker guarantees STRONG ORDERING, NO FURTHER PROGRESS CAN BE MADE. Brokers that allow reordering can continue — but WILL WASTE RESOURCES on messages that will never be acknowledged.**
>
> **DEAD LETTER QUEUES (DLQs) fix this: rather than retrying forever, the message is MOVED TO A DIFFERENT QUEUE TO UNBLOCK CONSUMERS. MONITORING IS USUALLY SET UP ON DLQs — ANY MESSAGE IN THE QUEUE IS AN ERROR.** An operator can then **permanently drop it, manually modify and reproduce it, or fix consumer code to handle it.**
***
# 12.9 Worked examples (/docs/ddia/stream-processing/worked-examples)
**① Polling cost vs notification.** 1,000 consumers polling a database every 100 ms, of which on average 1% of polls return data. That's **10,000 queries/s, 9,900 of which are wasted.** Halve the latency by polling every 50 ms → **20,000 queries/s, 19,800 wasted.** The cost is linear in freshness and the yield falls proportionally — §1's exact argument for why push beats poll.
**② Log retention as a safety margin.** Cluster ingests 200 MB/s; retention 7 days.
Storage = 200e6 × 86,400 × 7 ≈ **121 TB** (× replication factor 3 = **363 TB**).
A consumer processing at 180 MB/s that falls behind gains ground at only 20 MB/s. If it was down for 6 hours, it accumulated 200e6 × 21,600 = **4.3 TB** of lag, and needs **4.3e12 / 20e6 ≈ 60 hours** to catch up. **You have 7 days of retention and need 2.5 days to recover — the margin is real but not comfortable.** This is why §2.4's "monitor how far behind the head" matters, and why you must know your *catch-up rate*, not just your lag.
**③ Partition count vs parallelism.** 12 partitions, 20 consumer instances. **8 consumers sit idle** — §2.2's hard ceiling. Now one key carries 40% of traffic: that partition's consumer does 40% of the work while the other 11 share 60%. **Adding consumers changes nothing.** The only fixes are more partitions *and* a better key.
**④ Window state size.** Stream–stream join, 1-hour window, 50,000 events/s, 200 bytes each.
State = 50,000 × 3,600 × 200 = **36 GB** per side, **72 GB total**, held continuously. Extend the window to 24 hours (because some users click days later, §4.3a) → **1.7 TB**. This is why §4.2 warns that "large window sizes or high-throughput streams can cause stream processors to keep a lot of temporary state," and why the window length is a capacity decision, not a business decision alone.
**⑤ Straggler trade-off.** 1-minute tumbling windows; 99.9% of events arrive within 10 s, 0.09% within 5 minutes, 0.01% later.
* Close after 10 s → **0.1% of events dropped**; at 50,000 events/s that's **50 events/s discarded, forever, silently** unless you count them.
* Close after 5 minutes → **0.01% dropped**, but every result is **5 minutes stale.**
* Close after 10 s **and publish corrections** → fresh *and* eventually accurate, at the cost of downstream consumers having to handle **retractions.**
**⑥ Exactly-once, decomposed.** Processing a payment event must: (a) update operator state, (b) write to a database, (c) advance the consumer offset, (d) send a confirmation email.
* (a)+(c) — the framework handles atomically via checkpointing.
* (b) — idempotent if you store the offset alongside the row and check it.
* (d) — **nothing the framework can do.** The email provider must accept an **idempotency key**, or you must record "email sent for event X" transactionally with (b).
**Conclusion: "exactly-once" is always exactly-once *within a boundary*; every external effect needs its own idempotency story.**
***
# 1.2 Cloud vs Self-Hosting (/docs/ddia/trade-offs-data-systems/cloud-vs-self-hosting)
#### 2.1 The build/buy spectrum [#21-the-buildbuy-spectrum]
Two separable decisions: **who builds the software** and **who deploys it**.
Rule of thumb: **core competency / competitive advantage → in-house. Non-core, routine, commonplace → vendor.** (Most companies don't fabricate their own CPUs.)
#### 2.2 Honest scorecard [#22-honest-scorecard]
**Cloud wins when:**
* You **don't already know** how to deploy and operate the system. Hiring/training specialists is expensive.
* **Load varies a lot.** If you provision for peak and sit idle most of the time, you're wasting money. Analytical systems are the archetype: a big query needs lots of parallel compute, then the resources sit idle until the next query. Predefined daily reports can be enqueued and smoothed; **interactive queries are the variable ones — the faster you want them, the burstier the load**.
* The vendor's operational expertise (gained across many customers) exceeds yours.
**Self-hosting wins when:**
* You already have the operational skill **and** load is predictable → often just cheaper to buy machines.
* You need to **tune the system for your particular workload**. A cloud service won't customize for you.
* You need **full control of hardware** — e.g. very latency-sensitive high-frequency trading.
**The downsides of cloud, all of which reduce to "no control":**
| Risk | What you can actually do about it |
| -------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Missing feature | Politely ask the vendor. You cannot implement it. |
| Service goes down | Wait. |
| You trigger a bug or perf problem | Hard to diagnose — no OS metrics, no server logs, no internals |
| Vendor shuts down / raises price / changes product | Forced migration; running an old version isn't an option → **vendor lock-in** (mitigated only if alternatives expose a compatible API, and most cloud services have no standard API) |
| Geopolitics | Sanctions can lock you out of a provider in another country |
| Security & compliance | You must trust the provider with your data |
#### 2.3 Cloud native architecture — the *technical* consequence [#23-cloud-native-architecture--the-technical-consequence]
"Cloud native" = designed from the ground up to build on cloud services, not just to run on cloud VMs. Demonstrated advantages: **better performance on the same hardware, faster recovery from failures, faster scaling to match load, larger datasets.**
| Category | Self-hosted | Cloud native |
| ------------------ | --------------------------- | --------------------------------------------------------- |
| Operational / OLTP | MySQL, PostgreSQL, MongoDB | AWS Aurora, Azure SQL DB Hyperscale, Google Cloud Spanner |
| Analytical / OLAP | Teradata, ClickHouse, Spark | Snowflake, Google BigQuery, Azure Synapse Analytics |
**Layering.** Self-hosted software assumes generic resources: CPUs, RAM, a filesystem, an IP network. Cloud native services instead **build higher-level services on top of lower-level cloud services**:
* **Object storage (S3, Azure Blob, Cloudflare R2)** — more limited API than a filesystem (basic reads/writes of large files), but it **hides the underlying physical machines**: automatically distributes data across many machines so you never run out of disk on one, and survives machine/disk failure with no data loss.
* **Snowflake is a data warehouse built on S3.** Other services in turn build on Snowflake.
General abstraction rule: **higher-level abstractions are more oriented to particular use cases.** If your needs match, use the high-level thing. If nothing fits, compose from lower-level parts.
#### 2.4 Separation of storage and compute — the defining cloud-native move [#24-separation-of-storage-and-compute--the-defining-cloud-native-move]
Traditional model: disk is durable; **RAID** keeps copies across several disks on the same machine, transparent to applications.
Cloud model breaks that:
* **Local instance disks are treated as an ephemeral cache, not long-term storage** — the disk becomes inaccessible if the instance fails, or if the instance is resized (moving it to a different physical machine).
* **Virtual disks (EBS, Azure managed disks, GCP persistent disks)** can detach from one instance and attach to another. But a virtual disk is **not a physical disk** — it's a service run by a separate set of machines *emulating* a block device (typically 4 KiB blocks). Two costs: (a) block-device emulation overhead that a purpose-built cloud system avoids, and (b) **every I/O is a network call, making the application very sensitive to network glitches**.
* So cloud native services **avoid virtual disks** and build on **dedicated storage services optimized per workload**. Object stores are designed for **large files (hundreds of KB → several GB)**. Individual DB rows are far smaller — so cloud databases **manage small values in a separate service and pack larger blocks (containing many values) into the object store**.
**Multitenancy.** Cloud native systems typically share hardware across customers rather than one machine per customer. Benefits: better hardware utilization, easier scaling, easier management. Cost: **careful engineering so one customer's activity doesn't affect another's performance or security** (the noisy-neighbour problem).
#### 2.5 Operations in the cloud era [#25-operations-in-the-cloud-era]
Roles: DBAs/sysadmins → DevOps → **SRE** (Google's implementation of the idea). The role of operations: **deliver services reliably to users** (configuring infra, deploying apps) and **keep production stable** (monitoring and diagnosing anything that affects reliability).
Traditional self-hosted ops work is machine-level: capacity planning (watch disk space, add disks before you run out), provisioning machines, moving services between machines, OS patching.
Cloud services present an API that hides individual machines — e.g. cloud storage replaces fixed-size disks with **metered billing** (store without planning capacity, get charged for space used), and stays available even when individual machines fail.
**DevOps/SRE emphasis:**
* Automation, repeatable processes over manual one-off jobs
* **Ephemeral** VMs and services rather than long-running servers
* Frequent application updates
* **Learning from incidents**
* **Preserving organizational knowledge** as people come and go
**The bifurcation:** ops teams at infrastructure companies specialize in running a reliable service for many customers; customers of those services spend as little time on infrastructure as possible.
But cloud customers still need operations, just different operations:
* Choosing the right service for a task; integrating services; migrating between services
* **Capacity planning becomes financial planning; performance optimization becomes cost optimization** — metered billing removes capacity planning but you must still know what resources you use and why, or you burn money
* **Quotas and resource limits** (e.g. max concurrent processes) must be known and planned for *before* you hit them
* Integration between services is a growing challenge and **there are no standards** — it's manual effort
* Cannot be outsourced at all: application/library **security**, interactions between **your own** services, **monitoring load**, and **root-causing** performance degradations and outages
> **The need for operations is as great as ever — the cloud changed its shape, not its existence.**
***
# 1.6 Cross-cutting production failure catalog for this chapter (/docs/ddia/trade-offs-data-systems/cross-cutting-production-failure)
| Failure | Root trade-off it comes from |
| --------------------------------------------------- | -------------------------------------------------------------------- |
| Analytics query tanks production latency | Ran OLAP on the OLTP system — reason #3 for warehouses |
| Dashboards show stale data, nobody notices | Derived data with no freshness monitoring |
| "The numbers don't match between two dashboards" | Two derived paths from one system of record, no single definition |
| Vendor raises prices 4× / sunsets the product | Cloud lock-in with no compatible alternative API |
| Cloud service is slow and you can't tell why | No access to OS metrics, server logs, or internals |
| Application is intermittently slow, disks look fine | Virtual block device — every I/O is a network call |
| One tenant's heavy job degrades everyone | Multitenancy without proper resource isolation |
| Retry causes duplicate side effects | Timeout gives no information about whether the request was received |
| A cluster is slower than one big machine | Distributed by default; data movement cost exceeded parallelism gain |
| Deploy breaks 6 downstream clients | Microservice API evolution without schema management |
| Cannot honour a GDPR erasure request | Immutable logs + untracked derived copies |
| Surprise $40k cloud bill | Capacity planning became financial planning and nobody owned it |
***
# 1.4 Data Systems, Law, and Society (/docs/ddia/trade-offs-data-systems/data-systems-law-society)
Architecture is shaped by human and legal needs, not just technical ones.
* **GDPR** (2018, EU) and **CCPA** (California) give people control and legal rights over personal data. The **EU AI Act** adds restrictions on how personal data can be used.
* Automated systems make consequential decisions: who gets a loan or insurance, who gets a job interview, who is suspected of a crime. Social media changed news consumption → political opinion → election outcomes.
* **Legal requirements are reshaping system design foundations.** GDPR's **right to be forgotten** (erasure on request) collides head-on with designs built on **immutable append-only logs**. How do you delete data in the middle of a file that is supposed to be immutable? How do you handle data already baked into **derived datasets**, like the training data of an ML model? These are unsolved engineering challenges.
* Regulations **deliberately don't mandate technologies** (tech changes too fast) — they state high-level principles subject to interpretation. So there's no checklist for "GDPR-compliant architecture."
**Cost-benefit of storing data must include:** liability and reputational damage from a leak, legal costs and fines from non-compliance, and the fact that **governments or police may compel you to hand data over**. When data could reveal criminalized behavior (homosexuality in several countries; seeking an abortion in several US states), **storing it creates real safety risks for users** — travel to a clinic is revealed by location data, or even by a log of IP addresses over time.
**Data minimization (*Datensparsamkeit*)**: decide some data is not worth storing and delete it. This runs directly counter to the "big data" philosophy of hoarding speculatively. It aligns with GDPR: personal data may be collected only for a **specified, explicit purpose**, cannot later be used for another purpose, and must **not be kept longer than necessary**.
Industry compliance analogues: **PCI** (payment card industry) with frequent independent audits; **SOC 2 Type 2** for software vendors, also third-party audited.
***
# 1.7 Decision cheat sheet (/docs/ddia/trade-offs-data-systems/decision-cheat-sheet)
**Do I need a separate analytical system?**
Yes if any of: analysts need to join across ≥2 operational systems; analytical queries would compete with user traffic; analysts need ad-hoc SQL; or dataset is heading past a few hundred GB with aggregate-heavy access.
**Warehouse, lake, or real-time OLAP?**
* Analysts writing SQL over conformed business data → **warehouse**
* Data scientists needing raw/unstructured data and custom code → **lake** (or lakehouse)
* Aggregates served to *customers* at sub-second latency → **real-time OLAP**
* Both point lookups and scans in one request path → consider **HTAP**, knowing it's two engines behind one door
**Cloud or self-host?**
Cloud if you lack the operational skill, or your load is bursty. Self-host if load is predictable, you have the skill, and you need workload-specific tuning or hardware control. Always check: what's the exit path if the vendor changes?
**Distribute or stay on one node?**
Stay single-node unless you hit one of the nine reasons in §3.1. Modern single machines plus DuckDB/SQLite/Postgres cover far more than people assume. Distribution costs you: failure semantics, latency, debuggability, and cross-store consistency.
**Microservices?**
Only when team-coordination cost is the actual bottleneck. It's a people solution. In a small company it's overhead.
**Should we store this data at all?**
Cost = storage bill **+** breach liability **+** compliance fines **+** the risk to users if it's subpoenaed or leaked. Apply data minimization: collect for a specified explicit purpose, don't repurpose, don't keep longer than necessary.
***
# 1.3 Distributed vs Single-Node Systems (/docs/ddia/trade-offs-data-systems/distributed-vs-single-node)
A **distributed system** = several machines communicating over a network. Each participating process is a **node**.
#### 3.1 Nine legitimate reasons to go distributed [#31-nine-legitimate-reasons-to-go-distributed]
| Reason | Explanation |
| ----------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Inherent distribution** | Two or more interacting users on their own devices — communication *must* cross a network |
| **Requests between cloud services** | Data stored in one service, processed in another → network transfer. Cloud native & microservices are therefore distributed |
| **Fault tolerance / HA** | Redundancy so that a machine, several machines, the network, or a whole datacenter can fail and another takes over |
| **Scalability** | Data volume or compute exceeds one machine |
| **Latency** | Servers in multiple regions so each user is served from geographically nearby |
| **Elasticity** | Scale up/down with demand and pay only for what you use, instead of provisioning for peak |
| **Specialized hardware** | Object store = many disks/few CPUs; analytics = lots of CPU+RAM/no disks; ML = GPUs |
| **Legal compliance** | **Data residency laws** requiring data about people in a jurisdiction to be stored/processed inside it (scope varies — sometimes only medical or financial data) |
| **Sustainability** | Run jobs where/when renewable electricity is plentiful and the grid isn't strained; cuts carbon and cost |
#### 3.2 The costs, stated bluntly [#32-the-costs-stated-bluntly]
* **Every network request may fail or time out.** When a request times out, **you don't know whether the service received it**, so retrying may not be safe. (Ch 9.)
* **Network calls are vastly slower than in-process function calls**, even in fast datacenter networks. With large data volumes it's often faster to **bring the computation to the data** than to move the data to a separate machine.
* **More nodes are not always faster** — a simple single-threaded program on one computer can significantly beat a cluster of 100+ CPU cores.
* **Troubleshooting is hard.** If the system is slow, where is the problem? This is the domain of **observability**: collecting data about execution and querying it so both high-level metrics and individual events can be analyzed. Tracing: **OpenTelemetry, Zipkin, Jaeger** — track which client called which server for which operation and how long it took.
* **Cross-service consistency becomes the application's problem** when each service owns its database. **Distributed transactions** (Ch 8) exist but are **rarely used with microservices**: they cut against service independence, and many databases don't support them.
**Therefore:** doing a task on a single machine is often much simpler and cheaper. CPUs, memory, and disks have grown larger, faster, more reliable. With single-node engines like **DuckDB, SQLite, KùzuDB**, many workloads now fit on one node.
> Heuristic: **don't rush into distribution.** Prove you need it.
#### 3.3 Microservices [#33-microservices]
Client/server over HTTP is the common distribution style; the same process is often both a server (handling requests) and a client (making outbound ones). SOA → refined into **microservices**: a service has **one well-defined purpose** (S3's is file storage), exposes a network API, and has **one team responsible for it**.
**Advantages:** independent updates → less cross-team coordination; per-service hardware allocation; implementation hidden behind an API so owners can change internals freely. Each service typically **owns its own database** — sharing a database would make the entire schema part of the service's API (undeployable, unchangeable) and let one service's queries hurt another's performance.
**Costs:**
* **Testing** requires running all dependencies
* Each service needs infra for **releases, resource scaling, log collection, health monitoring, on-call alerting** — which is why **Kubernetes** became the standard foundation
* **API evolution is hard.** Clients expect certain fields; adding/removing them breaks clients, and the breakage is often discovered late, in staging or production. **OpenAPI and gRPC** help manage the client/server API relationship (Ch 5).
> **Microservices are primarily a technical solution to a people problem** — letting teams progress independently without coordinating. Valuable at a large company; in a small company with few teams it's likely unnecessary overhead, and the simplest possible implementation is preferable.
#### 3.4 Serverless / FaaS [#34-serverless--faas]
The cloud provider **automatically allocates and frees hardware based on incoming requests** — no explicit start/stop of instances. Just as cloud storage replaced capacity planning with metered billing, serverless **brings metered billing to code execution**: pay for the time your code runs.
Costs and caveats:
* **Time limits on function execution**, restricted runtime environments
* **Cold starts** on first invocation
* "Serverless" is a misleading name — each execution still runs on a server, just possibly a different one each time
* The term has been stretched: BigQuery and various Kafka offerings use "serverless" to mean **autoscaling + billing by usage rather than by machine instance**
#### 3.5 Cloud computing vs supercomputing (HPC) [#35-cloud-computing-vs-supercomputing-hpc]
Useful because large-scale analytical systems sometimes look like HPC.
| Dimension | Supercomputing / HPC | Cloud computing |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| Workload | Computationally intensive science: weather forecasting, climate modeling, molecular dynamics, optimization, PDEs | Online services, business data systems serving user requests with high availability |
| Fault handling | Large batch jobs **checkpoint to disk**; on node failure, **stop the whole cluster**, repair, restart from last checkpoint | Stopping the cluster is unacceptable — must serve users continuously with minimal interruption |
| Communication | **Shared memory and RDMA** — high bandwidth, low latency, **assumes high trust** among users | Network and machines shared by **mutually untrusting** organizations → VMs for resource isolation, encryption, authentication |
| Network topology | Specialized: multidimensional meshes and toruses, tuned to known communication patterns | IP + Ethernet in **Clos topologies** for high **bisection bandwidth** |
| Geography | All nodes assumed close together | Nodes distributed across geographic regions |
***
# 1.11 Forward links (/docs/ddia/trade-offs-data-systems/forward-links)
| Concept here | Where it's developed |
| ----------------------------------------------- | ------------------------------------------- |
| Why OLTP and OLAP use different storage layouts | **Ch 4** — Storage and Retrieval |
| Star/snowflake schemas | **Ch 3** — Data Models |
| Avro / Parquet, API evolution, OpenAPI/gRPC | **Ch 5** — Encoding and Evolution |
| Redundancy for fault tolerance | **Ch 6** — Replication |
| Distributed transactions across services | **Ch 8** — Transactions |
| Why network calls may fail ambiguously | **Ch 9** — Trouble with Distributed Systems |
| Data pipelines for integration | **Ch 11** — Batch Processing |
| Streams, CDC, second-level analytics | **Ch 12** — Stream Processing |
| Composing derived systems deliberately | **Ch 13** — Philosophy of Streaming Systems |
| Ethics, bias, privacy in depth | **Ch 14** — Doing the Right Thing |
# 1. Trade-Offs in Data Systems Architecture (/docs/ddia/trade-offs-data-systems)
> "There are no solutions; there are only trade-offs." — Thomas Sowell
**Thesis of the chapter:** every architectural choice in a data system is a trade-off along four axes. This chapter names the axes and gives you the vocabulary for the rest of the book.
The four axes:
1. **Operational (OLTP) vs Analytical (OLAP)** — what the data is *for*
2. **Cloud vs Self-hosted** — who builds it and who runs it
3. **Distributed vs Single-node** — how many machines, and why
4. **Business needs vs User rights** — law, ethics, and data minimization
***
# 1.0 The mental model for the whole book (/docs/ddia/trade-offs-data-systems/mental-model-whole-book)
**Definition worth memorizing:** an application is **data-intensive** if *data management* is the primary engineering challenge — storing large volumes, managing change, ensuring consistency under failure and concurrency, staying available. Contrast with **compute-intensive**, where the challenge is parallelizing a single big computation.
**Standard building blocks** every non-trivial app assembles:
| Block | Job |
| ----------------- | --------------------------------------------- |
| Databases | store data so it can be found again later |
| Caches | remember the result of an expensive operation |
| Search indexes | search by keyword / filter in arbitrary ways |
| Stream processing | react to events as they occur |
| Batch processing | periodically crunch accumulated data |
The hard part is never one block. It's **choosing between blocks with different characteristics, and gluing blocks together** when no single tool does the job.
***
# 1.1 Operational vs Analytical Systems (/docs/ddia/trade-offs-data-systems/operational-vs-analytical-systems)
#### 1.1 The people, because the split is a people-split first [#11-the-people-because-the-split-is-a-people-split-first]
| Role | What they do | Which system |
| ---------------------- | ---------------------------------------------------------- | ------------ |
| Backend engineer | builds services that read & **modify** data | Operational |
| Business analyst | reports for management (BI) | Analytical |
| Data scientist | novel insight, ML/AI features | Analytical |
| **Data engineer** | integrates operational ↔ analytical, owns the data infra | The bridge |
| **Analytics engineer** | models/transforms data so analysts & scientists can use it | Analytical |
Key observation: analysts and scientists both **read data that users and backend services generated**, and they **do not modify it** (they may create *derived* datasets). That read-only, derived nature is exactly why the systems can be split.
#### 1.2 OLTP vs OLAP — the access-pattern difference [#12-oltp-vs-olap--the-access-pattern-difference]
**OLTP (online transaction processing).** Historically a "transaction" was a *commercial* transaction (a sale, a payroll run). The name stuck even for social posts and game moves, because the *access pattern* stayed the same:
* **Point query** — look up a small number of records **by key**
* Insert / update / delete individual records driven by user input
* Interactive, so latency-sensitive
**OLAP (online analytical processing).** An analytical query **scans a huge number of records and computes aggregates** (count, sum, avg) rather than returning individual records.
Example analytical questions from a supermarket chain:
* What was the total revenue of each store in January?
* How many more bananas than usual did we sell during the promotion?
* Which brand of baby food is most often bought together with brand X diapers?
**The comparison table (worth internalizing — Table 1-1):**
| Property | Operational (OLTP) | Analytical (OLAP) |
| ------------------ | ---------------------------------------- | ---------------------------------- |
| Main read pattern | Point queries by key | Aggregate over many records |
| Main write pattern | Create/update/delete individual records | Bulk import (ETL) or event stream |
| Human user | End user of web/mobile app | Internal analyst, decision support |
| Machine user | Checking whether an action is authorized | Detecting fraud/abuse patterns |
| Query shape | Fixed, predefined in application code | Arbitrary, ad-hoc exploration |
| Query volume | Many small queries | Few queries, each complex |
| Data represents | **Latest state** (current point in time) | **History of events** over time |
| Dataset size | GB → TB | TB → PB |
Two consequences that follow directly:
* OLTP users are **not** allowed to write raw SQL (permission leakage + one bad query tanks everyone's latency). OLAP users *are* given arbitrary SQL, or a BI tool (Tableau, Looker, Power BI) that generates it.
* "Latest state" vs "history of events" is the deepest difference. It drives the storage layout in Ch 4, and the whole immutability philosophy in Ch 12–13.
#### 1.3 The third category: real-time / product analytics [#13-the-third-category-real-time--product-analytics]
Analytical *workload* (aggregates over many rows), but embedded in a **user-facing** product, so it needs OLTP-grade latency.
* Systems: **Apache Pinot, Apache Druid, ClickHouse**
* Ingest in **real time**, optimized for **low-latency query response**
* Contrast traditional OLAP: ingest in **batches**, optimized for **high-throughput** query processing
This is the category most modern "analytics dashboard in the product" features land in.
#### 1.4 Data warehousing [#14-data-warehousing]
**The problem it solves.** In the late 1980s/early 1990s companies stopped running analytics on their OLTP databases. Why running analytics directly on OLTP systems fails:
1. **Data silos** — the data of interest is spread across many operational systems, so you can't join it in one query. A large enterprise has dozens-to-hundreds of OLTP systems (website, point-of-sale, warehouse inventory, vehicle routing, supplier management, HR…), each complex, each with its own team, each operating independently.
2. **Wrong schema** — layouts good for OLTP are bad for analytics (see star schemas, Ch 3).
3. **Performance interference** — analytical queries are expensive; running them on OLTP hurts real users.
4. **Network/compliance isolation** — OLTP systems often sit in a network analysts may not access.
**A data warehouse is a separate database holding a read-only copy of data from all the OLTP systems**, which analysts can query to their hearts' content without affecting OLTP.
**ETL vs ELT.** ETL = extract, transform, then load. **ELT** = swap the last two: load raw, transform *inside* the warehouse. ELT is now more common because warehouse compute got cheap and elastic.
**ETL from SaaS.** When the source is an external SaaS product (CRM, email marketing, credit card processing) you have no DB access — only the vendor's API. Specialist connector services do this: **Fivetran, Singer, Airbyte**. Bringing SaaS data into your own warehouse enables analyses the SaaS API can't do.
#### 1.5 HTAP — and why it does not kill the warehouse [#15-htap--and-why-it-does-not-kill-the-warehouse]
**HTAP (hybrid transactional/analytical processing)** aims to serve OLTP and analytics in one system, no ETL.
Critical realism from the book: **many HTAP systems internally consist of an OLTP system coupled with a separate analytical system, hidden behind a common interface.** The distinction doesn't disappear; it gets hidden. So you still need to understand it to reason about performance.
And structurally HTAP can't replace a warehouse:
* Good practice = **each operational service owns its own database** → potentially hundreds of operational DBs.
* An enterprise wants **one** warehouse so analysts can join across systems in a single query.
HTAP's real niche: **one application that must both scan many rows analytically and read/update individual records at low latency.** Canonical example: **fraud detection**.
Wider trend this is an instance of: *the greater the scale, the more specialized systems become.* General-purpose systems handle small volumes fine; "one size fits all" stops being true as you scale (Stonebraker & Çetintemel).
#### 1.6 Data warehouse → data lake [#16-data-warehouse--data-lake]
Warehouses use a **relational model queried via SQL**. Great for analysts. Bad for data scientists, who need:
* **Feature engineering** — turning rows/columns into a vector or matrix of numbers (features) to train an ML model, in a way that maximizes model performance. Needs custom code that's awkward in SQL.
* **NLP on text** (e.g. extracting sentiment or topics from product reviews), **computer vision on images** — extracting structured info from unstructured data.
Despite efforts to add ML operators to SQL, data scientists largely prefer **Pandas, scikit-learn, R, Spark**.
**A data lake is a centralized repository holding a copy of any data that might be useful for analysis, obtained from operational systems via ETL — but it just contains *files*, imposing no particular file format, data model, or schema.**
* Files may be database records encoded as **Avro or Parquet** (Ch 5), or text, images, video, sensor readings, sparse matrices, feature vectors, genome sequences — anything.
* Cheaper than relational storage, because it uses commoditized **object storage**.
**The sushi principle: "raw data is better."** ETL generalized into **data pipelines**, and the lake became an intermediate stop on the way to the warehouse. The lake holds the raw form, so **each consumer transforms it into the shape that suits them** rather than being forced through one team's schema decision.
#### 1.7 Beyond the data lake [#17-beyond-the-data-lake]
Three forces reshaping the analytical side:
1. **Governance / privacy / compliance** — GDPR, CCPA; the **DataOps Manifesto** captures the operational maturity push.
2. **Streams, not just files and tables** — with files you rerun the analysis daily; **stream processing lets analytics respond in seconds**. That matters for e.g. identifying and blocking fraudulent or abusive activity.
3. **Reverse ETL** — pushing analytical outputs *back* into operational systems. Example: an ML model trained in the analytical system, deployed to production to generate "people who bought X also bought Y." Tools: **TFX, Kubeflow, MLflow**.
#### 1.8 Systems of record vs derived data — the single most useful lens in the book [#18-systems-of-record-vs-derived-data--the-single-most-useful-lens-in-the-book]
**System of record (source of truth):** holds the authoritative/canonical version. New data is written here **first**. Each fact appears **exactly once** (normalized). If another system disagrees, the system of record is **by definition** correct.
**Derived data system:** the result of transforming data from another system. **If you lose it, you can re-create it from the source.** Examples: caches, denormalized values, indexes, materialized views, transformed representations, **ML models trained on a dataset**.
Derived data is *technically redundant* — it duplicates information — but it's essential for read performance, and you can derive several datasets from one source to view the data from different angles.
Crucially: **most databases, storage engines, and query languages are not inherently one or the other.** A database is just a tool. Whether it's a system of record or derived data depends on **how you use it**. Being explicit about which data derives from which brings clarity to otherwise confusing architectures.
The gap the book keeps returning to: **many databases assume your app will only ever use that one database**, and make it hard to propagate updates to other systems. Data pipelines (Ch 11) and CDC (Ch 12) are the answer.
***
# 1.10 Self-test (/docs/ddia/trade-offs-data-systems/self-test)
Define "data-intensive" and contrast it with "compute-intensive."
Name the five standard building blocks and the job each does.
Give four differences between OLTP and OLAP access patterns. Which one drives the storage layout in Ch 4?
Why can't analysts just query the OLTP databases? Give all four reasons, and say which one is organizational rather than technical.
What is the difference between ETL and ELT, and why did ELT become more common?
Why doesn't HTAP eliminate the need for a data warehouse? Give the structural reason, not the performance one.
What is a data lake, and what specific data-science needs motivated it? What is the sushi principle?
Define system of record and derived data. Why is the distinction "not a property of the tool"?
Give three concrete downsides of a cloud service that all reduce to "you have no control."
What does "separation of storage and compute" actually mean, and why do cloud-native systems avoid virtual disks?
Why are local instance disks treated as an ephemeral cache?
In the cloud era, what does capacity planning become, and what does performance optimization become?
List five of the nine legitimate reasons to build a distributed system. Which two cannot be satisfied by a single machine at any price?
Why is "the request timed out" fundamentally different from "the request failed"?
Give three reasons a single machine may beat a cluster.
What problem are microservices actually solving? When are they unnecessary overhead?
Give four ways HPC differs from cloud computing in its assumptions about failure and trust.
Why does GDPR's right to erasure conflict with append-only log designs? Name two places a deletion must also reach.
State the data minimization principle and the philosophy it opposes.
a 15-person startup has one Postgres database, 200 GB of data, and analysts who currently run heavy queries against a read replica at 2 a.m. Growth projections say 10× data in 18 months. What would you change now, what would you defer, and what specific signal would tell you it's time to introduce a warehouse, a lake, or a real-time OLAP system? Justify each against a trade-off from this chapter.
# 1.5 Technology deep dives (/docs/ddia/trade-offs-data-systems/technology-deep-dives)
For each: *what problem, why wasn't the alternative enough, how it works internally, deployment, monitoring, scaling, backup, what breaks in production.*
Book content is the concepts; the ops sections are the practitioner layer the book leaves to you.
***
#### 5.1 Object storage (Amazon S3 / Azure Blob / GCS / Cloudflare R2) [#51-object-storage-amazon-s3--azure-blob--gcs--cloudflare-r2]
**Problem it solves.** Store an effectively unbounded amount of large files durably, without any single machine's disk capacity or failure being your problem.
**Why wasn't a filesystem / RAID / NAS enough?**
* A filesystem is bounded by one machine's disks. RAID survives a *disk* failure, not a *machine* or *datacenter* failure.
* NAS/SAN scale vertically and require capacity planning; you must decide the size in advance and pay for idle space.
* Object storage replaces capacity planning with **metered billing** and replaces machine-awareness with an API.
**How it works internally.**
* Flat **key → blob** namespace per bucket (the "directories" are a UI fiction over `/`-containing keys).
* Objects are split into chunks, chunks are replicated (or erasure-coded, e.g. Reed–Solomon k-of-n) across **independent failure domains** — different disks, machines, racks, and availability zones. Erasure coding gives similar durability to 3× replication at \~1.5× storage overhead.
* A separate **metadata/index service** maps keys → chunk locations. This index is the real scaling challenge, and it's why per-key operations are strongly consistent while *listing* is often weaker.
* **Immutable objects**: a PUT to an existing key writes a new version and flips a pointer. There is no in-place partial write — this is the single most important behavioral fact for system design on top of S3.
* Consistency: S3 has offered **strong read-after-write consistency for PUTs and DELETEs since Dec 2020**; before that it was eventually consistent, and enormous amounts of old design advice assumes the old model.
**Deployment.** No deployment — it is the substrate. Self-hosted equivalents: **MinIO, Ceph RADOS Gateway, SeaweedFS**. What you *do* deploy is policy: bucket per environment/tenant, IAM/bucket policies, encryption (SSE-S3 vs SSE-KMS), **versioning on**, **lifecycle rules** (transition to infrequent-access/Glacier at N days, expire noncurrent versions at M days), and **Block Public Access** at the account level.
**Monitoring.**
* Request rates and **4xx/5xx** split — a rising 503 `SlowDown` rate means you're hitting per-prefix request limits
* **First-byte and total latency** percentiles (p50/p99), not averages
* **Bytes stored by storage class** and by prefix (this is your bill)
* Replication lag if using cross-region replication
* **Access logs / CloudTrail data events** for who read what — required for most compliance stories
**Scaling.** Effectively unlimited capacity. The real limit is **request rate per prefix** (S3: \~3,500 PUT/COPY/POST/DELETE and \~5,500 GET/HEAD per second *per partitioned prefix*). You scale by **spreading keys across many prefixes** — historically by putting a hash at the *start* of the key. Read scaling beyond that: CloudFront/CDN in front, or S3 Transfer Acceleration for long-haul uploads.
**Backup.** Durability (11 nines) is **not** backup — it protects against disk failure, not against *you*. You need: **versioning** (undo overwrite/delete), **MFA Delete** or an Object Lock retention policy (undo malicious delete), **cross-region/cross-account replication** (undo account compromise or region loss). The account boundary matters: a backup in the same account that an attacker owns is not a backup.
**What actually breaks in production.**
* **Eventual-consistency assumptions in old code** — list-after-write still isn't strongly consistent in every provider/operation; jobs that "list the directory then process" silently miss new files.
* **Hot prefix throttling** — a batch job writing `2026-08-19/part-0000…` writes every key to one prefix and gets 503-throttled.
* **Cost explosions from small objects.** Object stores are for hundreds-of-KB-to-GB files. Millions of 2 KB objects means you pay per-request costs that dwarf storage, and any framework listing them takes forever. *This is exactly why cloud databases pack many small values into large blocks before writing to the object store.*
* **Lifecycle rules that delete data someone still depended on** — because the dependency was undocumented.
* **Silent public exposure** through a permissive bucket policy or a pre-signed URL with a year-long expiry.
* **Multipart upload leakage** — aborted multipart uploads keep consuming (billed) storage forever unless you add a lifecycle rule to abort incomplete uploads.
***
#### 5.2 Data warehouse (Snowflake / BigQuery / Redshift / Synapse) [#52-data-warehouse-snowflake--bigquery--redshift--synapse]
**Problem it solves.** One place where analysts can join data from every operational system and run arbitrary, expensive queries without touching production.
**Why wasn't running analytics on the OLTP DB enough?** The four reasons in §1.4: data silos, wrong schema, performance interference, network/compliance isolation.
**How it works internally.** (Full treatment in Ch 4; here's the shape.)
* **Columnar storage** — each column stored contiguously, so a query touching 2 of 200 columns reads 1% of the bytes.
* **Compression per column** (run-length, dictionary, bit-packing) — homogeneous data compresses hugely.
* **Vectorized execution** — operate on batches of column values in tight loops, exploiting SIMD, instead of row-at-a-time.
* **Disaggregated storage/compute** — Snowflake stores in S3 and spins up independent "virtual warehouses" (compute clusters) per workload, which is how one team's heavy query stops affecting another's. BigQuery separates storage (Colossus) from compute (Dremel) with the Jupiter network in between.
* **Micro-partitions + zone maps** — per-block min/max metadata lets the engine skip blocks entirely without reading them (data skipping / partition pruning).
**Deployment.** Warehouse itself is managed. What you deploy is the **modeling layer**: dbt (or equivalent) for transformations under version control, with CI running tests on models; separate dev/staging/prod databases; scheduled orchestration (Airflow/Dagster) for pipeline DAGs.
**Monitoring.**
* **Cost per query and per warehouse** — the #1 warehouse ops metric; bytes scanned (BigQuery) or credit-seconds (Snowflake)
* **Query queueing / concurrency saturation**
* **Freshness/SLA per table** — "how stale is this table?" is the metric analysts actually feel
* **Pipeline failure and retry counts**, and **row-count/schema drift** tests on each model
* **Spill to disk** (a query exceeding memory), which is the signal to resize the warehouse or fix the query
**Scaling.** Scale compute independently of data: resize a virtual warehouse (bigger machines) for a single heavy query, or add clusters (multi-cluster auto-scale) for concurrency. Isolate workloads on separate warehouses — ELT, BI, and data science shouldn't share compute.
**Backup.** Managed platforms give **time travel** (Snowflake: 1–90 days, then Fail-safe 7 days; BigQuery: 7-day time travel + snapshots). Time travel is not disaster recovery — for that you need **cross-region replication** and, for regulated data, exports to your own object storage. The **real** backup for a warehouse is the ability to **rebuild it from the lake/source systems**, because the warehouse is derived data. Test that rebuild.
**What actually breaks in production.**
* **A single analyst's exploratory query costs thousands of dollars** (an unfiltered scan of a petabyte table). Fix: partitioning + clustering + required partition filters + per-user byte quotas.
* **Silent pipeline failure** — a job fails, no one notices, dashboards show yesterday's numbers as if they were today's. Freshness alerting, not just job-failure alerting.
* **Schema drift upstream** — a source system renames a column; the ETL silently nulls it, and a KPI quietly drops to zero.
* **Timezone and late-arriving-data bugs** — the most common source of "the numbers don't match" between two dashboards.
* **Two teams computing "revenue" differently** — not a technical failure, but the one that destroys trust in the warehouse fastest. This is what the semantic/metrics layer exists to prevent.
***
#### 5.3 Data lake (S3/ADLS + Parquet + a table format) [#53-data-lake-s3adls--parquet--a-table-format]
**Problem it solves.** Store any data — including unstructured — cheaply, in raw form, so each consumer can transform it their own way (the sushi principle).
**Why wasn't a warehouse enough?** Feature engineering, NLP, and computer-vision workflows need custom code over arbitrary formats; data scientists prefer Pandas/scikit-learn/R/Spark to SQL; and relational storage is more expensive than object storage.
**How it works internally.** Files on an object store, typically **Parquet** (columnar, compressed, with per-row-group statistics) or **Avro** (row-oriented, great for streaming/CDC because it evolves schemas cleanly). A **table format** (Apache Iceberg, Delta Lake, Apache Hudi) layers over the files a metadata log giving: atomic commits, snapshot isolation, schema evolution, partition evolution, and **time travel** — this is what turned "a lake" into "a lakehouse" and made it queryable safely by multiple writers.
**Deployment.** Bucket layout by zone (`raw/`, `staging/`, `curated/`), partitioned by ingestion date; a catalog (Glue/Hive Metastore/Unity/Iceberg REST catalog) as the single source of table metadata; query engines (Athena/Trino/Spark/DuckDB) pointed at the catalog.
**Monitoring.** Small-file count per table (the canonical lake health metric), compaction job success, partition skew, catalog/metadata staleness, bytes scanned per query, and orphan-file accumulation.
**Scaling.** Storage scales for free. Query scaling = partition pruning + file sizing (aim for \~128 MB–1 GB Parquet files) + Z-ordering/clustering on high-selectivity predicates.
**Backup.** Object versioning + cross-region replication for the files; **the catalog is the fragile part** — back up the metastore/catalog database, because losing it turns a well-organized lake back into an undifferentiated pile of files.
**What actually breaks.**
* **The small-file problem** — a streaming job writing every minute produces millions of tiny Parquet files; query planning time exceeds query time. Requires scheduled compaction.
* **The lake becomes a swamp** — no catalog, no ownership, no schema, nobody knows which of the 40 `users_final_v3` datasets is real.
* **Concurrent writers corrupting a table** without a table format's atomic commits (partially-written partitions read as real data).
* **Schema-on-read mismatches** — one file has `user_id` as string, the next as int; the query engine fails, or worse, silently coerces.
* **GDPR erasure** — deleting one person's rows from immutable Parquet requires rewriting whole files. Table formats' row-level deletes exist precisely for this, and are slow and easy to forget to run against *all* derived copies.
***
#### 5.4 ETL/ELT connectors (Fivetran / Airbyte / Singer) [#54-etlelt-connectors-fivetran--airbyte--singer]
**Problem it solves.** Getting data out of SaaS products whose database you can never access — only a rate-limited, paginated, idiosyncratic API.
**Why not write it yourself?** You can — once. The cost is *maintenance*: every SaaS vendor changes their API, their pagination, and their rate limits independently, and each of your 40 connectors breaks on a different Tuesday.
**How it works internally.** Per-source connector with an **incremental cursor** (updated-at watermark, or a change/log endpoint), a normalization step into a target schema, and idempotent upserts into the destination keyed on primary key. Good connectors checkpoint the cursor so a failed sync resumes rather than restarting.
**Monitoring.** Sync success/failure per connector, **rows synced vs expected**, sync duration trend, API rate-limit consumption, and **replication lag per table** (the number analysts care about).
**Scaling.** Bounded by the source API, not by you. Scale = more frequent incremental syncs, never more parallel full-refreshes.
**What actually breaks.** Silent partial syncs after an API deprecation; hard deletes at the source that a watermark-based incremental sync can never see (you need soft deletes or periodic full refreshes); a full re-sync triggered by a schema change consuming the month's API quota in an hour; and PII arriving in the warehouse that nobody realized the SaaS was returning.
***
#### 5.5 Real-time OLAP (ClickHouse / Druid / Pinot) [#55-real-time-olap-clickhouse--druid--pinot]
**Problem it solves.** Analytical queries (aggregates over millions/billions of rows) at **interactive latency inside a user-facing product** — the dashboard your customers see, not the one your analysts see.
**Why wasn't a data warehouse enough?** Warehouses batch-ingest and optimize for throughput; query latency of seconds-to-minutes is fine for an analyst and fatal in a product. Why wasn't Postgres enough? It's row-oriented; scanning a billion rows to compute a sum reads every column of every row.
**Why wasn't a precomputed rollup enough?** It is, until users want a filter dimension you didn't precompute. These systems buy you arbitrary slicing.
**How it works internally (ClickHouse as the concrete case).**
* **MergeTree** family: data is written as immutable **parts**, sorted by the table's `ORDER BY` key; a background process **merges** parts into larger ones (an LSM-tree lineage — see Ch 4).
* Columnar storage with per-column codecs (`Delta`, `DoubleDelta`, `Gorilla`, `LZ4`, `ZSTD`); **sparse primary index** (one index entry per \~8,192-row granule) so the index fits in memory even for trillion-row tables.
* **Vectorized execution** over 65,536-row blocks; heavy SIMD use.
* **Materialized views that trigger on insert**, plus `AggregatingMergeTree`/`SummingMergeTree` to maintain rollups incrementally.
* Distribution: `Distributed` table fans a query out to shards; each shard replicated (ReplicatedMergeTree coordinates via Keeper/ZooKeeper).
**Deployment.** Shards × replicas; a coordination service (ClickHouse Keeper) with an odd number of nodes; ingestion via batched inserts (**never row-at-a-time**) or Kafka table engine. Separate the ingest path from the query path.
**Monitoring.** Parts-per-partition (the health metric — too many means merges are losing), merge queue depth and background pool saturation, insert batch size, `max_memory_usage` rejections, replication queue depth, mutation progress, and query p99 by query type.
**Scaling.** Vertical first — these engines exploit big machines very well. Then shard on a key with even distribution, and add replicas for read concurrency. Cluster reads and writes onto different replicas if query load is spiky.
**Backup.** `BACKUP`/`RESTORE` to object storage, or filesystem-level snapshots of parts plus schema DDL in version control. Replication is not backup — a bad `ALTER DELETE` replicates instantly.
**What actually breaks.**
* **"Too many parts" errors** — caused by frequent small inserts. The single most common ClickHouse production incident, and it's an application bug (not batching), not a database bug.
* **Choosing a bad `ORDER BY`** — irreversible without a full rewrite, and it determines every query's performance.
* **Memory-blown queries** — a `GROUP BY` on a high-cardinality column OOMs the server, taking down queries for everyone.
* **Distributed queries with no `GLOBAL IN`** — an innocuous subquery is re-executed per shard and melts the cluster.
* **Mutations** (`ALTER UPDATE/DELETE`) are asynchronous rewrites of whole parts; people treat them like OLTP updates and are surprised when a "delete one row" rewrites 200 GB.
***
#### 5.6 HTAP (SingleStore / TiDB / Spanner-style hybrids) [#56-htap-singlestore--tidb--spanner-style-hybrids]
**Problem it solves.** One application needing both low-latency single-record reads/writes and large scans — fraud detection is the archetype, where you must score a transaction *now* against aggregate history.
**Why wasn't OLTP + ETL + OLAP enough?** ETL latency. If the decision must be made in the request path, a warehouse refreshed every 15 minutes is useless.
**How it works internally.** Typically a **row store for recent/hot data + a column store for historical data**, with a background process converting row-format to column-format, and a query planner that can read both and union the results. This is the "internally two systems behind one interface" the book warns about.
**What actually breaks.** The seam between the two engines: queries that unexpectedly hit the row store scan and blow up; conversion lag making "analytics" silently exclude the last N minutes; and the operational reality that you now have one system with two very different tuning models and one team that understands neither fully.
***
#### 5.7 Kubernetes (as the microservices substrate) [#57-kubernetes-as-the-microservices-substrate]
**Problem it solves.** Every microservice independently needs: releases, resource allocation, log collection, health monitoring, and alerting. K8s provides that foundation once instead of N times.
**Why wasn't a VM per service with config management enough?** Bin-packing (one VM per service wastes capacity), rollout speed, and self-healing. Why wasn't a PaaS enough? Less control over networking, storage, and scheduling.
**How it works internally.** A declarative control loop: you write desired state to **etcd** via the API server; **controllers** continuously reconcile actual state toward desired; the **scheduler** binds Pods to Nodes by resource requests and constraints; **kubelet** on each node starts containers; **kube-proxy**/CNI provide service networking. Everything is a reconciliation loop — that's the whole design.
**Monitoring.** Pod restart counts and `CrashLoopBackOff`, **OOMKilled** counts, CPU throttling (a container hitting its CPU limit is throttled, not killed — and this is the silent latency killer), pending pods (= no node fits the request), node pressure conditions, etcd latency and DB size, and control-plane API latency.
**Backup.** **etcd snapshots** are the cluster backup; plus your manifests in Git (GitOps) so the cluster is reproducible. **PersistentVolumes need their own backup** (Velero, or CSI volume snapshots) — a common and painful gap.
**What actually breaks.**
* **Requests set too low** → nodes oversubscribed → everything throttles simultaneously under load.
* **Missing liveness/readiness distinction** — a liveness probe that fails under load restarts healthy-but-busy pods, converting a slowdown into an outage.
* **PodDisruptionBudgets missing** — a node drain during an upgrade takes all replicas of a service at once.
* **Stateful workloads** (databases) on K8s without understanding the storage layer — the classic way to lose data.
* **DNS**, always. CoreDNS saturation shows up as random, inexplicable timeouts across unrelated services.
***
#### 5.8 Distributed tracing (OpenTelemetry / Jaeger / Zipkin) [#58-distributed-tracing-opentelemetry--jaeger--zipkin]
**Problem it solves.** In a distributed system, "the system is slow" has no local answer. Tracing lets you ask *which* call, in *which* service, for *which* operation, took how long.
**Why weren't logs and metrics enough?** Metrics tell you *that* p99 rose; logs tell you what one service did. Neither reconstructs a **single request's path across service boundaries**, which is where the latency actually hides.
**How it works internally.** A **trace** = a tree of **spans**. Each span has a trace ID, span ID, parent span ID, timestamps, and attributes. **Context propagation** injects the trace ID into outgoing request headers (W3C `traceparent`), so downstream services attach their spans to the same trace. Spans are batched, exported to a collector, and **sampled** — because storing every span at scale is prohibitive. Two sampling models: **head-based** (decide at the root, cheap, may miss rare slow traces) and **tail-based** (buffer the whole trace, then keep the slow/erroring ones — much more useful, much more expensive, and requires all spans of a trace to reach the same collector).
**Monitoring the monitoring.** Span export failure/drop rate, collector queue saturation, and **cardinality of span attributes** — an attribute containing a user ID or a raw URL with IDs in it will destroy your tracing backend's cost model.
**Scaling.** Collector as a horizontally scaled deployment (with a load-balancing exporter in front when doing tail sampling), aggressive sampling of high-volume/low-value routes, and strict attribute allow-lists.
**What actually breaks.** Broken context propagation across an async boundary (a queue, a thread pool) — traces silently truncate at exactly the place you needed to see. Instrumenting only the HTTP layer, so all latency appears as "the database call" with no detail. And PII in span attributes, which turns your observability stack into a compliance liability.
***
# 1.8 Terminology introduced here (used for the rest of the book) (/docs/ddia/trade-offs-data-systems/terminology-introduced-here-used)
# 1.9 Worked examples (/docs/ddia/trade-offs-data-systems/worked-examples)
**① Is my data big enough to need a warehouse?** A company with 50 GB of operational data across 6 services.
Analysts need to join across 3 of them. **Data volume says "no warehouse needed" — but the *silo* argument says yes**, because you cannot join across three separately-owned OLTP databases in one query. **Reason #1 (silos) triggers long before reason #3 (query cost).** This is the most common mis-diagnosis: teams wait for the data to get big, when the real trigger was organizational.
**② The cost of an idle analytical cluster.** A 20-node cluster provisioned for peak query load, used 3 hours a day.
Utilization = 3/24 = **12.5%**. You pay for 24 hours of 20 nodes and use 3. In the cloud, the same workload with elastic compute costs \~1/8 as much. **This is precisely the "analytical systems have extremely variable load" argument for cloud** — and the counter-case is a *predictable* load, where owning hardware wins.
**③ Object store vs virtual disk, per I/O.** A query reading 10,000 blocks.
* **Local NVMe:** \~100 µs/block → **1 second**
* **Virtual block device (EBS):** every I/O is a *network call*, \~500 µs → **5 seconds**, and highly sensitive to network jitter
* **Object store, 10,000 separate GETs:** \~20 ms each → **200 seconds** ✗
* **Object store, 40 batched range reads of 250 blocks:** \~50 ms each → **2 seconds** ✔
**This is why cloud databases "pack many small values into large blocks before writing to the object store."** The access *pattern* matters more than the storage tier.
**④ Tail latency across services (a preview of Ch 2).** A request fans out to 30 microservices, each with p99 = 50 ms.
P(no service is slow) = 0.99³⁰ ≈ **74%** → **26% of user requests hit at least one p99-slow call.** Distribution didn't just add latency; it converted a 1-in-100 event into a 1-in-4 event.
**⑤ Data minimization as a cost calculation.** You're deciding whether to retain precise location history for 2 years.
* Storage cost: trivial, maybe $200/month
* **Breach liability:** location data revealing clinic visits, places of worship, or union meetings — for *every* user
* **Compliance exposure:** GDPR requires a specified explicit purpose and no longer than necessary
* **Compelled disclosure:** a subpoena reaches whatever exists
**The storage bill is the smallest term in the equation, and it's usually the only one anyone computes.**
***
# 8.6 THE ANOMALY TABLE — the single most useful table in the book (/docs/ddia/transactions/anomaly-table-single-most)
| Isolation level | Dirty reads | Read skew | Phantom reads | Lost updates | Write skew |
| ---------------------- | ----------- | ----------- | ------------- | ------------- | ----------- |
| **Read uncommitted** | ✗ Possible | ✗ Possible | ✗ Possible | ✗ Possible | ✗ Possible |
| **Read committed** | ✓ Prevented | ✗ Possible | ✗ Possible | ✗ Possible | ✗ Possible |
| **Snapshot isolation** | ✓ Prevented | ✓ Prevented | ✓ Prevented | **? Depends** | ✗ Possible |
| **Serializable** | ✓ Prevented | ✓ Prevented | ✓ Prevented | ✓ Prevented | ✓ Prevented |
*(**Dirty writes** are omitted because **almost all transaction implementations prevent them.**)*
**The six anomalies in one line each:**
| Anomaly | Definition |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Dirty read** | One client reads another client's writes **before they have been committed** |
| **Dirty write** | One client **overwrites data another client has written but not yet committed** |
| **Read skew** (nonrepeatable read) | A client **sees different parts of the database at different points in time** |
| **Phantom read** | A transaction reads objects matching a search condition; **another client makes a write that affects the results of that search.** SI prevents straightforward phantoms; **phantoms in the context of write skew require special treatment such as index-range locks** |
| **Lost update** | **Two clients concurrently perform a read-modify-write cycle; one overwrites the other's write without incorporating its changes** |
| **Write skew** | A transaction **reads something, decides based on what it saw, and writes the decision — but by the time the write is made, THE PREMISE OF THE DECISION IS NO LONGER TRUE.** **Only serializable isolation prevents this** |
***
# 8.9 Decision cheat sheet (/docs/ddia/transactions/decision-cheat-sheet)
**Which isolation level?**
**Which serializability implementation?**
| Choose | When |
| -------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Serial execution** | Active dataset fits in RAM; transactions are tiny; you can write stored procedures; writes fit one core or shard cleanly with **almost no cross-shard transactions** |
| **2PL** | You need serializability on a system that only offers it this way; contention is low; **and you can tolerate unstable tail latency** |
| **SSI** | **The default modern answer.** Read-heavy, moderate contention, spare capacity, short read/write transactions, and **an application that retries** |
**Do I need a distributed transaction?**
Prefer, in order: **(1) redesign so the transaction fits in one shard**; **(2) idempotency + a message-ID table** (§5.6) — this covers most "exactly-once" needs with only local transactions; **(3) database-internal distributed transactions** if your database has them; **(4) XA — essentially never for new systems.**
**How do I make retries safe?**
Retry only **transient** errors (40001, deadlock, timeout, failover). Use **exponential backoff with jitter**. Cap attempts. Make the operation **idempotent via a unique request ID** so a lost acknowledgment doesn't cause a double-apply. Keep **external side effects out of the transaction**, or make them idempotent too.
**Is my invariant enforceable?**
Single row → **check constraint**. Single column across rows → **uniqueness constraint**. Referential → **foreign key**. **Across multiple rows or an aggregate → nothing but serializable isolation (or a single-leader design that funnels the decision through one place).** And remember Ch 6: **multi-leader and leaderless replication cannot enforce these at all.**
***
# 8.5 Distributed Transactions (/docs/ddia/transactions/distributed-transactions)
**Single-node atomicity works because of one physical fact:**
**Distributed: you cannot just send commit to every node.**
> **Ensuring the nodes either ALL commit or ALL abort is the ATOMIC COMMITMENT PROBLEM.**
*(Note: **concurrency control in distributed transactions is broadly similar to single-node** — serial execution on sharded databases, 2PL works distributed, and there are **distributed serializability checkers for SSI.** **Achieving ATOMICITY is the new challenge.**)*
#### 5.1 Two-Phase Commit (2PC) [#51-two-phase-commit-2pc]
*(The marriage analogy: **the officiant asks each partner individually; after receiving both "I do"s, the couple is pronounced married — the transaction is committed — and the fact is broadcast. If either does not say yes, the ceremony is aborted.**)*
##### The system of promises — six steps [#the-system-of-promises--six-steps]
1. The application **requests a globally unique transaction ID** from the coordinator
2. It **begins a single-node transaction on each participant, attaching the global transaction ID.** All reads and writes happen in these single-node transactions. **If anything goes wrong at this stage, the coordinator or any participant can abort**
3. When ready to commit, the coordinator **sends `prepare` to all participants, tagged with the global ID.** If any request fails or times out, **the coordinator sends `abort` for that ID to all participants**
4. **On receiving `prepare`, a participant makes sure it CAN DEFINITELY COMMIT UNDER ALL CIRCUMSTANCES.** This includes **writing all transaction data to disk (A CRASH, A POWER FAILURE, OR RUNNING OUT OF DISK SPACE IS NOT AN ACCEPTABLE EXCUSE FOR REFUSING TO COMMIT LATER)** and checking for conflicts/constraint violations. **By replying yes, the node PROMISES to commit without error if requested — it SURRENDERS THE RIGHT TO ABORT, but without actually committing**
5. The coordinator **makes a definitive decision** (commit only if all voted yes) and **MUST WRITE THAT DECISION TO ITS TRANSACTION LOG ON DISK.** ← **the COMMIT POINT**
6. Once written, **send commit or abort to all participants. If this fails or times out, THE COORDINATOR MUST RETRY FOREVER. There is no more going back. If a participant crashed meanwhile, the transaction will be committed when it recovers — SINCE IT VOTED YES, IT CANNOT REFUSE TO COMMIT WHEN IT RECOVERS**
> **Two crucial POINTS OF NO RETURN: (1) when a participant votes yes, it promises it will definitely be able to commit later (though the coordinator may still choose to abort); (2) once the coordinator decides, that decision is IRREVOCABLE. Those promises ensure atomicity.**
>
> **Single-node atomic commit lumps these two events into one: writing the commit record to the transaction log.**
*(Continuing the analogy: **if you faint after saying "I do" and don't hear the pronouncement, that doesn't change the fact that the transaction was committed. When you recover, you can find out whether you are married by QUERYING THE OFFICIANT for the status of your global transaction ID — or wait for the officiant's next RETRY, since retries continued throughout your unconsciousness.**)*
##### Coordinator failure — the in-doubt problem [#coordinator-failure--the-in-doubt-problem]
**On recovery, the coordinator reads its transaction log; any transaction without a commit record is aborted.**
> **Thus THE COMMIT POINT OF 2PC COMES DOWN TO A REGULAR SINGLE-NODE ATOMIC COMMIT ON THE COORDINATOR.**
>
> **If the coordinator's DISK FAILS and its log is lost, the system HAS NO WAY TO AUTOMATICALLY RECOVER — only an administrator manually committing or aborting. And if only the most recent part of the log is lost, THE RECOVERING COORDINATOR MAY BELIEVE ALREADY-COMMITTED TRANSACTIONS WERE NOT COMMITTED AND TRY TO ABORT THEM, VIOLATING ATOMICITY.**
**Three-phase commit (3PC)** was proposed to make atomic commit **nonblocking. However, 3PC ASSUMES a network with BOUNDED DELAY and nodes with BOUNDED RESPONSE TIMES; in most practical systems with unbounded network delay and process pauses (Ch 9), 3PC CANNOT GUARANTEE ATOMICITY.**
> **A better solution in practice is to REPLACE THE SINGLE-NODE COORDINATOR WITH A FAULT-TOLERANT CONSENSUS PROTOCOL** (Ch 10).
#### 5.2 Two very different kinds of distributed transaction [#52-two-very-different-kinds-of-distributed-transaction]
> **Distributed transactions have a MIXED REPUTATION: seen as providing an important safety guarantee hard to achieve otherwise, but criticized for causing operational problems, killing performance, and PROMISING MORE THAN THEY CAN DELIVER. Many cloud services choose not to implement them.** *(Much of the cost is the ADDITIONAL `fsync` OPERATIONS required for crash recovery and the ADDITIONAL NETWORK ROUND TRIPS.)*
**But two things get conflated:**
| | **Database-internal** | **Heterogeneous** |
| ---------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- |
| **Participants** | **All nodes run THE SAME database software** | **Two or more TECHNOLOGIES** — databases from different vendors, or non-database systems like message brokers |
| **Examples** | YugabyteDB, TiDB, FoundationDB, Spanner, VoltDB, Cassandra, MySQL Cluster NDB, Kafka | XA transactions |
| **Freedom** | **Doesn't have to be compatible with anything else → can use ANY protocol and apply technology-specific optimizations. CAN OFTEN WORK QUITE WELL** | **A LOT MORE CHALLENGING** |
#### 5.3 Exactly-once message processing via 2PC [#53-exactly-once-message-processing-via-2pc]
> **A message from a queue can be acknowledged as processed IF AND ONLY IF the database transaction for processing it was successfully committed — by atomically committing the message acknowledgment and the database writes in a single transaction. With distributed transaction support this is possible even if the broker and the database are two unrelated technologies on different machines.**
>
> **If either fails, both are aborted, so the broker may safely REDELIVER later. The abort discards any side effects of the partially completed transaction. This is EXACTLY-ONCE SEMANTICS.**
> ⚠️ **Only possible if ALL affected systems can use the SAME atomic commit protocol. If a side effect is sending an email and the email server does not support 2PC, THE EMAIL COULD BE SENT TWO OR MORE TIMES.**
#### 5.4 XA transactions [#54-xa-transactions]
**X/Open XA (eXtended Architecture), introduced 1991.** Supported by **PostgreSQL, MySQL, Db2, SQL Server, Oracle**, and brokers **ActiveMQ, HornetQ, MSMQ, IBM MQ.**
> **XA is NOT A NETWORK PROTOCOL — it is merely a C API for interfacing with a transaction coordinator.** In Java: **JTA**, supported by **JDBC** drivers and **JMS** broker drivers.
**The architecture — and its fatal shape:**
##### Holding locks while in doubt — why in-doubt is so damaging [#holding-locks-while-in-doubt--why-in-doubt-is-so-damaging]
> **Database transactions acquire ROW-LEVEL EXCLUSIVE LOCKS on rows they modify. Under 2PL serializable, also SHARED LOCKS on rows they read. THE DATABASE CANNOT RELEASE THOSE LOCKS UNTIL THE TRANSACTION COMMITS OR ABORTS.**
>
> **Therefore a transaction must HOLD ITS LOCKS THROUGHOUT THE TIME IT IS IN DOUBT. If the coordinator crashed and takes 20 minutes to restart, those locks are held for 20 MINUTES. IF THE COORDINATOR'S LOG IS ENTIRELY LOST, THOSE LOCKS ARE HELD FOREVER — or until manually resolved.**
>
> **While held, no other transaction can modify those rows; depending on isolation level, others may even be blocked from READING them. THIS CAN CAUSE LARGE PARTS OF YOUR APPLICATION TO BECOME UNAVAILABLE.**
##### Recovering from coordinator failure [#recovering-from-coordinator-failure]
> **In practice, ORPHANED IN-DOUBT TRANSACTIONS DO OCCUR** — transactions whose outcome the coordinator cannot decide (log lost or corrupted by a software bug). **These cannot be resolved automatically, so THEY SIT FOREVER IN THE DATABASE, HOLDING LOCKS AND BLOCKING OTHER TRANSACTIONS.**
>
> **Even rebooting your database servers will not fix this, since a correct 2PC implementation MUST PRESERVE THE LOCKS OF AN IN-DOUBT TRANSACTION EVEN ACROSS RESTARTS** (otherwise it risks violating atomicity). **It's a sticky situation.**
**The only way out:** an administrator **examines the participants of each in-doubt transaction, determines whether any has already committed or aborted, and applies the same outcome to the others.** This requires **a lot of manual effort, most likely under high stress and time pressure during a serious production outage** (otherwise, why would the coordinator be in such a bad state?).
**Heuristic decisions** — the emergency escape hatch letting a participant **unilaterally decide** without the coordinator:
> **To be clear, "heuristic" here is a EUPHEMISM FOR PROBABLY BREAKING ATOMICITY, since the heuristic decision violates the system of promises in 2PC. Intended only for catastrophic situations, not regular use.**
##### The four fundamental problems with XA [#the-four-fundamental-problems-with-xa]
1. **A single-node coordinator is a SINGLE POINT OF FAILURE for the entire system**
2. **Making it part of the application server is problematic** — the coordinator's logs on local disk become **crucial durable system state, as important as the databases themselves**
3. **Even a replicated coordinator wouldn't fix it:** XA **provides NO WAY for coordinator and participants to communicate DIRECTLY — only via the application code and drivers. SO THE APPLICATION CODE WOULD BE THE SINGLE POINT OF FAILURE.** Solving this would require **totally redesigning how application code is run to make it replicated or restartable — perhaps similar to DURABLE EXECUTION (Ch 5). However, no tools seem to take this approach in practice**
4. **XA is a LOWEST COMMON DENOMINATOR**, because it must be compatible with a wide range of systems. **It cannot detect DEADLOCKS across different systems (that would require a standardized protocol for exchanging lock-wait information), and it DOES NOT WORK WITH SSI (that would require a protocol for identifying conflicts across systems)**
#### 5.5 Database-internal distributed transactions — why they're fine [#55-database-internal-distributed-transactions--why-theyre-fine]
**NewSQL databases use 2PC for cross-shard atomicity yet don't suffer XA's problems, because they don't need to interface with other technologies — they avoid the lowest-common-denominator trap.**
**The four fixes:**
1. **REPLICATE THE COORDINATOR, with automatic failover** if the primary crashes
2. **Let the coordinator and data shards COMMUNICATE DIRECTLY**, without intermediary application code
3. **REPLICATE THE PARTICIPATING SHARDS**, reducing the risk of aborting because of a fault in one shard
4. **COUPLE the atomic commitment protocol WITH a distributed CONCURRENCY CONTROL protocol** supporting deadlock detection and consistent reads across shards
**Consensus algorithms** replicate both coordinator and shards (Ch 10) — **tolerating faults by automatically failing over WITHOUT HUMAN INTERVENTION while continuing to guarantee strong consistency.** **Both snapshot isolation and SSI are possible across shards.**
#### 5.6 Exactly-once WITHOUT distributed transactions [#56-exactly-once-without-distributed-transactions]
> **You don't actually need distributed transactions to achieve exactly-once semantics.**
> **Achieving exactly-once processing requires ONLY TRANSACTIONS WITHIN THE DATABASE — atomicity across database and message broker is NOT NECESSARY. Recording the message ID makes the processing IDEMPOTENT, so it can be safely retried without duplicating side effects.** A similar approach is used in **Kafka Streams** (Ch 12).
>
> **That said, internal distributed transactions are still useful for the SCALABILITY of such patterns** — e.g. **message IDs on one shard and the main data on other shards, with atomicity across those shards.**
***
# 8.13 Forward links (/docs/ddia/transactions/forward-links)
| Concept here | Where it's developed |
| ---------------------------------------------------------- | ------------------------------------------- |
| Why timeouts and process pauses break 3PC's assumptions | **Ch 9** — Trouble with Distributed Systems |
| Determinism and its difficulty | **Ch 9** |
| Replacing the 2PC coordinator with consensus | **Ch 10** — Consistency and Consensus |
| Linearizability (the "consistency" of CAP) | **Ch 10** |
| State machine replication / shared logs | **Ch 10** |
| Exactly-once in stream processors (Kafka Streams) | **Ch 12** — Stream Processing |
| Idempotence at scale | **Ch 12** |
| Alternative approaches to constraints without transactions | **Ch 13** — Enforcing Constraints |
| Conflict detection in replicated systems | **Ch 6**, **Ch 13** |
# 8. Transactions (/docs/ddia/transactions)
> "Some authors have claimed that general two-phase commit is too expensive to support… We believe it is better to have application programmers deal with performance problems due to overuse of transactions as bottlenecks arise, rather than always coding around the lack of transactions."
> — James Corbett et al., *Spanner: Google's Globally-Distributed Database* (2012)
**The six things that go wrong, which transactions exist to hide:**
1. **The database software or hardware may fail at any time** — including in the middle of a write
2. **The application may crash at any time** — including halfway through a series of operations
3. **Network interruptions** can cut off the application from the database, or one node from another
4. **Several clients may write at the same time, overwriting one another's changes**
5. **A client may read data that doesn't make sense because it has only partially been updated**
6. **Race conditions between clients can cause surprising bugs**
> **A transaction groups several reads and writes into a logical unit. Conceptually all of them execute as ONE operation: either the entire transaction succeeds (COMMIT) or it fails (ABORT / ROLLBACK). If it fails, the application can safely RETRY.**
>
> **Transactions are not a law of nature; they were created with a purpose — to simplify the programming model for applications accessing a database.** Using them lets the application ignore certain error scenarios and concurrency issues because the database handles them instead. We call these **safety guarantees**.
**And the counterpoint, stated fairly:** *not every application needs transactions, and sometimes there are advantages to weakening or abandoning them (better performance, higher availability). Some safety properties can be achieved without transactions.* **On the other hand — the technical cause behind the Post Office Horizon scandal (Ch 2) was probably a lack of ACID transactions in the underlying accounting system.**
#### The historical arc, corrected [#the-historical-arc-corrected]
***
# 8.1 The Meaning of ACID (/docs/ddia/transactions/meaning-acid)
Coined **1983 by Theo Härder and Andreas Reuter**, to establish precise terminology for fault-tolerance mechanisms.
> **In practice, one database's implementation of ACID does not equal another's.** There is a lot of ambiguity around **isolation**. **The high-level idea is sound, but the devil is in the details. Today, when a system claims to be "ACID compliant," it's unclear what guarantees you can actually expect. "ACID" has unfortunately become mostly a MARKETING TERM.**
*(And **BASE** — basically available, soft state, eventual consistency — **is even more vague. The only sensible definition of BASE is "not ACID."**)*
#### 1.1 Atomicity — really "abortability" [#11-atomicity--really-abortability]
> **In ACID, atomicity is NOT about concurrency.** It does not describe what happens if several processes access the same data at once — **that's covered under I, for isolation.**
>
> **ACID atomicity describes what happens if a client wants to make several writes but a FAULT OCCURS AFTER SOME OF THE WRITES HAVE BEEN PROCESSED** — a process crashes, a network connection is interrupted, a disk becomes full, or an integrity constraint is violated. **The transaction is aborted and the database must discard or undo any writes it made.**
**Why it matters:** without atomicity, **if an error occurs partway through, it's difficult to know which changes took effect and which didn't. The application could try again — but that risks making some changes twice, leading to duplicate or incorrect data.**
> **The ability to abort a transaction on error and have all its writes discarded is THE DEFINING FEATURE of ACID atomicity. Perhaps ABORTABILITY would have been a better term.**
#### 1.2 Consistency — the letter that doesn't belong [#12-consistency--the-letter-that-doesnt-belong]
**"Consistency" has at least FIVE meanings in this book:**
| # | Meaning | Where |
| - | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------- |
| 1 | **Replica consistency / eventual consistency** | Ch 6 |
| 2 | **A consistent snapshot** — the whole DB as it existed at one moment; precisely, **consistent with the happens-before relation**: if it contains a value written at a time, it also reflects all writes that happened before that write | Ch 6, Ch 8 |
| 3 | **Consistent hashing** — a sharding/rebalancing approach | Ch 7 |
| 4 | **Linearizability** — what "C" means in the CAP theorem | Ch 10 |
| 5 | **ACID consistency** — an **application-specific notion of the database being in a "good state"** | Ch 8 |
**ACID consistency = INVARIANTS that must always be true** — e.g. **in an accounting system, credits and debits across all accounts must always be balanced.** If a transaction starts valid and its writes preserve validity, invariants are always satisfied. **(An invariant may be temporarily violated DURING execution, but must be satisfied again AT COMMIT.)**
**To have the database enforce them, declare them as constraints in the schema:** foreign-key constraints, uniqueness constraints, **check constraints** (restricting values in an individual row). **More complex requirements can sometimes be modeled with triggers or materialized views.**
> **However, complex invariants can be difficult or impossible to model with the constraints databases usually provide. Then it's the APPLICATION'S responsibility to define its transactions correctly. If you write bad data violating your invariants but you haven't declared those invariants, THE DATABASE CAN'T STOP YOU.**
>
> **As such, the C in ACID often depends on how the application uses the database and is NOT A PROPERTY OF THE DATABASE ALONE.**
#### 1.3 Isolation [#13-isolation]
> **Isolation means concurrently executing transactions are isolated from each other; they cannot step on each other's toes. The classic textbooks formalize isolation as SERIALIZABILITY: each transaction can pretend it is the only transaction running on the entire database. The database ensures that when the transactions have committed, THE RESULT IS THE SAME AS IF THEY HAD RUN SERIALLY, even though in reality they may have run concurrently.**
**But:** serializability has a performance cost. **Many databases use weaker isolation. Some popular databases, such as Oracle, DON'T EVEN IMPLEMENT IT — Oracle has an isolation level called "serializable," but it actually implements SNAPSHOT ISOLATION, which is weaker.**
#### 1.4 Durability [#14-durability]
> **The promise that after a transaction commits successfully, any data it wrote will not be forgotten, even if there is a hardware fault or the database crashes.**
* **Single-node:** written to nonvolatile storage. Regular file writes are **buffered in memory**, so databases use **`fsync`**. Plus a **write-ahead log** for crashes partway through a write, and **checksums** to detect corrupted or incomplete log entries.
* **Replicated:** data **successfully copied to a certain number of nodes.** The database must **wait until these replications complete before reporting success.**
> **Perfect durability does not exist; if all your hard disks and all your backups are destroyed at the same time, there's obviously nothing your database can do to save you.**
**The eight-point reality check on durability (this list is the most sobering in the chapter):**
***
# 8.8 Production failure catalog for this chapter (/docs/ddia/transactions/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ------------------------------------------------------ | ---------------------------------------------------------------------- |
| Two increments, counter went up by one | **Lost update** — read-modify-write race |
| User sees a new email but a zero unread count | **Dirty read** (or missing multi-object isolation) |
| Sale awarded to one buyer, invoice sent to another | **Dirty write** across two tables |
| $100 appears to vanish between two account reads | **Read skew** under read-committed |
| A backup taken over 3 hours contains inconsistent data | Backup without **snapshot isolation** |
| Both doctors went off call | **Write skew** — snapshot isolation is not enough |
| Two bookings for the same room at the same time | **Phantom** — nothing existed to lock |
| Two accounts registered with the same username | Write skew — **fixable with a uniqueness constraint** |
| Balance went negative despite a check | Write skew on an aggregate — double-spending |
| User's work discarded on a transient error | **ORM doesn't retry aborted transactions** |
| Retrying made the outage worse | Retrying a **contention/overload** error instead of backing off |
| Email sent twice | **Side effect outside the database** on a retried transaction |
| Everything hangs; one query holds a whole-table lock | **2PL with no suitable index** for a range lock |
| Latency p99 collapses under contention | **2PL** blocking; one slow transaction stalls the system |
| Massive abort rate after enabling serializable | **SSI under high contention**, or predicate lock escalation |
| Vacuum can't keep up; table bloats indefinitely | A **long-running transaction** pinning the MVCC horizon |
| Database refuses writes entirely | **Transaction ID wraparound** |
| Rows locked for 20 minutes, then forever | **In-doubt 2PC transaction**; coordinator crashed or lost its log |
| Two systems permanently disagree | **Heuristic decision** on an in-doubt XA transaction |
| Committed transactions get aborted on recovery | **Coordinator lost the most recent part of its log** |
| A node shuts itself down for no visible reason | **Clock offset** exceeded the configured maximum (Spanner/CockroachDB) |
| Cluster throughput collapses at 1,000 writes/s | **Cross-shard transactions** in a serial-execution system |
***
# 8.11 Self-test (/docs/ddia/transactions/self-test)
Why would "abortability" be a better name than "atomicity"? What is atomicity explicitly *not* about?
Give all five meanings of "consistency" in this book. Which one is the C in ACID, and why is it "not a property of the database alone"?
Name four reasons `fsync`-based durability is weaker than it sounds.
Why do storage engines universally provide atomicity and isolation for *single* objects? Give the three failure questions that motivate it.
Give three data-model situations that force you to need multi-object transactions.
What is the whole point of rolling back a transaction, and why do popular ORMs defeat it?
List the five caveats on retrying an aborted transaction. Which one can cause a double charge?
Define dirty read and dirty write. Which does read-committed prevent, and which race condition does it *not* prevent?
Why don't databases use read locks to prevent dirty reads? What do they do instead?
Walk through the read-skew example. Why is it "acceptable" under read-committed, and where is it *not* tolerable?
State the MVCC visibility rule for a row in one sentence with two clauses. Why are writes by *aborted* transactions convenient to ignore?
Why does PostgreSQL call snapshot isolation "repeatable read"? Give two other databases that use that name for something different.
Give five ways to prevent lost updates. Which two stop working under multi-leader/leaderless replication, and why?
Define write skew. Why is it a *generalization* of lost update?
Why can't `SELECT FOR UPDATE` fix the meeting-room-booking case, but it can fix the doctors case?
What is a phantom? Why does snapshot isolation prevent phantoms in read-only queries but not in read/write ones?
What is materializing conflicts, and why is it a last resort?
Why do serial-execution systems forbid interactive multistatement transactions? What replaces them?
Give the four constraints under which serial execution is viable, and the cross-shard throughput number.
State the 2PL rule for readers and writers, and contrast it with snapshot isolation's mantra. Which phase is "growing" and which is "shrinking"?
Why don't predicate locks perform well? Why is it *safe* to approximate a predicate with a broader one?
What happens under 2PL when there is no suitable index for a range lock?
Explain SSI's two detection mechanisms. Why does SSI wait until commit rather than aborting on a stale read?
Why is a "tripwire" different from a lock, and why does that make SSI's latency more predictable than 2PL's?
Why isn't sending commit to every node sufficient? Give the concrete reason a committed node can't retract.
List the six steps of 2PC and name the two points of no return.
A participant voted yes and the coordinator crashed. Enumerate what the participant may and may not do, and why a timeout doesn't help.
Why does 3PC not solve this in practice?
Give the four fundamental problems with XA. Which one survives even a replicated coordinator?
What is a "heuristic decision" a euphemism for?
List the four fixes that make database-internal distributed transactions work where XA doesn't.
Reproduce the four-step idempotency protocol for exactly-once processing, and show that no crash window produces a duplicate side effect.
Fill in the anomaly table from memory. Which cell is "depends," and on what?
you run a ticketing system. Invariants: (a) a seat is sold at most once, (b) a customer's balance never goes negative, (c) an event never oversells past capacity even across 200 concurrent purchase requests. The data is sharded by event ID; balances live in a separate service. Choose isolation levels, concurrency-control mechanisms, and a strategy for the cross-service part. State exactly which anomaly each choice prevents, and where you're relying on idempotency rather than atomicity.
# 8.4 Serializability (/docs/ddia/transactions/serializability)
**The bleak summary the chapter arrives at:**
> * **Isolation levels are hard to understand and inconsistently implemented** (the meaning of "repeatable read" varies significantly)
> * **It can be difficult to tell BY LOOKING AT THE APPLICATION CODE whether it is safe to run at a particular isolation level** — especially in a large application where you might not know everything happening concurrently
> * **There are NO GOOD TOOLS to help us detect race conditions.** Static analysis may help in principle, **but research techniques have not yet found their way into practical use. Testing is hard, because problems occur only if you get unlucky with the timing**
>
> **This has been the situation since the 1970s. All along, the answer from researchers has been simple: USE SERIALIZABLE ISOLATION.**
**Three implementations:**
#### 4.1 Actual Serial Execution [#41-actual-serial-execution]
**Two developments made it viable, only in the 2000s:**
1. **RAM became cheap enough to keep the entire active dataset in memory** — when all data is in memory, transactions execute much faster
2. **Designers realized OLTP transactions are usually SHORT and make only a SMALL number of reads and writes.** Long-running analytical queries are typically **read-only, so they can run on a consistent snapshot OUTSIDE the serial execution loop**
> **A system designed for single-threaded execution can sometimes PERFORM BETTER than one supporting concurrency, because it avoids the coordination overhead of locking. However, its throughput is LIMITED TO THAT OF A SINGLE CPU CORE.**
##### Why stored procedures are mandatory here [#why-stored-procedures-are-mandatory-here]
**The history:** designers originally intended a transaction to encompass an entire flow of user activity — **booking an airline ticket: search routes/fares/seats, decide, book seats on each flight, enter passenger details, pay.** **Unfortunately, HUMANS ARE VERY SLOW TO MAKE UP THEIR MINDS.** A transaction waiting for user input would require **a huge number of concurrent, mostly idle transactions.** So **almost all OLTP applications keep transactions short** — on the web, **a transaction is committed within the same HTTP request; a new request starts a new transaction.**
**But even with humans removed, the interactive client/server style remains:**
> **Therefore systems with single-threaded serial transaction processing DON'T ALLOW interactive multistatement transactions.** The application must **either limit itself to single-statement transactions or submit the entire transaction code ahead of time as a STORED PROCEDURE.**
**Stored procedures' bad reputation — four reasons, all fair:**
1. **Each vendor had its own language** (PL/SQL, T-SQL, PL/pgSQL) which **haven't kept up with general-purpose languages — ugly and archaic, lacking the ecosystem of libraries**
2. **Code running in a database is difficult to manage:** harder to debug, awkward to version control and deploy, trickier to test, **difficult to integrate with a metrics collection system**
3. **A database is much more performance-sensitive than an application server**, because one instance is shared by many app servers. **A badly written stored procedure can cause much more trouble than equivalent bad code in an app server**
4. **In a multitenant system allowing tenants to write stored procedures, it's a SECURITY RISK to execute untrusted code in the same process as the database kernel**
**But modern implementations fixed #1:** **VoltDB uses Java or Groovy, Datomic uses Java or Clojure, Redis uses Lua, MongoDB uses JavaScript.**
*(A modern use case the book calls out: **GraphQL proxies exposing the database directly. If the proxy doesn't support complex validation logic, you can embed it in a stored procedure — otherwise you must deploy a validation service between the proxy and the database.**)*
**VoltDB also uses stored procedures for REPLICATION:** instead of copying writes, **it executes the same stored procedure on each replica. This REQUIRES stored procedures to be DETERMINISTIC** — a transaction needing the current date/time must use **special deterministic APIs.** **This is STATE MACHINE REPLICATION** (Ch 10).
##### Sharding for serial execution [#sharding-for-serial-execution]
**To scale beyond one core, shard.** **If each transaction reads and writes within a SINGLE SHARD, each shard gets its own transaction processing thread — give each CPU core its own shard and throughput scales LINEARLY with cores.**
> **But any transaction touching multiple shards must be coordinated across all of them — the stored procedure must be performed IN LOCKSTEP across all shards.**
>
> **VoltDB reports about 1,000 CROSS-SHARD writes per second — ORDERS OF MAGNITUDE below its single-shard throughput, AND IT CANNOT BE INCREASED BY ADDING MORE MACHINES.**
**Whether transactions can be single-shard depends on the data:** **simple key-value data shards easily; data with multiple secondary indexes is likely to require a lot of cross-shard coordination** (Ch 7 §5).
**The four constraints, summarized:**
1. **Every transaction must be small and fast — IT TAKES ONLY ONE SLOW TRANSACTION TO STALL ALL TRANSACTION PROCESSING**
2. **The active dataset must fit in memory.** Rarely accessed data could go to disk, **but if a single-threaded transaction needed it, the system would get very slow**
3. **Write throughput must fit on one CPU core, or transactions must shard without cross-shard coordination**
4. **Cross-shard transactions are possible, but their throughput is hard to scale**
#### 4.2 Two-Phase Locking (2PL) [#42-two-phase-locking-2pl]
> ⚠️ **2PL IS NOT 2PC.** **2PL provides serializable ISOLATION; 2PC provides atomic COMMIT in a distributed database.** *Best to think of them as entirely separate concepts and ignore the unfortunate similarity in the names.*
**The rule:**
**Implementation — a shared/exclusive (multi-reader single-writer) lock on each object:**
* **Read** → acquire the lock in **shared mode.** Several transactions may hold it in shared mode; **if another holds it exclusively, wait**
* **Write** → acquire the lock in **exclusive mode.** No other transaction may hold it at all
* **Read then write** → **upgrade** the shared lock to exclusive
* **Hold every lock until the END of the transaction** (commit or abort)
> **This is where "two-phase" comes from: the first phase (GROWING) is when locks are acquired while the transaction executes; the second phase (SHRINKING) is when all locks are released at the end. THE TWO PHASES MUST NOT OVERLAP — once a lock is released, no new locks may be acquired.**
**Deadlock is frequent.** **The database automatically detects deadlocks and aborts one transaction; the application must retry.**
##### Performance — why it hasn't been the default since the 1970s [#performance--why-it-hasnt-been-the-default-since-the-1970s]
> **Transaction throughput and response times are SIGNIFICANTLY WORSE under 2PL. Partly the overhead of acquiring and releasing locks — but MORE IMPORTANTLY, REDUCED CONCURRENCY. By design, if two concurrent transactions try to do anything that MAY IN ANY WAY result in a race condition, one has to wait.**
**The pathological case:**
> **Databases running 2PL can have quite UNSTABLE LATENCIES, and can be VERY SLOW AT HIGH PERCENTILES if there is contention. JUST ONE SLOW TRANSACTION, or one that accesses a lot of data and acquires many locks, could cause the rest of the system to grind to a halt.** *Transaction timeouts and slow-query monitoring are used to detect and limit misbehaving queries.*
**And deadlocks occur MUCH more frequently under 2PL than under lock-based read-committed.** **When a deadlocked transaction is aborted and retried, IT NEEDS TO DO ITS WORK ALL OVER AGAIN — if deadlocks are frequent, significant wasted effort.**
##### Predicate locks — how 2PL handles phantoms [#predicate-locks--how-2pl-handles-phantoms]
> **A database with serializable isolation MUST prevent phantoms.** Conceptually you need a **PREDICATE LOCK** — like a shared/exclusive lock, but **rather than belonging to a particular object, it belongs to ALL OBJECTS THAT MATCH A SEARCH CONDITION.**
```sql
SELECT * FROM bookings
WHERE room_id = 123 AND end_time > '2026-01-01 12:00'
AND start_time < '2026-01-01 13:00';
```
* **A wants to READ objects matching a condition** → acquire a **shared-mode predicate lock on the query's conditions.** If B holds an exclusive lock on any matching object, A waits
* **A wants to INSERT/UPDATE/DELETE** → **first check whether EITHER THE OLD OR THE NEW VALUE matches any existing predicate lock.** If B holds a matching one, A waits
> **The key idea: A PREDICATE LOCK APPLIES EVEN TO OBJECTS THAT DO NOT YET EXIST IN THE DATABASE but might be added in the future (phantoms). If 2PL includes predicate locks, the database prevents all forms of write skew and other race conditions — its isolation becomes serializable.**
##### Index-range locks (next-key locking) — what's actually implemented [#index-range-locks-next-key-locking--whats-actually-implemented]
> **Predicate locks DO NOT PERFORM WELL: with many locks held by active transactions, CHECKING FOR MATCHING LOCKS BECOMES TIME-CONSUMING.** So most 2PL databases implement **index-range locking, a simplified APPROXIMATION.**
> **It's SAFE to simplify a predicate by making it match a GREATER set of objects.** A lock on bookings of room 123 between noon and 1 pm can be approximated by **locking bookings for room 123 at ANY time**, or **locking ALL ROOMS between noon and 1 pm.** **Any write matching the original predicate will definitely also match the approximation.**
> **Index-range locks are NOT AS PRECISE as predicate locks (they may lock a bigger range than strictly necessary), but since they have MUCH LOWER OVERHEADS, they are a GOOD COMPROMISE.**
>
> **If there is NO SUITABLE INDEX to attach a range lock to, the database FALLS BACK TO A SHARED LOCK ON THE ENTIRE TABLE. Not good for performance — it stops all other writes — but a SAFE fallback.**
*(Practical corollary: **a missing index doesn't just make a serializable query slow — it makes it lock the whole table.**)*
#### 4.3 Serializable Snapshot Isolation (SSI) [#43-serializable-snapshot-isolation-ssi]
> **Are serializable isolation and good performance fundamentally at odds? IT SEEMS NOT: SSI provides FULL SERIALIZABILITY with only a SMALL PERFORMANCE PENALTY compared to snapshot isolation.** First described **2008**.
##### Pessimistic vs optimistic [#pessimistic-vs-optimistic]
| | Principle |
| ------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **2PL — pessimistic** | **If anything MIGHT possibly go wrong (indicated by a lock held by another transaction), it's better to WAIT until the situation is safe.** Like mutual exclusion in multithreaded programming |
| **Serial execution — pessimistic to the extreme** | **Essentially equivalent to each transaction holding an exclusive lock on the ENTIRE DATABASE (or shard). We compensate by making each transaction very fast, so it holds the "lock" only briefly** |
| **SSI — optimistic** | **Instead of blocking when something potentially dangerous happens, transactions CONTINUE ANYWAY, in the hope that everything will turn out all right. At commit, the database checks whether isolation was violated; if so, ABORT AND RETRY. ONLY TRANSACTIONS THAT EXECUTED SERIALIZABLY ARE ALLOWED TO COMMIT** |
**When optimistic loses:** **it performs badly under HIGH CONTENTION (many transactions accessing the same objects), leading to a high proportion of aborts. If the system is already close to maximum throughput, THE ADDITIONAL LOAD FROM RETRIED TRANSACTIONS CAN MAKE PERFORMANCE WORSE.**
**When it wins:** **with enough spare capacity and not-too-high contention, optimistic tends to perform BETTER than pessimistic.**
**Reducing contention:** **commutative atomic operations** — several transactions incrementing a counter **don't care about order (as long as the counter isn't read in the same transaction), so the concurrent increments can all be applied without conflicting.**
##### The core idea: decisions based on an outdated premise [#the-core-idea-decisions-based-on-an-outdated-premise]
> **The recurring write-skew pattern: a transaction reads data, examines the result, and DECIDES to take an action based on what it saw. Under snapshot isolation, THE RESULT MAY NO LONGER BE UP TO DATE BY THE TIME THE TRANSACTION COMMITS.**
>
> **The transaction is acting on a PREMISE — "there are currently two doctors on call." Later, the premise may no longer be true.**
>
> **The database doesn't know how the application logic uses the query result. To be safe, IT MUST ASSUME THAT ANY CHANGE IN THE QUERY RESULT MEANS WRITES IN THAT TRANSACTION MAY BE INVALID** — there may be a causal dependency between the queries and the writes. **So it must detect situations where a transaction MAY HAVE ACTED ON AN OUTDATED PREMISE and abort.**
**Two detection cases:**
##### (a) Detecting stale MVCC reads — *the write happened BEFORE the read, but committed after* [#a-detecting-stale-mvcc-reads--the-write-happened-before-the-read-but-committed-after]
> **The database TRACKS when a transaction ignores another transaction's writes because of MVCC visibility rules. At commit, it checks whether any ignored writes have NOW been committed. If so, abort.**
**Why wait until commit rather than aborting immediately?** Three reasons:
1. **If transaction 43 were READ-ONLY, it wouldn't need to abort — there's no risk of write skew.** At read time the database doesn't yet know whether it will later write
2. **Transaction 42 may yet ABORT, or may still be uncommitted when 43 commits — so the read may turn out NOT to have been stale after all**
3. **By avoiding unnecessary aborts, SSI PRESERVES SNAPSHOT ISOLATION'S SUPPORT FOR LONG-RUNNING READS from a consistent snapshot**
##### (b) Detecting writes that affect prior reads — *the write happens AFTER the read* [#b-detecting-writes-that-affect-prior-reads--the-write-happens-after-the-read]
**This information is kept only for a while: after a transaction finishes and all concurrent transactions finish, the database can forget what it read.**
##### Performance of SSI [#performance-of-ssi]
**Granularity trade-off:** **detailed tracking → precise aborts but significant bookkeeping overhead. Less detailed → faster, but more transactions aborted than strictly necessary.**
**PostgreSQL reduces unnecessary aborts** using theory showing **it's sometimes OK for a transaction to read information overwritten by another — depending on what else happened, the execution may still be provably serializable.**
| Compared to | SSI's advantage |
| ---------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **2PL** | **One transaction doesn't block waiting for locks held by another. Writers don't block readers, and vice versa. This makes QUERY LATENCY MUCH MORE PREDICTABLE AND LESS VARIABLE. Read-only queries can run on a consistent snapshot WITHOUT ANY LOCKS — very appealing for read-heavy workloads** |
| **Serial execution** | **NOT limited to a single CPU core. FoundationDB DISTRIBUTES the detection of serialization conflicts across multiple machines**, scaling to very high throughput; **transactions can read and write in multiple shards while ensuring serializable isolation** |
| **Nonserializable SI** | Some overhead. **How significant is a matter of debate: some believe serializability checking is not worth it; others believe its performance is now so good that there is no need to use the weaker snapshot isolation anymore** |
> **The RATE OF ABORTS significantly affects overall performance. A transaction that reads and writes over a LONG PERIOD is likely to conflict and abort — SSI requires READ/WRITE transactions to be fairly SHORT (long-running READ-ONLY transactions are fine). However, SSI is LESS SENSITIVE TO SLOW TRANSACTIONS than 2PL or serial execution.**
***
# 8.2 Single-Object and Multi-Object Operations (/docs/ddia/transactions/single-object-multi-object)
**Recap of what A and I promise when a client makes several writes:**
* **Atomicity** — error halfway → abort, discard writes so far. **An ALL-OR-NOTHING guarantee.**
* **Isolation** — **another transaction should see either ALL or NONE of a transaction's writes, but not a subset.**
**The motivating example — an unread-message counter:**
```sql
SELECT COUNT(*) FROM emails WHERE recipient_id = 2 AND unread_flag = true
```
Too slow with many emails → **denormalize** into a separate counter field, incremented on new mail and decremented on read.
**How transactions are delimited:** in relational databases, **typically by the client's TCP connection** — everything between `BEGIN TRANSACTION` and `COMMIT` on a connection. **If the TCP connection is interrupted, the transaction must be aborted.**
> **Many nonrelational databases don't have such a grouping mechanism. Even a multi-object API (a multi-put updating several keys) DOESN'T NECESSARILY MEAN TRANSACTION SEMANTICS: the command may succeed for some keys and fail for others, leaving the database PARTIALLY UPDATED.**
#### 2.1 Single-object writes [#21-single-object-writes]
**Three questions that make the need obvious**, for writing a 20 kB JSON document:
* **Network interrupted after 10 kB — does the database store that unparseable fragment?**
* **Power fails mid-overwrite — do you get old and new values SPLICED TOGETHER?**
* **Another client reads during the write — does it see a partially updated value?**
> **Each outcome would be incredibly confusing, so storage engines ALMOST UNIVERSALLY provide atomicity and isolation at the level of a SINGLE OBJECT on ONE NODE.** Atomicity via **a log for crash recovery**; isolation via **a lock on each object.**
**Two richer single-object primitives:**
* **Atomic increment** — removes the read-modify-write cycle
* **Conditional write** — a write happens **only if the value has not been concurrently changed**; the database equivalent of **compare-and-set (CAS)**
*(Pedantic but useful: **"atomic increment" uses "atomic" in the multithreading sense. In ACID terms it should be called an ISOLATED or SERIALIZABLE increment.**)*
> **These are NOT transactions in the usual sense.** Aerospike's "strong consistency" mode and Cassandra/ScyllaDB's "lightweight transactions" offer **linearizable reads and conditional writes on a SINGLE object, but NO guarantees across multiple objects.**
#### 2.2 Why you need multi-object transactions [#22-why-you-need-multi-object-transactions]
| Data model | Why |
| -------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Relational** | A row **has a foreign-key reference to a row in another table**; in graphs, a vertex has edges. **Multi-object transactions ensure these references remain valid** — when inserting several records that refer to one another, **the foreign keys have to be correct and up to date, or the data becomes nonsensical** |
| **Document** | Fields updated together are often **within the same document** — no multi-object transaction needed. **BUT document databases lacking joins ENCOURAGE DENORMALIZATION, and when denormalized information needs updating you must update several documents in one go** |
| **Anything with secondary indexes** (almost everything but pure key-value) | **The indexes must be updated on every value change. Indexes are DIFFERENT DATABASE OBJECTS from a transaction point of view** — without isolation, **a record can appear in one index but not another** because the second index update hasn't happened yet |
#### 2.3 Handling errors and aborts [#23-handling-errors-and-aborts]
> **ACID databases are based on the philosophy that if the database is in danger of violating atomicity, isolation, or durability, IT WOULD RATHER ABANDON THE TRANSACTION ENTIRELY than allow it to remain half-finished.**
>
> **Datastores with leaderless replication work on more of a "BEST EFFORT" basis: "the database will do as much as it can, and if it runs into an error, IT WON'T UNDO SOMETHING IT HAS ALREADY DONE" — so it's the application's responsibility to recover.**
**The retry indictment:**
> **Popular ORM frameworks such as Rails ActiveRecord and Django DON'T RETRY aborted transactions — the error usually results in an exception bubbling up the stack, so any user input is thrown away and the user gets an error message. THIS IS A SHAME, BECAUSE THE WHOLE POINT OF ROLLING BACK TRANSACTIONS IS TO ENABLE SAFE RETRIES.**
**But retrying isn't perfect — five caveats:**
1. **The transaction actually succeeded but the acknowledgment was lost** → **retrying performs it twice** unless you have application-level deduplication
2. **If the error is due to overload or high contention, retrying MAKES IT WORSE.** Limit retries, use **exponential backoff**, and **handle overload errors differently from other errors** (Ch 2)
3. **Retry only after TRANSIENT errors** (deadlock, isolation violation, temporary network interruption, failover). **After a PERMANENT error (constraint violation) a retry is pointless**
4. **Side effects outside the database may happen even if the transaction is aborted** — e.g. **you wouldn't want to send the email again on every retry.** (2PC can help, §6)
5. **If the client process crashes while retrying, any data it was writing is lost**
***
# 8.7 Technology deep dives (/docs/ddia/transactions/technology-deep-dives)
***
#### 7.1 PostgreSQL MVCC + SSI [#71-postgresql-mvcc--ssi]
**Problem it solves.** Full serializability without readers blocking writers, on a single node, with predictable latency.
**Why not 2PL?** Unstable latencies, terrible high percentiles under contention, a table-scanning transaction blocking all writers, frequent deadlocks (§4.2).
**Why not serial execution?** Bounded by one core, requires stored procedures and an in-memory dataset.
**How it works internally.** §3.2's MVCC (`xmin`/`xmax` on every tuple, versions in the heap, visibility computed from a snapshot's transaction list) plus §4.3's SSI (SIREAD locks recording what each transaction read, tracking rw-dependencies, aborting a transaction that sits in a "dangerous structure" of two consecutive rw-edges). Serialization failures surface as `SQLSTATE 40001`.
**Deployment.** `default_transaction_isolation = 'serializable'` per-database or per-transaction; **the application MUST have a retry loop on 40001**, or serializable mode is worse than useless.
**Monitoring.**
* **Serialization failure rate** (40001) and deadlock rate (40P01) — the two abort classes, with different causes
* `max_pred_locks_per_transaction` exhaustion — when exceeded, **predicate locks are escalated to coarser granularity, causing a cascade of false-positive aborts**
* **Long-running transactions** (`pg_stat_activity` where `xact_start` is old) — these hold back the xmin horizon, block vacuum, and inflate SSI tracking
* **Table and index bloat**, autovacuum progress, and **transaction ID wraparound age** (`age(datfrozenxid)`) — the §3.2 32-bit txid hazard, and a genuine "database refuses writes" outage
* Lock waits by lock type
**Scaling.** Read replicas for read-only work at snapshot isolation. Write throughput is single-node.
**Backup.** Base backup + WAL archiving. **Note the §3.2 connection: a consistent backup is exactly a long-running snapshot-isolation read** — which is why backups and vacuum fight each other.
**What actually breaks.**
* **No retry loop.** The application gets 40001 and shows the user an error. **This is the single most common serializable-Postgres failure, and it's an application bug** — §2.3's ORM indictment made concrete.
* **A long-running analytics query on the primary** pins the xmin horizon, vacuum can't remove dead tuples, bloat grows unbounded, and eventually every query slows.
* **Predicate lock escalation** turning a moderate workload into an abort storm.
* **`SELECT FOR UPDATE` deadlocks** from inconsistent lock ordering across code paths.
* **Assuming "repeatable read" is repeatable read.** It's snapshot isolation, and **it does not prevent write skew.**
* **Transaction ID wraparound** on a high-churn table with autovacuum tuned too conservatively.
***
#### 7.2 MySQL / InnoDB (2PL for serializable, gap locks for phantoms) [#72-mysql--innodb-2pl-for-serializable-gap-locks-for-phantoms]
**Problem it solves.** Serializability using the classical, well-understood locking approach, with next-key locks handling phantoms.
**Why gap locks?** §4.2's index-range locking: a predicate lock is too expensive to evaluate, so InnoDB locks **the index record plus the gap before it** — the concrete implementation of "approximating a predicate by matching a greater set of objects."
**How it works internally.** MVCC for consistent reads at REPEATABLE READ; shared/exclusive record locks plus **gap locks** and **next-key locks** (record + preceding gap) for locking reads and for SERIALIZABLE. At SERIALIZABLE, plain `SELECT` is implicitly `SELECT … LOCK IN SHARE MODE`.
**Monitoring.** `SHOW ENGINE INNODB STATUS` for the latest deadlock; `Innodb_row_lock_time_avg` and `_max`; **lock wait timeouts** (`innodb_lock_wait_timeout`, default 50 s — long enough to look like a hang); history list length (the MVCC undo backlog, InnoDB's analogue of Postgres bloat).
**What actually breaks.**
* **Gap locks blocking inserts** into ranges nobody has rows in — the classic "why is my INSERT waiting on a SELECT" mystery. Especially bad when a query has **no useful index**, because then it locks far more than intended (§4.2's "fallback to a table lock" in a milder form).
* **`REPEATABLE READ` not detecting lost updates** — unlike Postgres/Oracle/SQL Server (§3.3). People assume MySQL RR ≡ Postgres RR. It isn't.
* **Deadlocks under 2PL far more frequent than under read committed**, each one wasting all the transaction's work.
* **A long-running read transaction growing the history list** until purge can't keep up.
* **Lock wait timeout treated as a permanent error** rather than a retryable one.
***
#### 7.3 VoltDB / H-Store (serial execution + stored procedures) [#73-voltdb--h-store-serial-execution--stored-procedures]
**Problem it solves.** Serializability with zero concurrency-control overhead, by removing concurrency.
**Why viable now?** §4.1: cheap RAM (whole active dataset in memory) and the realization that OLTP transactions are short.
**How it works internally.** One single-threaded execution engine per CPU core, each owning a shard; transactions are **stored procedures** submitted whole (no interactive round trips); replication by **executing the same procedure on each replica** — state machine replication, which **requires determinism** (special deterministic APIs for time and randomness).
**Monitoring.** Per-partition transaction latency; **the fraction of transactions that are cross-shard** (the metric that decides whether the architecture works at all); procedure execution time p99 (**one slow procedure stalls its entire partition**); snapshot completion.
**Scaling.** Linear in cores **as long as transactions are single-shard.** **Cross-shard writes: \~1,000/s, orders of magnitude below single-shard, and NOT improvable by adding machines.**
**What actually breaks.**
* **One slow stored procedure stalls everything on that partition** — §4.1's constraint #1, and it's absolute.
* **A partitioning scheme that produces many cross-shard transactions** collapses throughput by orders of magnitude.
* **Nondeterminism in a procedure** breaking replication (the same failure mode as Ch 5's durable execution and Ch 6's statement-based replication — the theme recurs everywhere replay is involved).
* **The working set exceeding memory.**
***
#### 7.4 Spanner / CockroachDB / FoundationDB (distributed transactions done properly) [#74-spanner--cockroachdb--foundationdb-distributed-transactions-done-properly]
**Problem it solves.** ACID across shards and regions, without XA's single-point-of-failure coordinator.
**Why not XA?** All four problems in §5.4: coordinator SPOF, coordinator log on an app server's local disk, no direct coordinator↔participant communication, lowest-common-denominator API.
**How they work internally.** §5.5's four fixes. Each shard is a **Raft/Paxos group** (so a participant failing doesn't abort the transaction — the group fails over). The **coordinator is itself replicated** by consensus, so its commit-point log survives a node loss. **Spanner** uses **TrueTime** (GPS + atomic clocks, exposing a bounded uncertainty interval) to assign globally meaningful commit timestamps and get external consistency, at the cost of a **commit-wait** of a few ms. **CockroachDB** uses HLCs plus an uncertainty window and read refreshes. **FoundationDB** uses **optimistic concurrency with a distributed conflict-detection tier** — SSI, scaled out, which is why §4.3 cites it as the counter-example to "serializable can't scale."
**Monitoring.** Transaction retry rate by reason (write-write conflict, uncertainty restart, read-refresh failure); **contention hotspots by key range** — the single most useful distributed-transaction metric; commit latency broken down by phase; **clock offset** (in Spanner, `TrueTime` epsilon; in CockroachDB, a node exceeding `max-offset` **shuts itself down**, deliberately, to preserve correctness); range/leaseholder distribution.
**Scaling.** Horizontal, genuinely — but **cross-shard transactions cost extra round trips**, so schema design that keeps transactions within one range still matters enormously.
**What actually breaks.**
* **Contention hotspots.** A sequential primary key or a single counter row funnels every transaction through one range, and optimistic concurrency turns that into an abort storm (§4.3's "performs badly under high contention").
* **Clock skew.** A node whose clock drifts past the configured max offset is killed to protect correctness — availability sacrificed for safety, by design.
* **Long-running read/write transactions** conflicting and aborting repeatedly.
* **Retries not implemented in the client**, same as §7.1.
* **Assuming cross-region transactions are cheap** — commit-wait and consensus round trips make them tens to hundreds of ms.
***
#### 7.5 XA / JTA heterogeneous transactions [#75-xa--jta-heterogeneous-transactions]
**Problem it solves.** Atomic commit across a database and a message broker from different vendors — §5.3's exactly-once.
**Why it's the wrong tool now.** §5.4's four problems, plus the fact that §5.6 shows **you can get exactly-once with idempotency and a message-ID table, using only local transactions.**
**Monitoring (if you must run it).** **Count of in-doubt/prepared transactions** (`pg_prepared_xacts` in Postgres, `XA RECOVER` in MySQL) — this should be zero at rest, and a nonzero steady state means orphans are accumulating; **age of the oldest prepared transaction**; lock waits attributable to prepared transactions; coordinator log disk health.
**What actually breaks.**
* **Orphaned in-doubt transactions holding locks forever**, surviving database restarts by design, requiring manual administrator resolution under outage pressure.
* **`max_prepared_transactions` set to 0** (the Postgres default) so XA silently doesn't work — or set high and forgotten, so a leaked prepared transaction blocks vacuum forever.
* **Heuristic decisions** silently breaking atomicity and leaving two systems permanently disagreeing.
* **The application server dying** and taking the coordinator's log with it.
* **No cross-system deadlock detection** — the deadlock just hangs until a timeout somewhere.
***
# 8.12 Terminology introduced here (/docs/ddia/transactions/terminology-introduced-here)
# 8.3 Weak Isolation Levels (/docs/ddia/transactions/weak-isolation-levels)
> **Concurrency bugs are hard to find by testing, because they are triggered only when you get unlucky with the timing. They occur rarely and are usually difficult to reproduce. Concurrency is also difficult to reason about, especially in a large application where you don't necessarily know which other pieces of code are accessing the database.**
**And they are not theoretical:**
> **Concurrency bugs caused by weak isolation have caused substantial loss of money, INCLUDING BANKRUPTING A BITCOIN EXCHANGE, led to investigation by financial auditors, and caused customer data to be corrupted.**
>
> **A popular comment is "Use an ACID database if you're handling financial data!" — but that MISSES THE POINT. Even many popular relational database systems (usually considered ACID) USE WEAK ISOLATION, so they wouldn't necessarily have prevented these bugs.**
**And the security framing, which people forget:**
> **Even if concurrency issues are rare in normal operation, you have to consider that AN ATTACKER MIGHT DELIBERATELY SEND A BURST OF HIGHLY CONCURRENT REQUESTS TO YOUR API in an attempt to exploit concurrency bugs.** To build applications that are reliable **and secure**, such bugs must be **systematically prevented.**
*(A great aside: **much of the banking system relies on text files exchanged via secure FTP. In this context, having an AUDIT TRAIL and human-level fraud prevention is actually more important than ACID properties.**)*
#### 3.1 Read Committed [#31-read-committed]
**Two guarantees:**
1. **No dirty reads** — you read only committed data
2. **No dirty writes** — you overwrite only committed data
**It is the DEFAULT in Oracle, PostgreSQL, SQL Server, and many others.**
##### No dirty reads [#no-dirty-reads]
**Why preventing dirty reads matters:**
* **A transaction updating several rows** → a dirty read means **another transaction sees SOME updates but not others** (the email/counter case). **Seeing a partially updated state is confusing to users and may cause other transactions to make INCORRECT DECISIONS.**
* **If a transaction aborts, its writes are rolled back.** With dirty reads, **a transaction may see data that is LATER ROLLED BACK — data that was never actually committed. Any transaction that read uncommitted data would ALSO need to be aborted → CASCADING ABORTS.**
##### No dirty writes [#no-dirty-writes]
**A dirty write = a later write overwrites a value written by a transaction that hasn't committed yet.** Prevented **usually by DELAYING the second write until the first transaction commits or aborts.**
> ⚠️ **BUT read-committed does NOT prevent the lost-update race of §1.3.** There, **the second write happens AFTER the first transaction committed — so it's not a dirty write. It's still incorrect, but for a different reason.**
##### Implementing read committed [#implementing-read-committed]
**Dirty writes → row-level locks.** A transaction wanting to modify a row **must acquire a lock and hold it until commit or abort.** Only one transaction holds it at a time. **Done automatically in read-committed mode and stronger.**
**Dirty reads — two options:**
| Option | Problem |
| ------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Read locks** (briefly acquire and release the same lock on read) | **Does not work well in practice: ONE LONG-RUNNING WRITE TRANSACTION CAN FORCE MANY OTHER TRANSACTIONS TO WAIT — even transactions that only read.** Harms read-only response times and is **bad for operability: a slowdown in one part of an application has a knock-on effect in a completely different part.** *(Still used by IBM Db2 and SQL Server with `read_committed_snapshot=off`)* |
| **Remember both values** ✔ | **For every row written, the database remembers BOTH the old committed value AND the new value set by the transaction holding the write lock. Readers are simply given the OLD value while the transaction is ongoing** |
**Read uncommitted** is even weaker: **prevents dirty writes but NOT dirty reads.** Better performance (no need to store two versions) and **reduces the probability of — but does not prevent — lost updates.**
#### 3.2 Snapshot Isolation and Repeatable Read [#32-snapshot-isolation-and-repeatable-read]
**The anomaly read-committed still permits — READ SKEW:**
*(Terminology note: **"skew" is overloaded** — in Ch 7 it meant an unbalanced workload with hot spots; **here it means a TIMING ANOMALY.**)*
**Where temporary inconsistency is NOT tolerable:**
| Case | Why |
| ------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Backups** | Copying the entire database **may take hours**, during which writes continue. **Some parts of the backup contain older data and other parts newer. If you restore from such a backup, the inconsistencies (such as disappearing money) BECOME PERMANENT** |
| **Analytical queries and integrity checks** | Queries scanning large parts of the database **return nonsensical results if they observe parts of the database at different points in time** |
> **Snapshot isolation: EACH TRANSACTION READS FROM A CONSISTENT SNAPSHOT — it sees all the data committed in the database at the start of that transaction. Even if the data is subsequently changed, each transaction sees only the old data from that particular point in time.**
>
> **A boon for long-running read-only queries. It is very hard to reason about the meaning of a query if the data it operates on is changing at the same time.**
**Supported by PostgreSQL, MySQL/InnoDB, Oracle, SQL Server. Some databases — Oracle, TiDB, Aurora DSQL — choose snapshot isolation as their HIGHEST isolation level. Cloud warehouses like BigQuery frequently use it**, since it provides a point-in-time view for analytical queries.
##### Multiversion concurrency control (MVCC) [#multiversion-concurrency-control-mvcc]
> **The key performance principle of snapshot isolation: READERS NEVER BLOCK WRITERS, AND WRITERS NEVER BLOCK READERS.** This lets a database handle long-running read queries on a consistent snapshot **at the same time as processing writes normally, without any lock contention between the two.** *(Writes still take write locks against each other, to prevent dirty writes.)*
**Instead of two versions per row, the database keeps SEVERAL COMMITTED VERSIONS, because various in-progress transactions may need to see different points in time.**
**PostgreSQL's implementation:**
*(Precision note: **PostgreSQL txids are 32-bit, so they overflow after \~4 billion transactions. The VACUUM process performs cleanup to ensure overflow doesn't affect the data.** This is the transaction-ID-wraparound hazard.)*
##### Visibility rules [#visibility-rules]
> **A row is visible if BOTH of these are true:**
>
> 1. **At the time the reader's transaction started, the transaction that INSERTED the row had already committed**
> 2. **The row is not marked for deletion, or if it is, the transaction that requested deletion HAD NOT YET COMMITTED at the time the reader's transaction started**
**Mechanically, four rules:**
1. **At transaction start, the database lists all OTHER transactions in progress at that time. Any writes those transactions make are IGNORED, EVEN IF THEY SUBSEQUENTLY COMMIT** — this is what makes the snapshot unaffected by another transaction committing
2. **Writes by transactions with a LATER txid are ignored, regardless of whether they committed**
3. **Writes by ABORTED transactions are ignored, regardless of when the abort happened** — which conveniently means **on abort we don't need to immediately remove rows; the visibility rule filters them out and GC removes them later**
4. All other writes are visible
> **By never updating values in place but instead inserting a new version every time, the database can provide a consistent snapshot while incurring only a SMALL overhead.** A long-running transaction may keep reading values that, from other transactions' points of view, **have long been overwritten or deleted.**
##### Indexes and MVCC [#indexes-and-mvcc]
* **Most common:** each index entry points at **one version of a row** (oldest or newest); each version references the next-oldest/newest. **A query using the index must ITERATE over the rows to find one that is visible AND matches.** When GC removes invisible versions, **the corresponding index entries can also be removed.**
* **PostgreSQL optimization:** avoid index updates **if different versions of the same row fit on the same page** (HOT updates).
* **Some databases store only DIFFERENCES between versions**, to save space.
* **CouchDB, Datomic, LMDB use an IMMUTABLE (copy-on-write) B-tree:** don't overwrite pages, **create a new copy of each modified page; parent pages up to the root are copied and updated to point to the new children. Pages unaffected by a write are SHARED with the new tree.**
> **With immutable B-trees, every write transaction creates a NEW B-TREE ROOT, and a particular root IS a consistent snapshot of the database at the point it was created. There is NO NEED to filter rows by transaction ID, because subsequent writes cannot modify an existing B-tree — they can only create new roots.** (Requires a background compaction/GC process.)
##### The naming disaster [#the-naming-disaster]
**Why:** **the SQL standard has no concept of snapshot isolation, because the standard is based on System R's 1975 definitions and snapshot isolation hadn't been invented.** It defines *repeatable read*, which **looks superficially similar.** PostgreSQL calls its level "repeatable read" **because it meets the standard's requirements and so it can claim standards compliance.**
> **The SQL standard's definition of isolation levels is FLAWED — ambiguous, imprecise, and not as implementation-independent as a standard should be. Even though several databases implement "repeatable read," there are big differences in the guarantees they provide. As a result, NOBODY REALLY KNOWS WHAT REPEATABLE READ ISOLATION MEANS.**
#### 3.3 Preventing Lost Updates [#33-preventing-lost-updates]
**The lost update problem: an application READS a value, MODIFIES it, and WRITES BACK the modified value. If two transactions do this concurrently, one modification can be LOST, because the second write doesn't include the first modification** — *the later write CLOBBERS the earlier write.*
**Where it shows up:**
* **Incrementing a counter or updating an account balance**
* **Making a local change to a complex value** — adding an element to a list within a JSON document (parse, change, write back)
* **Two users editing a wiki page**, each saving by **sending the ENTIRE page contents to the server**, overwriting whatever is currently there
**Five solutions:**
| # | Solution | Detail |
| ----- | -------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **1** | **Atomic write operations** | `UPDATE counters SET value = value + 1 WHERE key = 'foo';` — **usually the BEST solution if your code can be expressed in terms of them.** MongoDB has atomic ops for local JSON modifications; Redis for priority queues. **Implemented by exclusively locking the object when read, or by forcing all atomic operations onto a single thread.** ⚠️ **ORM frameworks make it easy to accidentally write UNSAFE read-modify-write cycles instead — a source of subtle bugs difficult to find by testing** |
| **2** | **Explicit locking** (`SELECT … FOR UPDATE`) | The application locks objects it will update; concurrent updaters must wait. **Needed when correctness requires logic you can't express as a database query** — e.g. a multiplayer game where a move must abide by the rules. ⚠️ **You must carefully think about application logic; it's EASY TO FORGET A LOCK somewhere and introduce a race condition.** ⚠️ **Locking multiple objects risks DEADLOCK** — databases usually detect and abort one; **you retry at the application level** |
| **3** | **Automatic lost-update detection** | Let them run in parallel; **if the transaction manager DETECTS a lost update, abort and force a retry.** Efficient in conjunction with snapshot isolation. **PostgreSQL's repeatable read, Oracle's serializable, and SQL Server's snapshot isolation DETECT this automatically. MySQL/InnoDB's repeatable read DOES NOT.** *Some authors argue a database must prevent lost updates to qualify as snapshot isolation — under that definition MySQL doesn't provide it.* **Big advantage: it doesn't require application code to use special features, so it's LESS ERROR-PRONE** |
| **4** | **Conditional writes (compare-and-set)** | For databases without transactions. `UPDATE wiki_pages SET content = 'new' WHERE id = 1234 AND content = 'old'` — **no effect if changed; check whether the update took effect and retry.** Better: **a VERSION NUMBER column incremented on every update** — **optimistic locking** |
| **5** | **Conflict resolution / replication** | See below |
> ⚠️ **A subtle MVCC interaction with conditional writes:** if another transaction concurrently modified `content`, **the new content may not be visible under the MVCC visibility rules. Many MVCC implementations have an EXCEPTION: values written by other transactions ARE visible to the evaluation of the `WHERE` clause of `UPDATE` and `DELETE` queries, even though those writes are not otherwise visible in the snapshot.**
**Replication changes everything:**
> **Locks and conditional writes ASSUME THERE IS A SINGLE UP-TO-DATE COPY OF THE DATA. Multi-leader and leaderless databases usually allow several writes concurrently and replicate asynchronously, so they CANNOT GUARANTEE a single up-to-date copy. Thus techniques based on locks or conditional writes DO NOT APPLY.**
>
> Instead: **allow concurrent writes to create SIBLINGS and merge them afterward.** Merging prevents lost updates **if the updates are COMMUTATIVE** — incrementing a counter, adding to a set. **That's the idea behind CRDTs.**
>
> **However, some operations (such as conditional writes) CANNOT be made commutative. And LWW — the default in many replicated databases — IS PRONE TO LOST UPDATES.**
#### 3.4 Write Skew and Phantoms [#34-write-skew-and-phantoms]
**The on-call doctors example:**
> **WRITE SKEW: neither a dirty write nor a lost update, because the two transactions are UPDATING TWO DIFFERENT OBJECTS.** It's definitely a race condition: **if they had run one after another, the second doctor would have been prevented from going off call.**
>
> **Write skew is a GENERALIZATION of the lost-update problem: it can occur if two transactions READ THE SAME OBJECTS and then UPDATE SOME OF THOSE OBJECTS (different transactions may update different objects). In the special case where they update the SAME object, you get a dirty write or lost update instead.**
**Your options are far more restricted than for lost updates:**
| Option | Verdict |
| ------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Atomic single-object operations | ✗ **Don't help — multiple objects are involved** |
| Automatic lost-update detection | ✗ **Doesn't help — write skew is NOT automatically detected in PostgreSQL's repeatable read, MySQL/InnoDB's repeatable read, Oracle's serializable, or SQL Server's snapshot isolation** |
| Database constraints | ✗ Mostly — **you'd need a constraint involving MULTIPLE OBJECTS. Most databases don't have built-in support**, though **triggers or materialized views** may work |
| **Explicit `SELECT … FOR UPDATE` on the rows the transaction depends on** | ✔ **The second-best option** if you can't use serializable |
| **True serializable isolation** | ✔ **The only automatic prevention** |
##### Four more examples of write skew [#four-more-examples-of-write-skew]
| Example | The check | Why it fails |
| ------------------------------ | --------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Meeting room booking** | `SELECT COUNT(*) FROM bookings WHERE room_id=123 AND end_time > … AND start_time < …` then `INSERT` | **Snapshot isolation does not prevent another user from concurrently inserting a conflicting meeting** |
| **Multiplayer game** | A lock prevents two players moving the SAME figure | **The lock doesn't prevent players moving TWO DIFFERENT figures to the SAME POSITION**, or other rule-violating moves. Sometimes a uniqueness constraint helps; **otherwise you're vulnerable** |
| **Claiming a username** | Check whether the name is taken, then create the account | Not safe under SI — **but a UNIQUENESS CONSTRAINT is a simple solution here** (the second transaction aborts) |
| **Preventing double-spending** | Insert a tentative spending item, list all items, check the sum is positive | **Two spending items inserted concurrently can together make the balance negative, with neither transaction noticing the other** |
##### Phantoms — the underlying mechanism [#phantoms--the-underlying-mechanism]
**The universal pattern:**
> **In the doctors example, the row modified in ③ was ONE OF THE ROWS RETURNED IN ①, so `SELECT FOR UPDATE` can lock it and make the transaction safe.**
>
> **The other four examples are DIFFERENT: they check for the ABSENCE of rows matching a condition, and the write ADDS a row matching that same condition. IF THE QUERY IN ① DOESN'T RETURN ANY ROWS, `SELECT FOR UPDATE` CAN'T ATTACH LOCKS TO ANYTHING.**
>
> **A PHANTOM is a write in one transaction that changes the result of a search query in another transaction. Snapshot isolation avoids phantoms in READ-ONLY queries, but in read/write transactions phantoms lead to particularly tricky cases of write skew.** *(SQL generated by ORMs is also prone to write skew.)*
##### Materializing conflicts [#materializing-conflicts]
**If there's no object to lock, artificially introduce one.**
> For the meeting rooms: **create a table of time slots and rooms, one row per room per 15-minute period, for all combinations ahead of time (e.g. the next six months).** A transaction locks (`SELECT FOR UPDATE`) the rows for the desired room and period, then checks and inserts.
>
> **The additional table isn't used to store booking information — it's PURELY A COLLECTION OF LOCKS.**
> ⚠️ **It can be hard and error-prone to figure out how to materialize conflicts, and IT'S UGLY TO LET A CONCURRENCY CONTROL MECHANISM LEAK INTO THE APPLICATION DATA MODEL. Materializing conflicts should be considered A LAST RESORT. A serializable isolation level is preferable in most cases.**
***
# 8.10 Worked examples (/docs/ddia/transactions/worked-examples)
**① Which anomaly is this?** Classify each:
| Scenario | Anomaly |
| ---------------------------------------------------------- | ---------------------------------- |
| T1 writes x=3 (uncommitted); T2 reads x=3 | **Dirty read** |
| T1 writes x (uncommitted); T2 overwrites x | **Dirty write** |
| T2 reads x=1, then later in the same transaction reads x=2 | **Read skew / nonrepeatable read** |
| T1 and T2 both read balance=100, both write 90 | **Lost update** |
| T1 counts 2 doctors, T2 counts 2 doctors, both remove one | **Write skew** |
| T1 counts 0 bookings, T2 counts 0 bookings, both insert | **Write skew via a phantom** |
**② Why SI can't stop write skew, precisely.** Under SI, both transactions read from snapshots taken before either wrote. Neither writes an object the other *read-and-wrote*, so there is **no write-write conflict** to detect. The conflict is **read-write**: T1 read a set that T2's write changed. **Only a mechanism that tracks reads — predicate/index-range locks (2PL) or SIREAD tracking (SSI) — can see it.**
**③ SSI abort probability.** N transactions/s each holding read-write conflict potential over a shared hot key, with average transaction duration D seconds. Roughly, the expected number of concurrent conflicting transactions is `N × D`. If `N=1000` and `D=5 ms`, that's **5 concurrent** — modest abort rate. If a code change makes `D=200 ms`, it's **200 concurrent**, and the abort rate goes superlinear. **This is why SSI requires short read/write transactions**, and why "just add a slow external API call inside the transaction" is catastrophic.
**④ 2PC in-doubt cost.** Coordinator crashes with 50 in-doubt transactions, each holding exclusive locks on \~20 rows. Restart takes 20 minutes. **1,000 rows are unwritable for 20 minutes**, and anything queuing on them backs up. If the coordinator's log is lost: **1,000 rows locked indefinitely, surviving database restarts, until an administrator manually resolves 50 transactions by inspecting every participant.**
**⑤ Serial execution throughput budget.** Single core, 3 GHz. A stored procedure doing 5 in-memory index lookups + 3 writes ≈ 50 µs → **\~20,000 txn/s per core**. Shard across 16 cores → **320,000 txn/s** — *if* every transaction is single-shard. Introduce 5% cross-shard transactions at 1,000/s ceiling: those 5% now cap total throughput at **20,000 txn/s** overall. **A 5% cross-shard rate costs you 94% of your throughput.**
**⑥ Exactly-once without 2PC — the four crash points.** Verify §5.6 covers every window:
**No window produces a double side effect**, and no window loses the message. This is why §5.6 is the right default.
***
# 9.10 Decision cheat sheet (/docs/ddia/trouble-distributed-systems/decision-cheat-sheet)
**How do I set a timeout?**
There is no correct constant. **Measure the RTT distribution across many machines over an extended period**, pick a target trade-off between detection delay and false-positive rate, and prefer an **adaptive detector (Phi Accrual)** over a constant. Then **tune the threshold per decision by the cost of being wrong** — evicting from a read pool is cheap; failing over a leader is not.
**Which clock do I use?**
**Monotonic for every duration** — timeouts, latency measurement, rate limiting, retry backoff. **Time-of-day only for human-facing points in time.** **Never order events across machines by wall clock.** For ordering, use **logical clocks** (Ch 10) or **version vectors** (Ch 6).
**Can I trust LWW?**
Only if **you never update existing records** (Ch 6 §3.4). Otherwise it is a data-loss mechanism whose loss rate is proportional to your clock skew — and **you will not get an error.**
**Do I need fencing?**
**If two holders of the "same" lease could corrupt data or lose writes — yes, always.** A lease alone is not mutual exclusion; **the resource must reject stale tokens.** If the wasted work from a double execution is merely inefficient (the third example in §5.2), you can skip it.
**Do I need Byzantine fault tolerance?**
**Almost certainly not.** Datacenter nodes are yours; multitenancy is handled by isolation, not BFT; and **BFT cannot protect against a shared software bug or a compromise that reaches all nodes.** Do add the **cheap "weak lying" defences**: application-level checksums, input sanitization and size limits, and multiple NTP servers with outlier rejection.
**What system model should I assume?**
**Partially synchronous + crash-recovery**, and design so that **safety properties hold unconditionally** while **liveness may be conditioned on a majority surviving and the network eventually recovering.** Then add explicit handling for **fail-slow** nodes, because the model doesn't cover them and they're the hardest case in practice.
**Should I go distributed at all?**
**"Distributed systems engineers will often regard a problem as trivial if it can be solved on a single computer — and indeed a single computer can do a lot nowadays. If you can avoid opening Pandora's box and simply keep things on a single machine, IT IS GENERALLY WORTH DOING SO."** Go distributed for **fault tolerance and low latency**, which a single node cannot provide — not reflexively for scale.
**How do I gain confidence in correctness?**
**Combine all three:** a **TLA+ spec** for the protocol, **Jepsen** against the real deployment, and **DST** if your architecture can support it. And **fix what they find.**
***
# 9.1 Faults and Partial Failures (/docs/ddia/trouble-distributed-systems/faults-partial-failures)
**The single computer is a lie we've agreed to believe:**
> **There is no fundamental reason software on a single computer should be flaky. When hardware works correctly, the same operation always produces the same result (it is DETERMINISTIC). If there is a hardware problem, the consequence is usually a TOTAL system failure — kernel panic, blue screen, failure to start.** An individual computer with good software is **either fully functional or entirely broken, but not something in between.**
>
> **This is a DELIBERATE CHOICE in the design of computers. If an internal fault occurs, we prefer a computer to CRASH COMPLETELY rather than returning a wrong result, because WRONG RESULTS ARE DIFFICULT AND CONFUSING TO DEAL WITH. Thus computers HIDE the fuzzy physical reality on which they are implemented and present an IDEALIZED SYSTEM MODEL that operates with mathematical perfection.**
>
> *(As Ch 2 showed, **this is not actually true** — data does get silently corrupted and CPUs do sometimes silently return the wrong result — **but it happens rarely enough that we can get away with ignoring it.**)*
**Across a network, that abstraction collapses.**
> *"In my limited experience I've dealt with long-lived network partitions in a single data center, PDU failures, switch failures, accidental power cycles of whole racks, whole-DC backbone failures, whole-DC power failures, and **a hypoglycemic driver smashing his Ford pickup truck into a DC's HVAC system**. And I'm not even an ops guy."* — Coda Hale
**PARTIAL FAILURE:** some parts of the system are broken in an unpredictable way while other parts work fine.
> **The difficulty is that partial failures are NONDETERMINISTIC: if you try to do anything involving multiple nodes and the network, IT MAY SOMETIMES WORK AND SOMETIMES UNPREDICTABLY FAIL. You may not even know whether something succeeded.**
**The compensation:** if a system CAN tolerate partial failures, that opens powerful possibilities — **rolling upgrades**, rebooting one node at a time while the system keeps working.
> **Fault tolerance therefore allows us to make distributed systems MORE RELIABLE THAN SINGLE-NODE SYSTEMS; we can build a reliable system from unreliable components.**
***
# 9.7 Formal Methods and Randomized Testing (/docs/ddia/trouble-distributed-systems/formal-methods-randomized-testing)
> **Because of concurrency, partial failures, and network delays, THERE ARE A HUGE NUMBER OF POTENTIAL STATES. We need to guarantee that the properties hold in EVERY possible state and that we haven't forgotten any edge cases.**
#### 7.1 Model checking [#71-model-checking]
> **Model checkers verify that INVARIANTS HOLD ACROSS ALL OF AN ALGORITHM'S STATES by systematically trying all the things that could happen.** Specifications are written in a purpose-built language (**TLA+, Gallina, FizzBee**) that lets you **focus on behavior without worrying about implementation details.**
**The honest limitations:**
> **Model checking CAN'T ACTUALLY PROVE that invariants hold for every possible state, since most real-world algorithms have AN INFINITE STATE SPACE.** True verification would require a **formal proof**, which is typically **more difficult than running a model checker.** Instead, model checkers **encourage you to REDUCE the model to an approximation that can be fully verified, or to LIMIT the execution to an upper bound** (e.g. a maximum number of messages). **Any bugs occurring only with longer executions would then NOT BE FOUND.**
>
> **They also DON'T RUN YOUR ACTUAL CODE — which makes state-space exploration tractable but RISKS THAT YOUR SPECIFICATION AND YOUR IMPLEMENTATION GO OUT OF SYNC.**
**Track record:** **CockroachDB, TiDB, Kafka** and many others use model specifications. **Using TLA+, researchers demonstrated the potential for DATA LOSS in viewstamped replication caused by AMBIGUITY IN THE PROSE DESCRIPTION of the algorithm.**
#### 7.2 Fault injection [#72-fault-injection]
**Inject faults into a running system's environment and see how it behaves:** network failures, machine crashes, disk corruption, paused processes.
**Mechanics:** deploy the system alongside **fault injection COORDINATORS** (deciding what faults to execute and when) and **SCRIPTS** (injecting failures into individual nodes). Tools: **`kill` to pause or kill a process, `umount` to unmount a disk, firewall settings to disrupt network connections.**
> **The myriad of tools required make fault injection tests CUMBERSOME TO WRITE. It's common to adopt a framework like JEPSEN, which comes with integrations for various operating systems and many prebuilt fault injectors. JEPSEN HAS BEEN REMARKABLY EFFECTIVE AT FINDING CRITICAL BUGS IN MANY WIDELY USED SYSTEMS.**
*(Production fault injection = **chaos engineering**, popularized by **Netflix's Chaos Monkey**.)*
#### 7.3 Deterministic simulation testing (DST) [#73-deterministic-simulation-testing-dst]
> **DST uses a similar state-space exploration process to a model checker, BUT IT TESTS YOUR ACTUAL CODE, NOT A MODEL.**
>
> **Network communication, I/O, and clock timing are all replaced with MOCKS that allow the simulator to CONTROL THE EXACT ORDER in which things happen. This lets it explore MANY MORE SITUATIONS than handwritten tests or fault injection.**
>
> ### **If a test fails, IT CAN BE RERUN, since the simulator knows the exact order of operations that triggered the failure — in contrast to fault injection, which does NOT have such fine-grained control.** [#if-a-test-fails-it-can-be-rerun-since-the-simulator-knows-the-exact-order-of-operations-that-triggered-the-failure--in-contrast-to-fault-injection-which-does-not-have-such-fine-grained-control]
**Three strategies for making code deterministic:**
| Level | How | Examples |
| --------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------- |
| **Application** | **Built from the ground up to execute deterministically.** FoundationDB uses **Flow**, an async communication library providing **a point to inject a deterministic network simulation**. TigerBeetle models state as **a state machine with all mutations in a SINGLE EVENT LOOP**, plus mock deterministic clocks | FoundationDB, TigerBeetle |
| **Runtime** | **A single-threaded runtime forces all asynchronous code to run sequentially.** FrostDB **patches Go's runtime to execute goroutines sequentially.** Rust's **MadSim** provides deterministic implementations of Tokio's async API, Amazon's S3 library, Kafka's Rust library — **applications swap in deterministic libraries WITHOUT CHANGING THEIR CODE** | FrostDB, MadSim |
| **Machine** | **A custom HYPERVISOR replaces normally nondeterministic operations with deterministic ones — everything from clocks to network and storage.** Then developers **run their entire distributed system in containers within the hypervisor and get a COMPLETELY DETERMINISTIC DISTRIBUTED SYSTEM** | Antithesis |
**Two bonus advantages beyond replayability:**
* **Antithesis BRANCHES a test execution into multiple subexecutions when it discovers less common behavior**, exploring many paths
* **Mocked clocks let tests run FASTER THAN WALL CLOCK TIME.** TigerBeetle's time abstraction **simulates network latency and timeouts WITHOUT ACTUALLY TAKING THE FULL LENGTH OF TIME** — exploring more code paths faster
#### 7.4 The power of determinism — the book's synthesis [#74-the-power-of-determinism--the-books-synthesis]
> **NONDETERMINISM IS AT THE CORE OF ALL THE DISTRIBUTED SYSTEMS CHALLENGES in this chapter: concurrency, network delay, process pauses, clock jumps, and crashes all happen in unpredictable ways that vary from one run to the next.**
>
> **Conversely, IF YOU CAN MAKE A SYSTEM DETERMINISTIC, THAT CAN HUGELY SIMPLIFY THINGS. Making things deterministic is a simple but powerful idea that ARISES AGAIN AND AGAIN in distributed system design.**
> ⚠️ **But making code FULLY deterministic requires care. Even once you have removed all concurrency and replaced I/O, network, clocks, and RNGs with deterministic simulations, elements of nondeterminism may remain: in some languages, THE ORDER IN WHICH YOU ITERATE OVER A HASH TABLE may be nondeterministic. WHETHER YOU RUN INTO A RESOURCE LIMIT (memory allocation failure, stack overflow) is also nondeterministic.**
***
# 9.14 Forward links (/docs/ddia/trouble-distributed-systems/forward-links)
| Concept here | Where it's developed |
| ------------------------------------------------------ | ------------------------------------- |
| How to actually reach agreement despite all this | **Ch 10** — Consistency and Consensus |
| Logical clocks and ID generators | **Ch 10** |
| Linearizability and its cost | **Ch 10** |
| Ballot/term numbers as fencing tokens inside consensus | **Ch 10** |
| Coordination services (ZooKeeper, etcd) | **Ch 10**, **Ch 7** |
| State machine replication / shared logs | **Ch 10** |
| Idempotence and exactly-once processing | **Ch 12** — Stream Processing |
| Why quorums can still return stale data | **Ch 6** — Leaderless Replication |
| Split brain and failover hazards | **Ch 6** — Handling Node Outages |
| Timeouts, retries, and metastable failure | **Ch 2** — Nonfunctional Requirements |
# 9. The Trouble with Distributed Systems (/docs/ddia/trouble-distributed-systems)
> "They're funny things, Accidents. You never have them till you're having them." — A.A. Milne
**The mindset shift this chapter demands:**
> **If you want your system to be reliable in the presence of faults, you have to RADICALLY CHANGE YOUR MINDSET and focus on what could go wrong, even though it may be unlikely. It doesn't matter whether there is only a one-in-a-million chance; IN A LARGE ENOUGH SYSTEM, ONE-IN-A-MILLION EVENTS HAPPEN EVERY DAY. Experienced systems operators will tell you that ANYTHING THAT CAN GO WRONG WILL GO WRONG.**
>
> **In distributed systems, SUSPICION, PESSIMISM, AND PARANOIA PAY OFF.**
**The chapter is deliberately all problems. Ch 10 is the solutions.**
***
# 9.5 Knowledge, Truth, and Lies (/docs/ddia/trouble-distributed-systems/knowledge-truth-lies)
> **A node in the network CANNOT KNOW ANYTHING FOR SURE about other nodes — it can only make guesses based on the messages it receives (or doesn't receive). A node can find out another node's state only by exchanging messages with it. If a remote node doesn't respond, THERE IS NO WAY OF KNOWING ITS STATE, because PROBLEMS IN THE NETWORK CANNOT RELIABLY BE DISTINGUISHED FROM PROBLEMS AT A NODE.**
>
> *"Discussions of these systems border on the philosophical: What do we know to be true or false in our system? How sure can we be of that knowledge, if the mechanisms for perception and measurement are unreliable?"*
>
> **Fortunately, we don't need to figure out the meaning of life. We can STATE THE ASSUMPTIONS we are making (the SYSTEM MODEL) and design the system to meet those assumptions. Algorithms can be PROVED to function correctly within a certain system model — meaning RELIABLE BEHAVIOR IS ACHIEVABLE EVEN IF THE UNDERLYING MODEL PROVIDES VERY FEW GUARANTEES.**
#### 5.1 The majority rules [#51-the-majority-rules]
**Three parables:**
> ### **The moral: A NODE CANNOT NECESSARILY TRUST ITS OWN JUDGMENT OF A SITUATION.** [#the-moral-a-node-cannot-necessarily-trust-its-own-judgment-of-a-situation]
>
> **A distributed system cannot exclusively rely on a single node, because a node may fail at any time, potentially leaving the system stuck. Instead, many algorithms rely on a QUORUM: decisions require a minimum number of votes from several nodes.**
>
> **That includes decisions about declaring nodes dead. IF A QUORUM OF NODES DECLARES ANOTHER NODE DEAD, THEN IT MUST BE CONSIDERED DEAD, EVEN IF THAT NODE STILL VERY MUCH FEELS ALIVE. The individual node MUST ABIDE BY THE QUORUM DECISION AND STEP DOWN.**
**Why a majority:** **allows the system to continue if a MINORITY are faulty (3 nodes → tolerate 1; 5 nodes → tolerate 2), and it is SAFE because THERE CAN BE ONLY ONE MAJORITY — there cannot be two majorities with conflicting decisions at the same time.**
#### 5.2 Distributed locks and leases [#52-distributed-locks-and-leases]
> **Locks and leases in distributed applications are PRONE TO MISUSE and are A COMMON SOURCE OF BUGS.**
**Where you need "only one of some thing":**
| Use | Consequence of two holders |
| ----------------------------------------- | ------------------------------------------------------------ |
| **Only one node is leader for a shard** | **Split brain — lost or corrupted data. SERIOUS** |
| **Only one client updates a resource** | **Corruption by concurrent writes. SERIOUS** |
| **Only one node processes an input file** | **Only some wasted computational resources. NOT A BIG DEAL** |
**Two ways it goes wrong — and note the second has NO pause at all:**
##### Fencing off zombies [#fencing-off-zombies]
> **A ZOMBIE is a former leaseholder that has not yet found out that it lost the lease and is still acting as if it were the current leaseholder. SINCE WE CANNOT RULE OUT ZOMBIES ENTIRELY, we have to instead ensure THEY CAN'T DO ANY DAMAGE. This is called FENCING OFF the zombie.**
**The bad approach — STONITH** ("shoot the other node in the head"): disconnect it from the network, shut down the VM, or physically power down the machine.
> **It is NOT PARTICULARLY EFFECTIVE: it does not protect against the LARGE NETWORK DELAYS of case ②; ALL THE NODES COULD SHUT ONE ANOTHER DOWN; and by the time a zombie has been detected and shut down, IT MAY BE TOO LATE AND DATA MAY ALREADY BE CORRUPTED.**
**The good approach — FENCING TOKENS:**
**Fencing tokens under other names:**
| System | Name |
| ---------------------------------- | --------------------------------------------------------------------------- |
| **Chubby** (Google's lock service) | **sequencers** |
| **Kafka** | **epoch numbers** |
| **Paxos** | **ballot number** |
| **Raft** | **term number** |
| **ZooKeeper** | the transaction ID **`zxid`** or node version **`cversion`** |
| **etcd** | the **revision number** along with the **lease ID** |
| **Hazelcast** | **`FencedLock`** API generates one explicitly |
> **Fencing is similar to OPTIMISTIC CONCURRENCY CONTROL (Ch 8) — except that FENCING IS PERMANENT, while concurrency control failures can be retried.**
**What the storage service must support:** either a check for an outdated token, **or simply a write that succeeds only if the object has not been written by another client since the current client last read it — an atomic CAS.** **Object stores support this: S3 calls it CONDITIONAL WRITES, Azure Blob Storage CONDITIONAL HEADERS, Google Cloud Storage REQUEST PRECONDITIONS.**
**The subtle point about needing a lock service at all:**
> **If your clients write to only ONE storage service that supports conditional writes, THE LOCK SERVICE IS SOMEWHAT REDUNDANT — the lease assignment could have been implemented directly on that storage service. However, ONCE YOU HAVE A FENCING TOKEN, YOU CAN USE IT WITH MULTIPLE SERVICES OR REPLICAS and ensure the old leaseholder is fenced off ON ALL OF THEM.**
**Fencing a leaderless replicated store:**
#### 5.3 Byzantine faults [#53-byzantine-faults]
> **Fencing tokens can detect and block a node acting in error INADVERTENTLY. However, if the node DELIBERATELY wanted to subvert the system's guarantees, it could easily do so by SENDING MESSAGES WITH A FAKE FENCING TOKEN.**
>
> **In this book we assume that nodes are UNRELIABLE BUT HONEST. They may be slow or never respond, and their state may be outdated — but we assume that IF A NODE DOES RESPOND, IT IS TELLING THE "TRUTH."**
>
> **A BYZANTINE FAULT is a node "lying" — sending arbitrary faulty or corrupted responses, e.g. casting MULTIPLE CONTRADICTORY VOTES IN THE SAME ELECTION.**
*(Etymology: it generalizes the **two generals problem**. In the Byzantine version, **n generals need to agree, hampered by TRAITORS in their midst; it is not known in advance who the traitors are.** The name comes from **"byzantine" in the sense of excessively complicated, bureaucratic, devious** — used in politics long before computers. **Lamport wanted a nationality that would not offend readers, and was advised that calling it The Albanian Generals Problem was not such a good idea.**)*
**Where Byzantine fault tolerance IS relevant:**
* **Aerospace: data in memory or a CPU register could be corrupted by RADIATION, leading a node to respond in arbitrarily unpredictable ways.** Since failure is very expensive (an aircraft crashing, a rocket colliding with the ISS), **flight control systems must tolerate Byzantine faults**
* **Multiple mutually untrusting parties: cryptocurrencies and blockchains are a way of getting mutually untrusting parties to agree on whether a transaction happened, WITHOUT RELYING ON A CENTRAL AUTHORITY**
**Where it is NOT, and why:**
> **In a datacenter, all nodes are controlled by your organization (so they can hopefully be trusted), and RADIATION LEVELS ARE LOW ENOUGH that memory corruption is not a major problem.** *(Although datacenters in orbit are being considered.)* **Multitenant systems have mutually untrusting tenants, but they are ISOLATED VIA FIREWALLS, VIRTUALIZATION, AND ACCESS CONTROL POLICIES, not Byzantine fault tolerance. BFT protocols are QUITE EXPENSIVE — in most server-side data systems, THE COST MAKES THEM IMPRACTICABLE.**
**Two things BFT explicitly CANNOT save you from:**
1. **A software bug.** *"A bug could be regarded as a Byzantine fault, but IF YOU DEPLOY THE SAME SOFTWARE TO ALL NODES, THEN A BYZANTINE FAULT-TOLERANT ALGORITHM CANNOT SAVE YOU."* Most BFT algorithms need **a supermajority of more than two-thirds** functioning correctly — **to use this against bugs you would need FOUR INDEPENDENT IMPLEMENTATIONS of the same software and hope a given bug appears in only one.**
2. **Security compromise.** *"In most systems, IF AN ATTACKER CAN COMPROMISE ONE NODE, THEY CAN PROBABLY COMPROMISE ALL OF THEM, because the nodes are probably running the same software. Thus, TRADITIONAL MECHANISMS — authentication, access control, encryption, firewalls — CONTINUE TO BE THE MAIN PROTECTION."*
*(Web applications **do** need to expect arbitrary and malicious behavior from clients under end-user control — **which is why input validation, sanitization, and output escaping matter.** But we **don't use BFT protocols; we simply make the server THE AUTHORITY on what client behavior is allowed.** In **peer-to-peer networks with no central authority, BFT is more relevant.**)*
**Weak forms of lying — cheap, pragmatic, worth doing:**
| Guard | Against |
| --------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Application-level checksums** (or TLS) | **Corrupted packets that EVADE TCP/UDP checksums** — this does happen |
| **Input sanitization**: escaping to prevent SQL injection, range checks, **limiting string size to prevent denial of service through large memory allocations** | Malicious or malformed input. **An internal service behind a firewall may get away with less, but basic checks in protocol parsers are still a good idea** |
| **Multiple NTP servers**, estimating errors and checking that a **majority agree on a time range** | **A misconfigured NTP server reporting incorrect time is detected as an OUTLIER and excluded** |
***
# 9.4 Process Pauses (/docs/ddia/trouble-distributed-systems/process-pauses)
**The broken leader-lease loop:**
```java
while (true) {
request = getIncomingRequest();
// Ensure that the lease always has at least 10 seconds remaining
if (lease.expiryTimeMillis - System.currentTimeMillis() < 10000) {
lease = lease.renew();
}
if (lease.isValid()) {
process(request); // ← what if the thread pauses for 15 s HERE?
}
}
```
**Two bugs:**
1. **It relies on SYNCHRONIZED CLOCKS:** the expiry time was set by a **different machine** and is compared to the **local system clock.** **If clocks are out of sync by more than a few seconds, the code will start doing strange things.**
2. **Even using only the local monotonic clock:** the code assumes **very little time passes between checking the time and processing the request.** **If the thread stops for 15 seconds around `lease.isValid`, the lease will likely have expired by the time the request is processed, and another node will already have taken over. THERE IS NOTHING TO TELL THIS THREAD IT WAS PAUSED, so it won't notice until the next loop iteration — BY WHICH TIME IT MAY HAVE ALREADY DONE SOMETHING UNSAFE.**
#### 4.1 Nine reasons a thread pauses for a long time [#41-nine-reasons-a-thread-pauses-for-a-long-time]
> **The problem is similar to making multithreaded code on a single machine thread-safe: YOU CAN'T ASSUME ANYTHING ABOUT TIMING. But the tools we have for that — mutexes, semaphores, atomic counters, lock-free data structures, blocking queues — DON'T DIRECTLY TRANSLATE, because a distributed system has NO SHARED MEMORY, ONLY MESSAGES SENT OVER AN UNRELIABLE NETWORK.**
>
> ### **A node must assume its execution can be paused for a significant length of time AT ANY POINT, EVEN IN THE MIDDLE OF A FUNCTION. During the pause, the rest of the world keeps moving and may even declare the paused node dead. Eventually the node may continue running, WITHOUT EVEN NOTICING THAT IT WAS ASLEEP until it checks its clock sometime later.** [#a-node-must-assume-its-execution-can-be-paused-for-a-significant-length-of-time-at-any-point-even-in-the-middle-of-a-function-during-the-pause-the-rest-of-the-world-keeps-moving-and-may-even-declare-the-paused-node-dead-eventually-the-node-may-continue-running-without-even-noticing-that-it-was-asleep-until-it-checks-its-clock-sometime-later]
#### 4.2 Could we eliminate pauses? — hard real-time systems [#42-could-we-eliminate-pauses--hard-real-time-systems]
**Where failure to respond by a deadline causes serious damage:** computers controlling **aircraft, rockets, robots, cars.** *"If your car's onboard sensors detect a crash, you wouldn't want the airbag release to be delayed because of an inopportune GC pause."*
> ⚠️ **In embedded systems, REAL-TIME means the system is carefully designed and tested to meet specified timing guarantees IN ALL CIRCUMSTANCES. This is in contrast to the vaguer use of "real-time" on the web** (servers pushing data to clients, stream processing).
**What real-time requires at every level:**
* **A real-time operating system (RTOS)** scheduling processes with **a guaranteed allocation of CPU time in specified intervals**
* **Library functions must document their WORST-CASE EXECUTION TIMES**
* **Dynamic memory allocation may be restricted or disallowed entirely** (real-time GCs exist, but the application must not give the collector too much work)
* **An enormous amount of testing and measurement**
> **This severely restricts the range of programming languages, libraries, and tools. Developing real-time systems is VERY EXPENSIVE, and they are most commonly used in safety-critical embedded devices.**
>
> **Also: "REAL-TIME" IS NOT THE SAME AS "HIGH-PERFORMANCE" — in fact, real-time systems may have LOWER THROUGHPUT, since they prioritize timely responses above all else.**
>
> **For most server-side data processing systems, real-time guarantees are simply NOT ECONOMICAL OR APPROPRIATE. Consequently, these systems must suffer the pauses and clock instability that come from operating in a non-real-time environment.**
#### 4.3 Limiting the impact of garbage collection [#43-limiting-the-impact-of-garbage-collection]
**GC has genuinely improved:** *"A properly tuned collector will now usually pause processes for NO MORE THAN A FEW MILLISECONDS."* Java offers **CMS, G1, ZGC, Epsilon, Shenandoah**, each optimized for different memory profiles; **Go offers a simpler concurrent mark-and-sweep collector that attempts to optimize itself.**
**Four mitigation strategies:**
| Strategy | Detail |
| ---------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Use a language without a GC** | **Swift** (automatic reference counting), **Rust** and **Mojo** (object lifetimes tracked via the type system, so the compiler determines how long memory must be allocated) |
| **Reduce garbage** | **Pool and reuse objects rather than discarding them; allocate data OFF-HEAP** |
| **Treat GC as a planned outage** ✔ | **If the runtime can warn the application that a GC pause is coming, the application STOPS SENDING NEW REQUESTS to that node, waits for outstanding requests to finish, and THEN performs the GC while no requests are in progress. This HIDES GC PAUSES FROM CLIENTS and reduces high percentiles of response time** |
| **Restart before a full GC** | **Use the GC only for SHORT-LIVED objects (fast to collect) and RESTART PROCESSES PERIODICALLY before they accumulate enough long-lived objects to require a full GC. One node at a time, with traffic shifted away first — as in a ROLLING UPGRADE** |
***
# 9.9 Production failure catalog for this chapter (/docs/ddia/trouble-distributed-systems/production-failure-catalog-chapter)
| Symptom | Underlying mechanism |
| ---------------------------------------------------------------- | ----------------------------------------------------------- |
| Request timed out; did it happen or not? | **Six indistinguishable cases** (§2.1) — you cannot know |
| Retry caused a duplicate charge/email | Timeout ambiguity + non-idempotent operation |
| TCP said "delivered," the work never happened | **ACK means the KERNEL received it**, not the application |
| Data duplicated after a reconnect | **TCP dedup applies to one connection only** |
| Redundant switches, still an outage | **Redundancy doesn't guard against human misconfiguration** |
| A can reach B, B can reach C, A can't reach C | **Partial/asymmetric network fault** |
| Node sends fine but receives nothing (or vice versa) | **One-directional link failure** |
| Cluster deadlocked and stayed broken after the network recovered | **Untested network-fault error handling** |
| Healthy node evicted during a traffic spike | Timeout too short; **detecting overload as death** |
| All nodes declared each other dead | **Cascading failure** from premature eviction |
| Latency fine at 60% load, chaotic at 90% | **Queueing delays explode near capacity** |
| Random slowness with no code change | **Noisy neighbour** in a multitenant environment |
| Timestamps out of order across nodes | **Clock skew**; ordering by wall clock |
| Writes silently disappear, no errors | **LWW + a node with a lagging clock** |
| Elapsed time computed as negative | **Time-of-day clock stepped backward** |
| Everything hung at midnight | **Leap second** |
| Clock drifted for weeks, nobody noticed | **NTP firewalled off** — the silent failure |
| A microsecond timestamp that's wrong by 40 ms | **No confidence interval exposed** |
| Leader kept writing after losing its lease | **Process pause** → zombie; no fencing |
| Ancient write arrives and corrupts a file | **Delayed packet from a crashed former leaseholder** |
| STONITH fired and both nodes died | **Mutual shutdown**; STONITH doesn't handle delayed packets |
| Node responds to health checks but does no work | **Fail-slow / gray failure / limping node** |
| Quorum algorithm violated after a disk wipe | **Node amnesia** breaks the stable-storage assumption |
| Verified model, buggy system | **Spec/implementation drift** |
| Simulation replay isn't reproducible | An **uncontrolled nondeterminism source** (hash order, OOM) |
***
# 9.12 Self-test (/docs/ddia/trouble-distributed-systems/self-test)
Why do computers deliberately crash rather than return wrong results? What does that design choice hide?
Define partial failure. Why is nondeterminism the thing that makes it hard?
List the six things that may have happened when a request times out. Which of them mean the operation already took effect?
Give four reasons TCP's "reliability" doesn't give you application-level reliability. What do you actually need?
Why doesn't redundant network hardware reduce faults as much as expected?
Give two examples of asymmetric or partial network faults from the chapter.
Enumerate the four points at which a packet can be queued, and say which one is invisible to both endpoints.
Why does a system near capacity have far worse delay variance than one with spare capacity?
Why is there no "correct" timeout value? What should you do instead of a constant?
Explain the cascading-failure loop caused by a timeout that's too short.
Contrast a telephone circuit with a TCP connection. Why did the internet choose packet switching?
"Variable delays are not a law of nature." Explain the cost/benefit trade-off in one sentence.
Give three properties of monotonic clocks and three of time-of-day clocks. Which is safe for measuring a timeout, and why?
Why do bad clocks cause *silent* data loss rather than visible crashes?
Walk through the LWW example where a causally later write gets an earlier timestamp. Name three separate problems with client-clock LWW.
Why can NTP never be accurate enough to guarantee correct event ordering?
What is a clock confidence interval? Why do most APIs not expose one, and what do TrueTime and ClockBound do differently?
Explain Spanner's commit-wait. Why are atomic clocks helpful but not strictly necessary?
Find both bugs in the lease-renewal loop. Which one survives switching to a monotonic clock?
List six causes of a multi-second process pause. Which two can occur without any code of yours running?
Why don't mutexes and semaphores translate to distributed systems?
What does "hard real-time" require at each layer, and why is real-time not the same as high-performance?
Describe the technique of treating a GC pause as a planned outage.
Tell the three "majority rules" parables. What is the moral, in one sentence?
Why can there only ever be one majority? What does that buy you?
Draw both distributed-lock failure modes. Which one involves no pause at all?
Why is STONITH insufficient? Give three reasons.
Explain fencing tokens. What must the *storage* side do, and what must a newly-elected leaseholder do immediately?
How do you fence a leaderless replicated store with LWW?
Define a Byzantine fault. Give two contexts where BFT is warranted and two things it cannot protect against.
Name three cheap defences against "weak lying."
Define the three timing models and the four node-failure models. Which combination is most useful, and which failure mode does it fail to capture?
Distinguish safety from liveness precisely (not "bad" vs "good"). Which one may be conditioned on caveats, and what caveats are standard?
Why does node amnesia break quorum correctness? What does that tell you about system models?
Compare model checking, fault injection, and DST on what each finds and misses. Why is DST's replayability so valuable?
List six places determinism has appeared in this book so far. Name two residual sources of nondeterminism even after mocking I/O and clocks.
you operate a sharded database with single-leader replication per shard, leases held via etcd, storage in S3, running on VMs in three availability zones. Enumerate every failure mode from this chapter that could produce *two nodes writing as leader for the same shard*, and specify the mechanism you'd use to make each one harmless. State explicitly which of your defences are safety properties and which are liveness properties.
# 9.6 System Model and Reality (/docs/ddia/trouble-distributed-systems/system-model-reality)
**A SYSTEM MODEL is an abstraction describing an algorithm's assumptions.**
#### 6.1 Timing models [#61-timing-models]
| Model | Assumption | Realism |
| --------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Synchronous** | **Bounded network delay, bounded process pauses, bounded clock error.** *(Not exactly synchronized clocks or zero delay — just that they NEVER EXCEED A FIXED UPPER BOUND)* | **NOT a realistic model of most practical systems, because unbounded delays and pauses DO occur** |
| **Partially synchronous** ✔ | **Behaves like a synchronous system MOST of the time, but SOMETIMES exceeds the bounds.** When it happens, delay, pauses, and clock error may become **arbitrarily large** | **A REALISTIC MODEL OF MANY SYSTEMS.** *"Most of the time, networks and processes are quite well behaved — otherwise we would never be able to get anything done — but we have to reckon with the fact that ANY TIMING ASSUMPTIONS MAY BE SHATTERED OCCASIONALLY"* |
| **Asynchronous** | **No timing assumptions at all — it does not even have a clock, so it cannot use timeouts** | Some algorithms can be designed for it, **but it is VERY RESTRICTIVE** |
#### 6.2 Node failure models [#62-node-failure-models]
| Model | Assumption |
| ------------------------------------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Crash-stop (fail-stop)** | **A node can fail in ONLY ONE WAY: by crashing. It suddenly stops responding and THEREAFTER IS GONE FOREVER — IT NEVER COMES BACK** |
| **Crash-recovery** ✔ | Nodes may crash at any moment and **perhaps start responding again after an unknown time.** Nodes have **STABLE STORAGE preserved across crashes**, while **in-memory state is lost** |
| **Degraded performance / partial functionality** | Nodes **SLOW DOWN. They may still respond to health checks while being TOO SLOW TO GET ANY REAL WORK DONE.** Called a **LIMPING NODE, GRAY FAILURE, or FAIL-SLOW — and it can be EVEN MORE DIFFICULT TO DEAL WITH THAN A CLEANLY FAILED NODE** |
| **Byzantine (arbitrary)** | **Nodes may do absolutely anything, including trying to trick and deceive other nodes** |
**Real causes of fail-slow, worth memorizing:** *"A Gigabit network interface could suddenly drop to **1 Kb/s** throughput because of a driver bug; a process under memory pressure may spend most of its time performing garbage collection; worn-out SSDs can have erratic performance; and hardware can be affected by **high temperature, loose connectors, mechanical vibration, power supply problems, firmware bugs.**"* Plus: **a process stops doing SOME of the things it is supposed to do while other aspects continue working — because a background thread has crashed or deadlocked.**
> ### **For modeling real systems, the PARTIALLY SYNCHRONOUS model with CRASH-RECOVERY faults is generally the most useful.** [#for-modeling-real-systems-the-partially-synchronous-model-with-crash-recovery-faults-is-generally-the-most-useful]
#### 6.3 Safety vs liveness [#63-safety-vs-liveness]
**Example properties for a fencing-token generator:**
| Property | Definition | Kind |
| ---------------------- | ------------------------------------------------------------------------------------------------- | ------------ |
| **Uniqueness** | **No two requests for a fencing token return the same value** | **SAFETY** |
| **Monotonic sequence** | **If request x returned tₓ and y returned t\_y, and x completed before y began, then tₓ \< t\_y** | **SAFETY** |
| **Availability** | **A node that requests a token and does not crash EVENTUALLY receives a response** | **LIVENESS** |
> **A giveaway: LIVENESS PROPERTIES OFTEN INCLUDE THE WORD "EVENTUALLY."** *(And yes — **eventual consistency is a liveness property.**)*
**The precise definitions:**
> **Why the distinction matters:** for distributed algorithms it is common to require **SAFETY PROPERTIES ALWAYS HOLD, IN ALL POSSIBLE SITUATIONS OF A SYSTEM MODEL. Even if all nodes crash, or the entire network fails, THE ALGORITHM MUST NEVER RETURN A WRONG RESULT.**
>
> **With LIVENESS properties we ARE allowed to make caveats — e.g. a request needs to receive a response only if a MAJORITY OF NODES HAVE NOT CRASHED, and only if THE NETWORK EVENTUALLY RECOVERS.** *(The partially synchronous model requires exactly this: any period of interruption lasts only a finite duration and is then repaired.)*
#### 6.4 Where the model meets reality [#64-where-the-model-meets-reality]
> **Algorithms in the crash-recovery model assume DATA IN STABLE STORAGE SURVIVES CRASHES. But what if the data on disk is CORRUPTED OR WIPED OUT by hardware error or misconfiguration? What if a server has a FIRMWARE BUG and FAILS TO RECOGNIZE ITS HARD DRIVES ON REBOOT even though they're correctly attached?**
>
> **Quorum algorithms rely on A NODE REMEMBERING THE DATA IT CLAIMS TO HAVE STORED. IF A NODE MAY SUFFER FROM AMNESIA and forget previously stored data, THAT BREAKS THE QUORUM CONDITION AND THUS BREAKS THE CORRECTNESS OF THE ALGORITHM.** Perhaps a new model is needed in which stable storage *mostly* survives — **but that model then becomes harder to reason about.**
> **A real implementation may still have to include code to handle the case of something happening that was ASSUMED TO BE IMPOSSIBLE — even if that handling boils down to `printf("Sucks to be you"); exit(666);` — that is, LETTING A HUMAN OPERATOR CLEAN UP THE MESS. (THIS IS ONE DIFFERENCE BETWEEN COMPUTER SCIENCE AND SOFTWARE ENGINEERING.)**
>
> **That is not to say abstract system models are worthless — QUITE THE OPPOSITE. They are incredibly helpful for DISTILLING DOWN the complexity of real systems to a manageable set of faults we can reason about.**
***
# 9.8 Technology deep dives (/docs/ddia/trouble-distributed-systems/technology-deep-dives)
***
#### 8.1 Failure detectors (timeouts, heartbeats, Phi Accrual) [#81-failure-detectors-timeouts-heartbeats-phi-accrual]
**Problem it solves.** Decide whether a node is dead, so a load balancer can remove it or a follower can be promoted — **without any reliable way to observe the difference between "dead," "slow," and "unreachable."**
**Why wasn't an explicit signal enough?** §2.4: RST/FIN, crash scripts, switch queries, and ICMP all exist and **none of them can be relied upon.** In general you get **no response at all.**
**Why wasn't a fixed timeout enough?** §2.5: there is **no correct value.** Too short → false positives → load transfer → cascading failure. Too long → slow recovery. And **the delay distribution is not stationary** — it depends on load, on noisy neighbours, on whether a switch is being upgraded.
**How Phi Accrual works internally.** Instead of a boolean "up/down," it maintains a **sliding window of recent heartbeat inter-arrival times**, fits a distribution (normal in the classic formulation), and computes
`φ = −log₁₀(P(heartbeat arrives later than the current elapsed time))`.
φ rises continuously as the silence lengthens. The application picks a **threshold**: φ ≥ 8 means "the probability we are wrong is about 10⁻⁸." **The value is that the detector adapts automatically to a network whose latency distribution changes** — exactly the §2.5 recommendation.
**Deployment.** Heartbeats between all peers (Cassandra's gossip) or from a leader; thresholds per-role — **be more conservative about leader failover than about removing a node from a read pool**, because the cost of a false positive differs by an order of magnitude.
**Monitoring.** False-positive rate (nodes marked dead that recover within seconds — **the single most useful signal that your threshold is too aggressive**); time-to-detection distribution; heartbeat inter-arrival p99; **flapping count per node**; the correlation between detection events and load spikes (if they correlate, you are detecting overload, not death).
**Scaling.** All-to-all heartbeating is O(n²); above a few hundred nodes you need gossip or a hierarchy.
**What actually breaks.**
* **Detecting overload as death**, then transferring load away, then detecting more overload — §2.5's cascading failure, and the reason Ch 7 warned against automatic rebalancing plus automatic failure detection.
* **A GC pause exceeding the timeout**, so a perfectly healthy leader is demoted mid-write.
* **Asymmetric faults** (§2.3): the node is alive and receiving, but its outbound path is broken, so it is declared dead while it keeps believing it is the leader — the §5.1 parable, and precisely the situation that requires fencing.
* **Health checks that only prove the process is running**, not that it can do work — so a **fail-slow / limping node** stays in rotation forever, which §6.2 notes is *harder* to handle than a clean crash.
* **Flapping** when the threshold sits right at the network's p99.
***
#### 8.2 NTP, PTP, and clock monitoring [#82-ntp-ptp-and-clock-monitoring]
**Problem it solves.** Give every machine roughly the same notion of wall-clock time, so timestamps in logs, expiries, and cross-machine reasoning are comparable.
**Why wasn't the hardware clock enough?** §3.2 ①: **up to 200 ppm drift**, temperature-dependent — **17 seconds per day** if never resynced.
**Why isn't NTP enough for ordering?** §3.4: **NTP's accuracy is bounded by network round-trip time**, so you cannot make clock error smaller than network delay, which is exactly what correct ordering would require.
**How it works internally.** NTP samples several servers, measures offset and round-trip delay for each, discards outliers (**the "weak lying" defence of §5.3**), and either **slews** (gradually adjusts the rate, ≤0.05%) or **steps** (jumps) the clock. **PTP** pushes accuracy to sub-microsecond by using **hardware timestamping in the NIC** and **transparent clocks in switches** that account for their own queueing delay — which is why PTP needs switch support and NTP doesn't.
**Deployment.** Stratum-1 sources (GPS/atomic) in each datacenter; **at least 4 upstream servers so outlier rejection works**; `chrony` over `ntpd` on modern Linux (faster convergence, better on VMs); **leap-second smearing configured consistently across the fleet** — a mixed fleet where half smear and half step is worse than either.
**Monitoring — and this is the section people skip.**
* **Clock offset per node vs the reference** — with an alert threshold well below any application-level assumption
* **Whether the NTP daemon is actually running and synchronized** (`chronyc tracking`: stratum, root dispersion, last offset). §3.2 ③: **a node firewalled off from NTP looks perfectly healthy while drifting**
* **Step (jump) events** — every one of these is a potential correctness incident
* **Root dispersion / estimated error** — this is your confidence interval, and §3.5 says most systems never look at it
* On VMs: **steal time**, since §3.2 ⑦ means the guest's own accuracy estimate can be wrong
**Backup.** Not applicable, but: **have a plan for GPS jamming** (§3.2), which is a real and locality-dependent risk.
**What actually breaks.**
* **Silent drift after a firewall change** — the canonical §3.3 failure: nothing errors, and data quietly disappears via LWW.
* **A leap second** hanging every JVM in the fleet simultaneously (the Ch 2 example) — a *correlated software fault*.
* **VM live migration** making the clock jump forward mid-transaction.
* **Half the fleet smearing and half stepping** during a leap second, so nodes disagree by a full second for a day.
* **Using `System.currentTimeMillis()` to measure a duration**, so a backward step produces a negative elapsed time — and code that divides by it.
* **Trusting microsecond digits** when root dispersion is 40 ms (§3.5).
***
#### 8.3 Distributed lock services (ZooKeeper, etcd, Chubby) and fencing [#83-distributed-lock-services-zookeeper-etcd-chubby-and-fencing]
**Problem it solves.** Make exactly one node the leader / lock holder, and — crucially — **make it safe when that guarantee is inevitably violated.**
**Why wasn't a lock in a database enough?** A lock with no lease expires never (the holder crashes and blocks forever); a lock with a lease can be held by two nodes at once (§5.2 cases ① and ②).
**Why isn't STONITH enough?** §5.2: it doesn't stop a delayed packet, all nodes can shoot each other, and it may act too late.
**How it works internally.** A consensus-replicated state machine (Ch 10). Locks are **ephemeral nodes / leases** tied to a session with a heartbeat; when the session lapses, the lock is released automatically. **Every state change carries a monotonically increasing version** — ZooKeeper's `zxid`/`cversion`, etcd's `revision` — which is what makes it a fencing-token generator, satisfying §6.3's *uniqueness* and *monotonic sequence* safety properties.
**The critical design rule:** **the lock service alone is not sufficient. The RESOURCE must check the token.** A lock service without token-checking on the storage side gives you the illusion of mutual exclusion and none of the substance.
**Deployment.** 3 or 5 nodes (odd, for majority quorums), **spread across availability zones but NOT across regions** — cross-region consensus latency makes every lock acquisition painful. Session timeout tuned above your worst realistic GC pause.
**Monitoring.** **Session expiry rate** (each one is a potential zombie); leader elections per hour (frequent elections = your timeouts fight your GC pauses); **request latency p99**; **fsync latency on the consensus log** — this bounds everything; watch-count and znode-count (ZooKeeper falls over on unbounded watches); **quorum health and whether any member is lagging**.
**Backup.** **etcd snapshots** / ZooKeeper transaction-log + snapshot backups. Losing the coordination store loses the shard map, the leader identity, and every lock — and Ch 7 §4 notes it's the authority for routing too.
**What actually breaks.**
* **A lock service used without fencing tokens.** The system looks correct in testing and corrupts data in production, exactly as HBase did.
* **A session timeout shorter than a GC pause** → the leader is demoted while alive → split brain until fencing catches it.
* **A client that acquires a lease and then does a long operation** without re-checking, so §5.2 case ① is guaranteed rather than merely possible.
* **`fsync` latency spikes on the consensus log** stalling all lock operations cluster-wide.
* **Cross-region deployment** turning every lock acquisition into a 150 ms round trip.
* **Treating the lock service as available**: it is CP, not AP — during a partition the minority side cannot acquire locks, by design.
***
#### 8.4 Jepsen, TLA+, and deterministic simulation (Antithesis, FoundationDB) [#84-jepsen-tla-and-deterministic-simulation-antithesis-foundationdb]
**Problem it solves.** Establish that an implementation actually satisfies its claimed safety properties under the faults of §§2–4 — which normal testing cannot do, because the bugs live in orderings you never happened to produce.
**Why wasn't unit/integration testing enough?** §7: the state space is enormous, and concurrency bugs **only manifest when you get unlucky with the timing** (Ch 8 §3).
**Why isn't one technique sufficient?**
| Technique | Finds | Misses |
| ---------------------------- | ------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------- |
| **TLA+ / model checking** | Design-level bugs, protocol ambiguities (the **viewstamped-replication data-loss** finding) | Anything where **the implementation diverges from the spec** |
| **Jepsen / fault injection** | Real bugs in the real system under real faults | **Not replayable**; coarse control; only explores orderings it happens to hit |
| **DST** | Real bugs in real code, **replayable**, systematic | Requires **controlling every source of nondeterminism** — an architectural commitment |
**How Jepsen works internally.** Runs the real cluster; a **generator** issues concurrent client operations while a **nemesis** injects faults (partition, clock skew, process kill, pause); every operation is recorded with invoke/complete times into a **history**; a **checker** (Knossos/Elle) then searches for a linearization or a serializable schedule consistent with that history — **if none exists, it produces a concrete counterexample.** Elle goes further and infers **transaction dependency cycles**, naming the exact Ch 8 anomaly observed.
**Monitoring/practice.** Run in CI on a schedule, not once before launch; **keep the failing histories** — they're the regression suite; treat a Jepsen finding as a *design* review trigger, not just a patch.
**What actually breaks.**
* **Spec/implementation drift** — a beautifully verified TLA+ model of code that no longer matches it (§7.1).
* **DST that misses a nondeterminism source** — hash iteration order, allocation failure (§7.4) — so a "deterministic" replay isn't.
* **Testing only clean crashes** and never latency, partial partitions, or clock skew — which are the faults that actually break systems.
* **Running chaos/fault injection and not fixing what it finds**, converting the practice into theatre.
***
# 9.13 Terminology introduced here (/docs/ddia/trouble-distributed-systems/terminology-introduced-here)
# 9.3 Unreliable Clocks (/docs/ddia/trouble-distributed-systems/unreliable-clocks)
**Eight questions applications ask of clocks — note the split:**
> **In a distributed system, time is a tricky business, because communication is not instantaneous. The time when a message is received is ALWAYS LATER than when it is sent — but because of variable delays, WE DON'T KNOW HOW MUCH LATER. This makes it difficult to determine THE ORDER in which things happened.**
>
> **Each machine has its own clock — usually a QUARTZ CRYSTAL OSCILLATOR — not perfectly accurate, so each machine has its own notion of time.** Synchronized via **NTP**, whose servers get time from a more accurate source such as a **GPS receiver**.
#### 3.1 The two clocks [#31-the-two-clocks]
| | **Time-of-day clock** | **Monotonic clock** |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------- |
| **API** | `clock_gettime(CLOCK_REALTIME)`, `System.currentTimeMillis` | `clock_gettime(CLOCK_MONOTONIC/BOOTTIME)`, `System.nanoTime` |
| **Returns** | **Seconds since the epoch** (midnight UTC 1 Jan 1970, Gregorian, **not counting leap seconds**) | **A meaningless absolute value** — perhaps nanoseconds since boot |
| **Guarantee** | Synchronized with NTP, so a timestamp on one machine ideally means the same as on another | **Guaranteed to ALWAYS MOVE FORWARD** |
| **Can jump?** | **YES — if the local clock is too far ahead of NTP, it may be FORCIBLY RESET and appear to JUMP BACK.** Also leap seconds, and DST (avoidable by always using UTC) | **No jumps. NTP may SLEW it** — adjust the rate at which it moves forward — **by up to 0.05% by default, but cannot make it jump** |
| **Use for** | **Points in time** | **DURATIONS — timeouts, response times.** *"More like a stopwatch than a wall clock"* |
| **Comparable across machines?** | Yes (approximately) | **NO — the values don't mean the same thing** |
| **Resolution** | Historically coarse (**10 ms steps on older Windows**); less of a problem now | **Usually quite good — microseconds or less** |
> ⚠️ **On a server with multiple CPU sockets, there may be a SEPARATE TIMER PER CPU, not necessarily synchronized with other CPUs.** The OS compensates and tries to present a monotonic view — **but it is wise to take this guarantee of monotonicity WITH A PINCH OF SALT.**
**In a distributed system, using a monotonic clock for elapsed time is usually fine, because it doesn't assume any synchronization between nodes.**
#### 3.2 Clock synchronization is worse than you think — eight failure modes [#32-clock-synchronization-is-worse-than-you-think--eight-failure-modes]
**High accuracy IS achievable if you invest:** **MiFID II requires high-frequency trading funds to synchronize within 100 MICROSECONDS of UTC**, to help debug flash crashes and detect market manipulation. Achieved with **GPS receivers and/or atomic clocks, PTP (Precision Time Protocol), and careful deployment and monitoring.**
> ⚠️ **Relying on GPS alone can be risky because GPS SIGNALS CAN EASILY BE JAMMED. In some locations (close to military facilities) this happens FREQUENTLY.**
#### 3.3 Why bad clocks are especially dangerous [#33-why-bad-clocks-are-especially-dangerous]
> **Part of the problem is that INCORRECT CLOCKS EASILY GO UNNOTICED. If a CPU is defective or the network misconfigured, it most likely WON'T WORK AT ALL, so the issue is quickly spotted. If the quartz clock is defective or NTP misconfigured, MOST THINGS WILL SEEM TO WORK FINE, even as the clock drifts further from reality.**
>
> ### **If software relies on an accurately synchronized clock, the result is more likely to be SILENT AND SUBTLE DATA LOSS than a dramatic crash.** [#if-software-relies-on-an-accurately-synchronized-clock-the-result-is-more-likely-to-be-silent-and-subtle-data-loss-than-a-dramatic-crash]
>
> **Therefore: CAREFULLY MONITOR THE CLOCK OFFSETS between all machines. Any node whose clock drifts too far from the others should be DECLARED DEAD AND REMOVED.**
#### 3.4 Timestamps for ordering events — where it bites [#34-timestamps-for-ordering-events--where-it-bites]
**Three serious problems with client-clock LWW** (Cassandra and ScyllaDB do this deliberately, to avoid the extra read round-trip needed to find the greatest existing timestamp):
1. **Database writes can mysteriously disappear.** **A node with a LAGGING clock is UNABLE TO OVERWRITE values previously written by a node with a FASTER clock UNTIL THE CLOCK SKEW HAS ELAPSED. This can cause ARBITRARY AMOUNTS OF DATA TO BE SILENTLY DROPPED WITHOUT ANY ERROR BEING REPORTED.**
2. **LWW cannot distinguish sequential-in-quick-succession from truly concurrent writes.** **Additional causality tracking (VERSION VECTORS, Ch 6) is needed.**
3. **Two nodes could independently generate writes with the SAME timestamp**, especially at millisecond resolution. **A tiebreaker (a large random number) is required — but this can ALSO lead to violations of causality.**
> **Even with tightly NTP-synchronized clocks, you could send a packet at timestamp 100 ms (sender's clock) and have it arrive at timestamp 99 ms (recipient's clock) — SO IT APPEARS THE PACKET ARRIVED BEFORE IT WAS SENT, WHICH IS IMPOSSIBLE.**
>
> **Could NTP be made accurate enough? PROBABLY NOT, because NTP's accuracy is itself limited by the NETWORK ROUND-TRIP TIME, plus quartz drift. To guarantee correct ordering you would need THE CLOCK ERROR TO BE SIGNIFICANTLY LOWER THAN THE NETWORK DELAY, WHICH IS NOT POSSIBLE.**
**The alternative: LOGICAL CLOCKS** — based on **incrementing counters rather than an oscillating quartz crystal.** They **do not measure the time of day or seconds elapsed, ONLY THE RELATIVE ORDERING of events.** Time-of-day and monotonic clocks, which measure actual elapsed time, are **PHYSICAL CLOCKS.** (Ch 10.)
#### 3.5 Clock readings with a confidence interval [#35-clock-readings-with-a-confidence-interval]
> **It doesn't make sense to think of a clock reading as A POINT IN TIME. It is more like A RANGE OF TIMES, within a CONFIDENCE INTERVAL — a system may be 95% confident the time is between 10.3 and 10.5 seconds past the minute.**
>
> **IF WE KNOW ONLY THE TIME ±100 ms, THE MICROSECOND DIGITS IN THE TIMESTAMP ARE ESSENTIALLY MEANINGLESS.**
**Computing the bound:** with a GPS receiver or atomic clock attached, **the error range is determined by the device and the signal quality.** From a server: **expected quartz drift since the last sync + the NTP server's uncertainty + the network round-trip time.**
> **Unfortunately, MOST SYSTEMS DON'T EXPOSE THIS UNCERTAINTY. When you call `clock_gettime`, the return value doesn't tell you the expected error — SO YOU DON'T KNOW WHETHER ITS CONFIDENCE INTERVAL IS FIVE MILLISECONDS OR FIVE YEARS.**
>
> **The exceptions: Google Spanner's TrueTime API and Amazon ClockBound, which return `[earliest, latest]`.**
#### 3.6 Synchronized clocks for global snapshots — Spanner's trick [#36-synchronized-clocks-for-global-snapshots--spanners-trick]
**The problem:** MVCC (Ch 8) requires a **monotonically increasing transaction ID.** On one node, a counter suffices. **Distributed across many machines, a global monotonically increasing ID is difficult, because it requires COORDINATION — and the ID must reflect CAUSALITY (if B reads or overwrites a value written by A, B must have a higher ID). With lots of small, rapid transactions, creating such IDs becomes AN UNTENABLE BOTTLENECK.**
**Spanner's observation:**
> **The atomic clocks and GPS receivers are NOT STRICTLY NECESSARY. THE IMPORTANT THING IS TO HAVE A CONFIDENCE INTERVAL — accurate clock sources only help keep that interval SMALL.**
*(**YugabyteDB can leverage ClockBound on AWS**, and several other systems now rely on clock synchronization to various degrees.)*
***
# 9.2 Unreliable Networks (/docs/ddia/trouble-distributed-systems/unreliable-networks)
**Shared-nothing systems: each machine has its own memory and disk, and one machine cannot access another's except by making requests over the network.** *(Even with shared object storage, machines communicate with it over the network.)*
**The internet and most datacenter networks (Ethernet) are ASYNCHRONOUS PACKET NETWORKS: one node can send a packet, but the network gives NO GUARANTEES as to WHEN it will arrive or WHETHER it will arrive at all.**
#### 2.1 The six things that could have happened [#21-the-six-things-that-could-have-happened]
> **The sender can't even tell whether the packet was delivered. The only option is for the recipient to send a response message — which may in turn be lost or delayed.**
>
> **The usual way of handling this is a TIMEOUT. However, when a timeout occurs, YOU STILL DON'T KNOW WHETHER THE REMOTE NODE GOT YOUR REQUEST** — and if the request is still queued somewhere, **it may still be delivered, even though you've given up.**
#### 2.2 Why TCP doesn't save you [#22-why-tcp-doesnt-save-you]
**What TCP genuinely provides:** detects and retransmits dropped packets; detects reordered packets and restores order; detects corruption via a **simple checksum**; and figures out how fast it can send — **congestion control / flow control / backpressure.**
**How a "send" actually works:**
**Four reasons TCP's "reliability" is not your reliability:**
1. **TCP decides a packet must have been lost if no ACK arrives within a timeout — but it CAN'T TELL whether the outbound packet or the ACK was lost.** It can resend, but **can't guarantee the new packet gets through** (*if the network cable is unplugged, TCP can't plug it back in for you*). **Eventually it gives up and signals an error.**
2. **TCP's deduplication and retransmission apply only to A SINGLE CONNECTION.** If the application reconnects and retransmits, **data could be duplicated.**
3. **If a TCP connection closes with an error, you have NO WAY OF KNOWING HOW MUCH DATA WAS ACTUALLY PROCESSED by the remote node.**
4. **Even if you receive an ACK, that means only that the OS KERNEL on the remote node received it; THE APPLICATION MAY HAVE CRASHED BEFORE HANDLING THAT DATA.**
> **If you want to be sure a request was successful, you need A POSITIVE RESPONSE FROM THE APPLICATION ITSELF.**
*(Most of this applies equally to **QUIC**, **SCTP** (WebRTC), and **BitTorrent uTP**.)*
#### 2.3 Network faults in practice — the numbers [#23-network-faults-in-practice--the-numbers]
| Finding | Detail |
| -------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Frequency** | One study of a medium-sized datacenter: **\~12 network faults per month — half disconnected a single machine, half disconnected an ENTIRE RACK** |
| **Redundancy doesn't help as much as you'd think** | **Adding redundant networking gear doesn't reduce faults as much as expected, since it DOESN'T GUARD AGAINST HUMAN ERROR (misconfigured switches), which is a MAJOR CAUSE OF OUTAGES** |
| **Physical causes** | Wide-area fiber interruptions blamed on **cows, beavers, and sharks** (shark bites now rarer with better cable shielding). **Humans too: accidental misconfiguration, scavenging, sabotage** |
| **Delay magnitude** | **Across cloud regions, round-trip times of UP TO SEVERAL MINUTES at high percentiles.** Even within one datacenter, **packet delay of more than a minute** during a topology reconfiguration triggered by a switch software upgrade. **We have to assume messages might be delayed ARBITRARILY** |
| **Asymmetric and partial faults** | **A and B can communicate, B and C can communicate, but A and C cannot.** Or: **a network interface that DROPS ALL INBOUND PACKETS BUT SENDS OUTBOUND PACKETS SUCCESSFULLY.** *"Just because a network link works in one direction doesn't guarantee it's also working in the opposite direction."* |
| **Aftershocks** | **Even a BRIEF network interruption can have repercussions lasting MUCH LONGER than the original issue** |
> **If the error handling of network faults is not defined and tested, ARBITRARILY BAD THINGS could happen — the cluster could become DEADLOCKED and permanently unable to serve requests even when the network recovers, or it could POTENTIALLY DELETE ALL OF YOUR DATA. If software is put in an unanticipated situation, it may do arbitrary unexpected things.**
**Important nuance:**
> **Handling network faults doesn't necessarily mean TOLERATING them. If your network is normally fairly reliable, a valid approach may be to SIMPLY SHOW AN ERROR MESSAGE to users. However, you DO need to know how your software reacts and ensure the system can RECOVER.**
*(Terminology: **network partition / netsplit** — one part of the network cut off from the rest. **Not fundamentally different from other network interruptions, and NOT related to sharding**, Ch 7.)*
#### 2.4 Fault detection [#24-fault-detection]
**Systems that must detect faulty nodes:** a **load balancer** taking a dead node out of rotation; a **single-leader database** promoting a follower.
**Four cases where you get explicit feedback:**
| Signal | Limitation |
| ---------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------- |
| **TCP RST / FIN** — machine reachable, no process listening on the port | Only if the machine is up and the process crashed |
| **A script notifying other nodes of a crash** (HBase does this) — process died but the OS is running | Only if the OS survives |
| **Querying network switch management interfaces** for link failures | **Ruled out if you're on the internet, in a shared datacenter with no switch access, or if a network problem blocks the management interface** |
| **ICMP Destination Unreachable** from a router | **The router doesn't have a magic failure detection capability either; it is SUBJECT TO THE SAME LIMITATIONS as other participants** |
> **Rapid feedback is useful, but YOU CAN'T COUNT ON IT. In general you have to assume you will get NO RESPONSE AT ALL.**
>
> **You need to strike a balance between FALSE POSITIVES and FALSE NEGATIVES: too short a timeout causes alive nodes to be incorrectly suspected dead; too long a timeout causes unnecessary delays waiting for dead nodes.**
#### 2.5 Timeouts and unbounded delays [#25-timeouts-and-unbounded-delays]
**The cost of a premature declaration of death:**
**The fictitious world where timeouts are easy:**
> If **every packet is either delivered within time `d` or lost**, and **a non-failed node always handles a request within time `r`**, then **every successful request receives a response within `2d + r`** — a reasonable timeout.
>
> **Unfortunately, most systems have NEITHER guarantee.** Asynchronous networks have **UNBOUNDED DELAYS** (no upper limit on arrival time), and **most server implementations cannot guarantee handling requests within a maximum time.**
>
> **For failure detection, it's NOT SUFFICIENT for the system to be fast MOST of the time: if your timeout is low, it takes only A TRANSIENT SPIKE in round-trip times to throw the system off balance.**
##### Network congestion and queueing — the four queues [#network-congestion-and-queueing--the-four-queues]
**Plus:** when TCP retransmits a lost packet, **the application doesn't see the packet loss directly — it sees THE RESULTING DELAY** (waiting for the timeout, then for the retransmitted packet's ACK).
**TCP vs UDP:** latency-sensitive applications (videoconferencing, VoIP) use UDP — **a trade-off between reliability and variability of delays.** UDP **avoids some causes of variable delay** (no flow control, no retransmission) though **still susceptible to switch queues and scheduling delays.**
> **UDP is a good choice WHEN DELAYED DATA IS WORTHLESS.** In a VoIP call there isn't time to retransmit before the data is due to play; **the application fills the missing slot with silence and moves on. THE RETRY HAPPENS AT THE HUMAN LAYER INSTEAD** — *"Could you repeat that please? The sound just cut out."*
**Two aggravating factors:**
* **Queueing delays have an especially wide range when a system is close to maximum capacity. A system with plenty of spare capacity can easily DRAIN queues; in a highly utilized system, LONG QUEUES BUILD UP VERY QUICKLY.** *(Same curve as Ch 2 §2.1.)*
* **Multitenancy:** network links, switches, NICs, and CPUs are shared. **Because you have no control over or insight into other customers' usage, network delays can be highly variable if a NOISY NEIGHBOUR is using a lot of resources.**
**Therefore, timeouts must be chosen experimentally:**
> **Measure the distribution of round-trip times over an extended period and over many machines to determine the expected variability. Then determine an appropriate trade-off between failure-detection delay and risk of premature timeouts.**
>
> **Even better: rather than configured constants, CONTINUALLY MEASURE response times and their variability (JITTER) and AUTOMATICALLY ADJUST timeouts.** The **Phi Accrual failure detector** (used in Akka and Cassandra) does this. **TCP retransmission timeouts work similarly.**
#### 2.6 Synchronous vs asynchronous networks — why we chose unreliability [#26-synchronous-vs-asynchronous-networks--why-we-chose-unreliability]
**The telephone network is the counter-example:** a call **establishes a CIRCUIT — a fixed, guaranteed amount of bandwidth allocated along the entire route, remaining in place until the call ends.** ISDN runs at **4,000 frames/second; a call is allocated 16 bits within each frame in each direction — so each side is guaranteed to send exactly 16 bits of audio every 250 microseconds.**
> **This network is SYNCHRONOUS: even passing through several routers, it does not suffer from queueing, BECAUSE THE SPACE HAS ALREADY BEEN RESERVED IN THE NEXT HOP. And because there is no queueing, the maximum end-to-end latency is fixed — a BOUNDED DELAY.**
**Why datacenters and the internet don't do this:**
**The deeper framing — variable delay as dynamic resource partitioning:**
> A wire carrying 10,000 simultaneous calls divides the resource **STATICALLY: even if you're the only call and 9,999 slots are unused, your circuit gets the same fixed bandwidth.**
>
> **The internet shares bandwidth DYNAMICALLY. Senders push and jostle to get packets over the wire, and switches decide which packet to send from one moment to the next. The downside is queueing; the advantage is that IT MAXIMIZES UTILIZATION OF THE WIRE. The wire has a fixed cost, so if you utilize it better, EACH BYTE IS CHEAPER.**
>
> **The same applies to CPUs:** sharing a core dynamically among threads means a thread **can be paused for varying lengths of time — but it utilizes the hardware better than a static allocation. Better utilization is also why cloud platforms run several VMs from different customers on the same physical machine.**
> ### **Variable delays in networks are NOT A LAW OF NATURE but simply the result of a COST/BENEFIT TRADE-OFF.** [#variable-delays-in-networks-are-not-a-law-of-nature-but-simply-the-result-of-a-costbenefit-trade-off]
>
> **Latency guarantees ARE achievable if resources are STATICALLY PARTITIONED — but at the cost of reduced utilization, i.e. MORE EXPENSIVE. Multitenancy with dynamic partitioning gives better utilization, so it is cheaper, with the downside of variable delays.**
**Hybrid attempts:** **ATM** (an Ethernet competitor in the 1980s, little adoption outside telephone core switches); **InfiniBand** (end-to-end flow control at the link layer, reducing queueing — though it can still suffer link-congestion delays); **QoS mechanisms** (packet prioritization/scheduling, admission control/rate-limiting) which can **emulate circuit switching or provide statistically bounded delay**; **L4S**; Linux **TC**.
> **However, such QoS mechanisms are NOT currently enabled in multitenant datacenters and public clouds, or on the internet. Currently deployed technology does not allow us to make ANY guarantees about delays or reliability. CONSEQUENTLY, THERE'S NO "CORRECT" VALUE FOR TIMEOUTS — they need to be determined experimentally.**
***
# 9.11 Worked examples (/docs/ddia/trouble-distributed-systems/worked-examples)
**① Why you can't distinguish the six cases.** Client sends request at t=0, timeout at t=5 s, no response. Enumerate the posterior: request lost (no effect), request queued (**effect may still happen at t=90 s**), node dead before processing (no effect), node paused (**effect happens when it resumes**), response lost (**effect already happened**), response delayed (**effect already happened**). **Three of six cases mean the write already applied; two mean it may apply later.** Hence: **every non-idempotent operation over a network needs a deduplication key** (Ch 8 §5.6).
**② Drift budget.** 200 ppm = 200 µs per second.
* Resync every 30 s → **6 ms** max drift
* Resync every 60 s → **12 ms**
* Resync once per hour → **720 ms**
* Resync once per day → **17.3 s**
**If your application assumes clocks agree within 100 ms, you need to resync at least every \~8 minutes** — *and that's the drift term alone, before NTP's own 35 ms+ error.*
**③ Why LWW loses data proportionally to skew.** Node A's clock is 200 ms ahead. A writes `x` at true time T (stamped T+200 ms). B writes `x` at true time T+150 ms (stamped T+150 ms). B's write is **genuinely later** but has the **smaller** timestamp, so it loses. **Every write B makes in the 200 ms window after any A write is silently discarded** — and at 1,000 writes/s that's **200 lost writes per A-write**, with no error anywhere.
**④ Spanner's commit-wait cost.** Uncertainty ε = 7 ms ⇒ commit-wait ≈ 2ε ≈ **14 ms added to every read/write transaction's commit.** Halve ε to 3.5 ms (better clock hardware) and you halve the wait. **This is the direct, quantifiable reason Google puts atomic clocks in datacenters** — it's not for accuracy per se, it's to shrink a latency tax on every write.
**⑤ Lease safety margin.** Lease duration L = 30 s, renewed when \< 10 s remain, so the renewal deadline gives 10 s of slack. Worst realistic pause: a 15 s GC or VM suspend. **10 s slack \< 15 s pause ⇒ the zombie window is real.** Options: raise the margin above the worst pause (slower failover), reduce the worst pause (GC tuning, §4.3), or — **the only sound answer — make the pause harmless with fencing tokens**, since you cannot bound the pause.
**⑥ Quorum sizes and majorities.** n=5, majority=3. Two disjoint majorities would need ≥6 nodes, so **there can only ever be one** — that's the entire safety argument. Now suppose one node suffers amnesia after a disk wipe and rejoins claiming to have data it lost: the majority intersection guarantee still holds *numerically* but **the value it returns is wrong**, which is §6.4's point that the model's stable-storage assumption is doing real work.
***
# 12.10 What actually breaks in production — Ch. 12 consolidated (/docs/kafka/administering-kafka/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------ |
| 1 | **Anyone with shell access changed prod topics; no audit trail** | *"Default configurations DO NOT RESTRICT the use of these tools"* | Restrict tool access to administrators; ACLs alone don't cover ZK-writing tools |
| 2 | **A CLI tool corrupted cluster state** | Tool version ≠ broker version; some tools write to ZooKeeper | **Run tools on the brokers themselves**, using the deployed version |
| 3 | **Two topics merged in dashboards** | **Periods in topic names** become underscores in metrics | Never use `.` in topic names |
| 4 | **Confusion with Kafka internals** | Topic name starts with `__` | Reserved by convention for internal topics |
| 5 | **Automation silently skipped a config change** | `--if-exists` on `--alter` masked a missing topic | *"Using it is NOT RECOMMENDED"* — it hides real problems |
| 6 | **Producers failing with NotEnoughReplicas; nobody noticed the buildup** | No monitoring on the ISR ladder | Alert on **`--at-min-isr-partitions`** (before it becomes under-min-ISR) |
| 7 | **A partition is completely offline** | No leader available | `--unavailable-partitions`; possibly unclean election (Ch. 7 §3.2) |
| 8 | **URP alerts fire constantly during deploys** | URPs are **expected** during maintenance/rebalance | *"Isn't necessarily bad"* — alert on duration/trend, not presence |
| 9 | **Keyed consumers break after adding partitions** | `hash(key) % N` changed | *"Set the number of partitions ONCE... and avoid resizing"* |
| 10 | **Need fewer partitions; can't** | Reducing partitions is impossible | Delete + re-create, or create `topic-v2` and migrate producers |
| 11 | **`--delete` appeared to do nothing** | `delete.topic.enable=false` — the request is **ignored** | Enable it (accepting the risk), or delete manually with full downtime |
| 12 | **Controller overwhelmed; cluster sluggish after a cleanup script** | Deleted many topics at once | *"NOT MORE THAN ONE OR TWO TOPICS AT A TIME"*; consider `controller_mutations_rate` |
| 13 | **Deleted the wrong topic** | *"NOT A REVERSIBLE OPERATION"* — and no success/failure output | `delete.topic.enable=false` as a guardrail; verify with `--list`/`--describe` |
| 14 | **Cluster slow with thousands of unused topics** | Even empty topics consume disk, filehandles, memory, **and controller metadata** | Delete unused topics (carefully, per #12) |
| 15 | **`--delete --group` failed: "The group is not empty"** | Group has active members | Shut down all consumers first |
| 16 | **Offsets reset accidentally** | Ran the export command **without `--dry-run`** | The export/destroy commands differ by one flag — script it carefully |
| 17 | **Imported offsets had no effect** | Consumers were **running** and overwrote them | *"ALL consumers in the group are STOPPED"* first |
| 18 | **A client's quota is 5× smaller than expected** | Quotas are **per-broker**; all leadership landed on one broker | Keep leadership balanced (preferred leader election / Cruise Control) |
| 19 | **Two unrelated consumer groups share a quota** | Same `client.id` across groups | *"Best practice to set the client ID for each consumer group to something unique"* |
| 20 | **Automation misread a topic's effective config** | `--describe` shows **only overrides**, never cluster defaults | Keep separate knowledge of defaults, or use **AdminClient's `describeConfigs`** (Ch. 5), which reports `isDefault()` |
| 21 | **Leadership badly skewed after a rolling restart** | Original leaders **do not automatically resume leadership** | `auto.leader.rebalance.enable`, or run preferred leader election, or Cruise Control |
| 22 | **A wrapper script around the console consumer lost messages** | *"Difficult to interact with the console consumer in a way that does not lose messages"* | Use the real client libraries |
| 23 | **Console producer messages all have null keys** | No tab separator, or `parse.key=false` | Set `key.separator` and `parse.key` via **`--property`** (not `--producer-property`) |
| 24 | **A config passed to the client had no effect** | Used `--property` (formatter) instead of `--producer-property`/`--consumer-property` (client) | Know which is which |
| 25 | **Cluster performance tanked during a reassignment** | Reassignment copies whole partitions; **disrupts page cache** + network + disk I/O | **`--throttle`**, combinable with **`--additional`** to throttle an in-flight move |
| 26 | **Reassignment off a decommissioning broker is glacially slow** | That broker is the leader for everything being copied — one NIC serving all reads | **Move leadership off first** (Cruise Control demotion, or just **bounce the broker**, temporarily disabling auto-rebalance) |
| 27 | **Can't verify reassignment progress** | Lost the JSON file used in `--execute` | Keep both generated files (`revert-` and `expand-`) |
| 28 | **A reassignment proposal is impossible to satisfy** | Rack-awareness constraints | `--disable-rack-aware` (understanding the durability cost) |
| 29 | **RF didn't increase after changing the cluster default** | *"Existing topics will NOT automatically be increased"* | Reassign with an extra broker ID in each replica set |
| 30 | **A cancelled reassignment left the cluster worse off** | `--cancel` reverts to the **prior** replica set — which may include a dead/overloaded broker; and **replica order isn't guaranteed** (so preferred leaders may change) | Understand before cancelling; re-run preferred leader election afterward |
| 31 | **A poison-pill message breaks a consumer and you can't see it** | — | **`kafka-dump-log.sh --print-data-log`** on the right segment |
| 32 | **Consumption errors that look like corruption** | Corrupt **index** file | `--index-sanity-check` / `--verify-index-only`; indexes regenerate |
| 33 | **Replicas silently diverged (gaps in a follower)** | *"previously replicated log segments can get deleted from a broker, and the follower WILL NOT FILL IN THE GAPS"* | `kafka-replica-verification.sh` — but see #34 |
| 34 | **Running replica verification destroyed cluster performance** | It reads **all messages from the oldest offset, from all replicas, in parallel, in a loop** | Use sparingly, off-peak, scoped by topic regex |
| 35 | **Controller alive but non-functional** | Controller thread hit an exception | Delete `/admin/controller` znode → forces resignation and re-election (**cannot choose the successor**) |
| 36 | **A topic is stuck "marked for deletion" forever** | Deletion requested with deletion disabled, or replicas went offline mid-delete | Delete `/admin/delete_topic/` (**not the parent**), then force a controller move to clear cached requests |
| 37 | **Cluster unstable after editing ZooKeeper** | Modified topic metadata **while brokers were online** | *"NEVER attempt to delete or modify topic metadata in ZooKeeper while the cluster is online"* — manual deletion requires **full shutdown** |
***
# 12.2 Consumer groups — `kafka-consumer-groups.sh` (/docs/kafka/administering-kafka/consumer-groups-kafka-consumer)
> Can *"list consumer groups, describe specific groups, delete consumer groups or specific group info, or reset consumer group offset information."*
> ### ZOOKEEPER-BASED CONSUMER GROUPS [#zookeeper-based-consumer-groups]
>
> *"In older versions, consumer groups could be managed in ZooKeeper. **This behavior was DEPRECATED in versions 0.11.0.\* and later, and old consumer groups are no longer used.** Some versions of the provided scripts **may still show deprecated `--zookeeper` connection string commands, but it is NOT RECOMMENDED to use them** unless you have an old environment."*
#### 2.1 List and describe [#21-list-and-describe]
```bash
kafka-consumer-groups.sh --bootstrap-server localhost:9092 --list
# console-consumer-95554 ← ad hoc consumers appear as
# console-consumer-9581 ← console-consumer-
# my-consumer
```
```bash
kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--describe --group my-consumer
GROUP TOPIC PARTITION CURRENT-OFFSET LOG-END-OFFSET LAG CONSUMER-ID HOST CLIENT-ID
my-consumer my-topic 0 2 4 2 consumer-1-029af... /127.0.0.1 consumer-1
my-consumer my-topic 1 2 3 1 consumer-1-029af... /127.0.0.1 consumer-1
my-consumer my-topic 2 2 3 1 consumer-2-42c1a... /127.0.0.1 consumer-2
```
**The fields, precisely defined (Table 12-1):**
| Field | Description |
| -------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| `GROUP` | The consumer group name |
| `TOPIC` | Topic being consumed |
| `PARTITION` | Partition ID |
| **`CURRENT-OFFSET`** | *"**The NEXT offset to be consumed** by the group for this topic partition. **This is the POSITION of the consumer within the partition.**"* |
| **`LOG-END-OFFSET`** | *"The current **HIGH-WATER MARK** offset from the broker... **the offset of the next message to be produced** to this partition."* |
| **`LAG`** | *"The **difference** between `CURRENT-OFFSET` and `LOG-END-OFFSET`"* |
| `CONSUMER-ID` | *"A generated unique consumer-id **based on the provided client-id**"* |
| `HOST` | *"Address of the host the consumer group is reading from"* |
| `CLIENT-ID` | *"String provided by the client identifying the client"* |
*(Note `LOG-END-OFFSET` is the **high-water mark**, not the log-end offset in the Ch. 6 §5.6 sense — so this LAG is measured against what's *readable*, which is the right thing.)*
#### 2.2 Delete a group [#22-delete-a-group]
```bash
kafka-consumer-groups.sh --bootstrap-server localhost:9092 --delete --group my-consumer
# Deletion of requested consumer groups ('my-consumer') was successful.
```
> *"This will remove the **entire group, including ALL STORED OFFSETS for ALL topics** the group is consuming. **To perform this action, ALL CONSUMERS IN THE GROUP SHOULD BE SHUT DOWN** as the group must not have any active members. **If you attempt to delete a group that is not empty, an error stating 'The group is not empty' will be thrown and NOTHING WILL HAPPEN.**"*
> 💡 *"It is also possible to use the same command to **delete offsets for a SINGLE TOPIC** that the group is consuming **without deleting the entire group** by adding the `--topic` argument."*
#### 2.3 Offset management [#23-offset-management]
**Why:** *"useful for **resetting the offsets for a consumer when there is a problem that requires messages to be reread**, or for **advancing offsets and skipping past a message that the consumer is having a problem with** (e.g., if there is a **badly formatted message** that the consumer cannot handle)."*
##### Export (via `--dry-run`) [#export-via---dry-run]
```bash
kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--export --group my-consumer --topic my-topic \
--reset-offsets --to-current --dry-run > offsets.csv
cat offsets.csv
# my-topic,0,8905
# my-topic,1,8915
# ...
```
**CSV format:** `,,`
> ### ⚠️ *"Running the same command WITHOUT the `--dry-run` option **will RESET THE OFFSETS COMPLETELY, so be careful.**"* [#️-running-the-same-command-without-the---dry-run-option-will-reset-the-offsets-completely-so-be-careful]
That's a genuinely dangerous ergonomic: the *export* command and the *destructive reset* command differ by one flag.
##### Import [#import]
```bash
kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--reset-offsets --group my-consumer \
--from-file offsets.csv --execute
TOPIC PARTITION NEW-OFFSET
my-topic 0 8905
...
```
> **The recommended workflow:** *"A common practice is to **export the current offsets, MAKE A COPY OF THE FILE (so that you preserve a backup), and EDIT THE COPY** to replace the offsets with the desired values."*
> ### ⚠️ STOP CONSUMERS FIRST [#️-stop-consumers-first]
>
> *"Before performing this step, **it is important that ALL consumers in the group are STOPPED. THEY WILL NOT READ THE NEW OFFSETS IF THEY ARE WRITTEN WHILE THE CONSUMER GROUP IS ACTIVE. THE CONSUMERS WILL JUST OVERWRITE THE IMPORTED OFFSETS.**"*
*(Same constraint as Ch. 5 §6.4 — and there it explains *why*: groups only read offsets on assignment or startup.)*
***
# 12.6 Dumping log segments — `kafka-dump-log.sh` (/docs/kafka/administering-kafka/dumping-log-segments-kafka)
**Why:** *"perhaps because you ended up with a **'POISON PILL' message in your topic that is CORRUPTED and your consumer cannot handle it.** ... **This will allow you to VIEW INDIVIDUAL MESSAGES WITHOUT NEEDING TO CONSUME AND DECODE THEM.**"*
**Path convention:** `/-/`
**Metadata only:**
```bash
kafka-dump-log.sh --files /tmp/kafka-logs/my-topic-0/00000000000000000000.log
Starting offset: 0
baseOffset: 0 lastOffset: 0 count: 1 baseSequence: -1 lastSequence: -1
producerId: -1 producerEpoch: -1 partitionLeaderEpoch: 0
isTransactional: false isControl: false position: 0
CreateTime: 1623034799990 size: 77 magic: 2
compresscodec: NONE crc: 1773642166 isvalid: true
...
```
**With payloads — add `--print-data-log`:**
```bash
kafka-dump-log.sh --files .../00000000000000000000.log --print-data-log
baseOffset: 0 lastOffset: 0 count: 1 ...
| offset: 0 CreateTime: 1623034799990 keysize: -1 valuesize: 9
sequence: -1 headerKeys: [] payload: Message 1
...
```
💡 **This output is Chapter 6 §6.5's batch header, made concrete.** Map them:
**Index validation options:**
> *"The index is used for **finding messages within a log segment**, and **if corrupted, will cause ERRORS IN CONSUMPTION. Validation is performed WHENEVER A BROKER STARTS UP IN AN UNCLEAN STATE** (i.e., it was not stopped normally), **but it can be performed manually as well.**"*
| Option | Depth |
| ----------------------- | ---------------------------------------------------------------------------------------- |
| `--index-sanity-check` | *"just check that the index is **in a usable state**"* |
| `--verify-index-only` | *"check for **mismatches** in the index **without printing out all the index entries**"* |
| `--value-decoder-class` | *"allows serialized messages to be **deserialized by passing in a decoder**"* |
*(Ch. 6 §6.6: indexes are derived data — safe to delete, auto-regenerated. This tool tells you whether you need to.)*
***
# 12.3 Dynamic configuration changes — `kafka-configs.sh` (/docs/kafka/administering-kafka/dynamic-configuration-changes-kafka)
> *"There is **a plethora of configurations** for topics, clients, brokers, and more that **can be updated DYNAMICALLY DURING RUNTIME without having to shut down or redeploy a cluster.**"*
**Four entity types:** `topics`, `brokers`, `users`, `clients`.
> *"New dynamic configs are being added constantly with each release, so **it is good to ensure you have the same version of this tool that matches the version of Kafka you are running.**"*
>
> 💡 *"For ease of setting up these configs consistently **via automation, the `--add-config-file` argument can be used with a preformatted file of all the configs** you want to manage and update."*
#### 3.1 Topic configuration overrides [#31-topic-configuration-overrides]
```bash
kafka-configs.sh --bootstrap-server localhost:9092 \
--alter --entity-type topics --entity-name my-topic \
--add-config retention.ms=3600000
# Updated config for topic: "my-topic".
```
**The point:** *"we can **override the cluster-level defaults for individual topics to accommodate DIFFERENT USE CASES WITHIN A SINGLE CLUSTER.**"* (Ch. 7's bank example: strict defaults, relaxed complaints topic.)
**Valid topic config keys (Table 12-2), grouped by purpose:**
**Note `message.downconversion.enable`** — you can *disable* down-conversion per topic, forcing old clients to fail loudly instead of silently burning broker CPU (Ch. 6 §6.5).
#### 3.2 Client and user overrides — all quotas [#32-client-and-user-overrides--all-quotas]
> *"For Kafka clients and users, there are **only a few configurations that can be overridden, which are all essentially types of QUOTAS.**"*
| Config key | Description |
| ------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `consumer_bytes_rate` | *"bytes a single client ID is allowed to **consume from a single broker** in one second"* |
| `producer_bytes_rate` | *"bytes a single client ID is allowed to **produce to a single broker** in one second"* |
| **`controller_mutations_rate`** | *"The rate at which mutations are accepted for the **create topics** request, the **create partitions** request, and the **delete topics** request. **The rate is accumulated by the NUMBER OF PARTITIONS created or deleted.**"* |
| `request_percentage` | *"The percentage per quota window (**out of a total of (num.io.threads + num.network.threads) × 100%**) for requests from the user or client"* |
`controller_mutations_rate` is the guard against the controller-overload problem from §1.7 — someone scripting mass topic creation/deletion.
#### ⚠️ Quotas are PER-BROKER — the balance dependency [#️-quotas-are-per-broker--the-balance-dependency]
> *"**Because throttling occurs on a PER-BROKER basis, EVEN BALANCE OF LEADERSHIP of partitions across a cluster becomes PARTICULARLY IMPORTANT to enforce this properly.**"*
> ### CLIENT ID vs CONSUMER GROUP [#client-id-vs-consumer-group]
>
> \*"**The client ID is NOT necessarily the same as the consumer group name.** Consumers can set their own client ID, and **you may have many consumers that are in DIFFERENT GROUPS that specify the SAME client ID.**
>
> 💡 *"It is considered **a best practice to set the client ID for each consumer group to something unique that identifies that group.** This allows **a single consumer group to SHARE A QUOTA**, and it makes it easier to identify in logs what group is responsible for requests."*
**Combining user and client in one command:**
```bash
kafka-configs.sh --bootstrap-server localhost:9092 \
--alter --add-config "controller_mutations_rate=10" \
--entity-type clients --entity-name \
--entity-type users --entity-name
```
#### 3.3 Broker overrides [#33-broker-overrides]
> *"**More than 80 overrides** can be altered with `kafka-configs.sh` for brokers."* Three called out specifically:
| Config | Purpose |
| ------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **`min.insync.replicas`** | *"Adjusts the minimum number of replicas that need to acknowledge a write for a produce request to be successful **when producers have set acks to all (or –1)**"* |
| **`unclean.leader.election.enable`** | *"Allows replicas to be elected as leader **even if it results in data loss.** Useful when it is **permissible to have some lossy data, or to turn on FOR SHORT TIMES to UNSTICK a Kafka cluster if unrecoverable data loss cannot be avoided.**"* |
| **`max.connections`** | *"maximum connections allowed to a broker at any time. We can also use **`max.connections.per.ip`** and **`max.connections.per.ip.overrides`** for more fine-tuned throttling."* |
💡 **The fact that `unclean.leader.election.enable` is dynamically changeable is important** — Ch. 7 §3.2 described it as requiring a restart, but as a dynamic broker config you can flip it on, recover, and flip it back **without bouncing the cluster.**
#### 3.4 Describing and removing overrides [#34-describing-and-removing-overrides]
```bash
kafka-configs.sh --bootstrap-server localhost:9092 \
--describe --entity-type topics --entity-name my-topic
# Configs for topics:my-topic are
# retention.ms=3600000
```
> ### ⚠️ TOPIC OVERRIDES ONLY [#️-topic-overrides-only]
>
> *"The configuration description **will ONLY SHOW OVERRIDES — it does NOT include the cluster DEFAULT configurations. THERE IS NOT A WAY TO DYNAMICALLY DISCOVER THE CONFIGURATION OF THE BROKERS THEMSELVES.** This means that **when using this tool to discover topic or client settings in AUTOMATION, THE USER MUST HAVE SEPARATE KNOWLEDGE OF THE CLUSTER DEFAULT CONFIGURATION.**"*
*(Contrast Ch. 5 §5: the AdminClient's `describeConfigs` **does** return defaults with `isDefault()` — so the programmatic API is strictly more capable here than the CLI.)*
**Removing an override reverts to the cluster default:**
```bash
kafka-configs.sh --bootstrap-server localhost:9092 \
--alter --entity-type topics --entity-name my-topic \
--delete-config retention.ms
```
***
# 12. Administering Kafka (/docs/kafka/administering-kafka)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 12*
> **Scope, stated honestly:** *"While these tools provide basic functions, **you may find they are LACKING for more complex operations or are UNWIELDY TO USE AT LARGER SCALES.** This chapter will describe **only the basic tools that are part of the Apache Kafka open source project.**"*
**Mechanics:** *"The tools are implemented in **Java classes**, and a set of **scripts** are provided natively to call those classes properly."* All examples assume `/usr/local/kafka/bin/` is your working directory or on `$PATH`.
***
# 12.11 Operational quick reference (/docs/kafka/administering-kafka/operational-quick-reference)
#### The chapter's own closing advice [#the-chapters-own-closing-advice]
> *"As you begin to scale your Kafka clusters larger, **EVEN THE USE OF THESE TOOLS MAY BECOME ARDUOUS AND DIFFICULT TO MANAGE. IT IS HIGHLY RECOMMENDED TO ENGAGE WITH THE OPEN SOURCE KAFKA COMMUNITY and take advantage of the many other open source projects in the ecosystem TO HELP AUTOMATE many of the tasks outlined in this chapter.**"*
***
# 12.8 Other tools in the distribution (/docs/kafka/administering-kafka/other-tools-distribution)
| Tool | Purpose |
| ---------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`kafka-acls.sh`** | *"interacting with access controls for Kafka clients. Includes full features for **authorizer properties, set up for deny or allow principles, cluster- or topic-level restrictions, ZooKeeper TLS file configuration**"* (Ch. 11) |
| **`kafka-mirror-maker.sh`** | *"A **lightweight** script for mirroring data"* (Ch. 10) |
| **`kafka-broker-api-versions.sh`** | *"helps to easily **identify different versions of usable API elements when UPGRADING** from one Kafka version to another and **check for compatibility issues**"* (Ch. 6 §5.9) |
| Producer/consumer **performance test** scripts | benchmarking (used for MirrorMaker tuning, Ch. 10 §8.2) |
| ZooKeeper administration scripts | — |
| **`trogdor.sh`** | *"a **test framework** designed to run benchmarks and other workloads to attempt to **STRESS TEST** the system"* (Ch. 7 §6.2) |
***
# 12.5 Partition management (/docs/kafka/administering-kafka/partition-management)
#### 5.1 Preferred replica election [#51-preferred-replica-election]
**The problem it solves:**
> *"**Leadership is defined within Kafka as THE FIRST IN-SYNC REPLICA IN THE REPLICA LIST.** However, **when a broker is stopped or loses connectivity, leadership is transferred to another in-sync replica, and THE ORIGINAL DOES NOT RESUME LEADERSHIP OF ANY PARTITIONS AUTOMATICALLY. THIS CAN CAUSE WILDLY INEFFICIENT BALANCE AFTER A DEPLOYMENT ACROSS A FULL CLUSTER if automatic leader balancing is not enabled.**"*
> *"it is recommended to **ensure that this setting is enabled** or to use **other open source tooling such as Cruise Control** to ensure that a good balance is maintained at all times."*
**The fix — a cheap, safe operation:**
> *"a **lightweight, GENERALLY NON-IMPACTING** procedure called **preferred leader election.** This tells the cluster controller to **select the ideal leader** for partitions. **Clients can track leadership changes automatically**, so they will be able to move to the new broker."*
```bash
# all topics
kafka-leader-election.sh --bootstrap-server localhost:9092 \
--election-type PREFERRED --all-topic-partitions
# specific partitions from a JSON file
kafka-leader-election.sh --bootstrap-server localhost:9092 \
--election-type PREFERRED --path-to-json-file partitions.json
```
```json
{ "partitions": [
{ "partition": 1, "topic": "my-topic" },
{ "partition": 2, "topic": "foo" }
] }
```
Also supports `--topic` + `--partition` directly.
> *"An older version of this tool called **`kafka-preferred-replica-election.sh`** is also available **but has been DEPRECATED in favor of the new tool, which allows for more customization, such as specifying whether we want a 'PREFERRED' or 'UNCLEAN' election type.**"*
#### 5.2 Reassigning replicas — `kafka-reassign-partitions.sh` [#52-reassigning-replicas--kafka-reassign-partitionssh]
**Four reasons you'd do it:**
**The three-step process:**
##### Step 1 — generate a proposal [#step-1--generate-a-proposal]
Scenario: *"a four-broker cluster. You've recently added two new brokers, bringing the total up to six, and you want to move two of your topics onto brokers 5 and 6."*
```json
// topics.json
{ "topics": [ { "topic": "foo1" }, { "topic": "foo2" } ], "version": 1 }
```
```bash
kafka-reassign-partitions.sh --bootstrap-server localhost:9092 \
--topics-to-move-json-file topics.json --broker-list 5,6 --generate
```
Output gives **two** JSON blocks:
> ⚠️ *"You'll notice in the output that **there isn't a good balance of leadership, as the proposal will result in ALL LEADERSHIP MOVING TO BROKER 5.** We will ignore this for now and **presume the cluster automatic leadership balancing is enabled**, which will help distribute it later."*
>
> 💡 *"the first step **can be SKIPPED if you know exactly where you want to move your partitions** to and you manually craft the JSON."*
##### Step 2 — execute [#step-2--execute]
```bash
kafka-reassign-partitions.sh --bootstrap-server localhost:9092 \
--reassignment-json-file expand-cluster-reassignment.json --execute
```
> *"**Save this to use as the `--reassignment-json-file` option during rollback**"* — the tool reminds you.
**What actually happens under the hood:**
**Three useful flags:**
| Flag | Purpose |
| -------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`--additional`** | *"allows you to **add to the EXISTING reassignments** so they can continue **without interruption and WITHOUT the need to wait until the original movements have completed** in order to start a new batch"* |
| **`--disable-rack-aware`** | *"There may be times when, due to rack awareness settings, **the end-state of a proposal may not be POSSIBLE.** This can be overridden"* |
| **`--throttle`** | bytes/sec. *"Partition reassignments have **a big impact on the performance of your cluster**, as they will cause **changes in the consistency of the MEMORY PAGE CACHE** and use **network and disk I/O.** ... Can be combined with `--additional` **to throttle an ALREADY-STARTED reassignment process that may be causing issues.**"* |
**That last point is the emergency brake:** `--throttle` + `--additional` lets you slow down a reassignment you already started and which is hurting the cluster.
*(And note the page-cache remark — reassignment reads cold data from the beginning of partitions, evicting the hot pages that serve your real consumers. Ch. 6 §6.2's tiered-storage isolation argument, again.)*
> ### 💡 IMPROVING NETWORK UTILIZATION WHEN REASSIGNING REPLICAS [#-improving-network-utilization-when-reassigning-replicas]
>
> \*"When removing many partitions from a single broker — such as if that broker is being removed from the cluster — **it may be useful to REMOVE ALL LEADERSHIP FROM THE BROKER FIRST.**
>
> Doing that manually *"is arduous."* Options:
>
> * **Cruise Control** includes broker **"demotion,"** which *"safely moves leadership off a broker and is **probably the simplest way** to do this."*
> * **Without such tools: A SIMPLE RESTART OF A BROKER WILL SUFFICE.** *"As a broker is preparing to shut down, all leadership for its partitions will move to other brokers. **This can SIGNIFICANTLY INCREASE THE PERFORMANCE of reassignments and REDUCE THE IMPACT on the cluster, as the replication traffic will be DISTRIBUTED TO MANY BROKERS.**"*
> * ⚠️ *"**However, if automatic leader reassignment is enabled after the broker is bounced, LEADERSHIP MAY RETURN to this broker, so it may be beneficial to TEMPORARILY DISABLE this feature.**"*
##### Step 3 — verify [#step-3--verify]
```bash
kafka-reassign-partitions.sh --bootstrap-server localhost:9092 \
--reassignment-json-file expand-cluster-reassignment.json --verify
Reassignment of partition [foo1,0] completed successfully
Reassignment of partition [foo1,1] is in progress
...
```
> *"This will show which reassignments are **currently in progress**, which have **completed**, and (if there was an error) which have **failed. To do this, you MUST HAVE THE FILE with the JSON object that was used in the execute step.**"*
#### 5.3 Changing the replication factor [#53-changing-the-replication-factor]
**Why:** *"a partition was created with the wrong RF, you want increased redundancy as you expand your cluster, or you want to decrease redundancy for cost savings."*
> 💡 *"One clear example is that **if a cluster RF DEFAULT setting is adjusted, EXISTING TOPICS WILL NOT AUTOMATICALLY BE INCREASED.**"*
**Method: craft the JSON with an extra broker ID in the replica set.** RF 2 → RF 3 by adding broker 4 to `[5,6]`:
```json
{ "version":1,
"partitions":[{"topic":"foo1","partition":1,"replicas":[5,6,4]},
{"topic":"foo1","partition":2,"replicas":[5,6,4]},
{"topic":"foo1","partition":3,"replicas":[5,6,4]}] }
```
Then `--execute` and confirm via `--verify` or `kafka-topics.sh --describe`:
```txt
Topic:foo1 PartitionCount:3 ReplicationFactor:3 Configs:
Topic: foo1 Partition: 0 Leader: 5 Replicas: 5,6,4 Isr: 5,6,4
```
#### 5.4 Canceling reassignments [#54-canceling-reassignments]
> *"Canceling a replica reassignment **in the past was a DANGEROUS process that required unsafe manual manipulation of ZooKeeper nodes** (deleting the `/admin/reassign_partitions` znode). **Fortunately, this is NO LONGER THE CASE.**"*
**`--cancel`** *"will cancel the active reassignments that are ongoing in a cluster... designed to **RESTORE THE REPLICA SET TO THE ONE IT WAS PRIOR TO reassignment being initiated.**"*
***
# 12.4 Producing and consuming from the console (/docs/kafka/administering-kafka/producing-consuming-console)
> ### ⚠️ PIPING OUTPUT TO ANOTHER APPLICATION [#️-piping-output-to-another-application]
>
> *"While it is possible to write applications that wrap around the console consumer or producer... **this type of application is QUITE FRAGILE and SHOULD BE AVOIDED. IT IS DIFFICULT TO INTERACT WITH THE CONSOLE CONSUMER IN A WAY THAT DOES NOT LOSE MESSAGES.** Likewise, **the console producer DOES NOT ALLOW FOR USING ALL FEATURES, and properly sending bytes is tricky. IT IS BEST TO USE EITHER THE JAVA CLIENT LIBRARIES DIRECTLY or a third-party client library.**"*
#### 4.1 Console producer [#41-console-producer]
```bash
kafka-console-producer.sh --bootstrap-server localhost:9092 --topic my-topic
>Message 1
>Test Message 2
>^D ← EOF (Ctrl-D) closes the client
```
> *"By default, messages are read **one per line, with a TAB CHARACTER separating the key and the value** (if no tab character is present, **the key is null**). ... the producer reads in and produces **raw bytes** using the default serializer (**`DefaultEncoder`**)."*
**Passing producer configs — two ways:**
> ### ⚠️ CONFUSING COMMAND-LINE OPTIONS [#️-confusing-command-line-options]
>
> *"The **`--property`** option is available for both the console producer and consumer, **but this should NOT be confused with `--producer-property` or `--consumer-property`. THE `--property` OPTION IS ONLY USED FOR PASSING CONFIGURATIONS TO THE MESSAGE FORMATTER, AND NOT THE CLIENT ITSELF.**"*
**Useful producer options:**
| Option | Meaning |
| ------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--batch-size` | *"number of messages sent in a single batch if they are **not** being sent synchronously"* |
| `--timeout` | *"If a producer is running in **asynchronous** mode, max time waiting for the batch size before producing — **to avoid long waits on low-producing topics**"* |
| `--compression-codec ` | `none`, `gzip`, `snappy`, `zstd`, or `lz4`. **Default: `gzip`** |
| `--sync` | *"Produce messages **synchronously**, waiting for each message to be acknowledged before sending the next"* |
**Line-reader options** (via `--property`, for `kafka.tools.ConsoleProducer$LineMessageReader`):
| Property | Meaning |
| --------------- | ---------------------------------------------------------------------------------------------------------------------------------------- |
| `ignore.error` | *"Set to `false` to **throw an exception** when `parse.key` is `true` and a key separator is not present. **Defaults to `true`.**"* |
| `parse.key` | *"Set to `false` to always set the key to `null`. Defaults to `true`."* |
| `key.separator` | *"delimiter between key and value. **Defaults to a tab character.**"* |
> *"the `LineMessageReader` will **split the input on the FIRST instance of the `key.separator`.** If there are no characters remaining after that, **the value of the message will be EMPTY.** If no key separator is present, or if `parse.key` is `false`, **the key will be `null`.**"*
> ### CHANGING LINE-READING BEHAVIOR [#changing-line-reading-behavior]
>
> *"You can provide **your own class**... must **extend `kafka.common.MessageReader`** and will be responsible for creating the `ProducerRecord`. Specify it with **`--line-reader`**, and **make sure the JAR containing your class is in the classpath.**"*
#### 4.2 Console consumer [#42-console-consumer]
```bash
kafka-console-consumer.sh --bootstrap-server localhost:9092 \
--whitelist 'my.*' --from-beginning
```
> *"messages are printed in standard output, **delimited by a new line.** By default, it outputs **the raw bytes in the message, WITHOUT THE KEY, with no formatting** (using `DefaultFormatter`)."*
**Two mutually exclusive topic selectors:**
**Useful consumer options:**
| Option | Meaning |
| ------------------------- | ---------------------------------------------------------------------------------------------------------- |
| `--formatter ` | Message formatter class. Default `kafka.tools.DefaultMessageFormatter` |
| `--from-beginning` | Consume from the **oldest** offset; otherwise from the **latest** |
| `--max-messages ` | Max messages before exiting |
| `--partition ` | Consume only from that partition |
| `--offset` | An offset ``, or `earliest`, or `latest` |
| `--skip-message-on-error` | *"Skip a message if there is an error when processing instead of halting. **Useful for debugging.**"* |
Configs via `--consumer.config ` or `--consumer-property =`.
**The four message formatters:**
**`DefaultMessageFormatter` properties (Table 12-4), via `--property`:**
> *"The deserializer classes must implement **`org.apache.kafka.common.serialization.Deserializer`**, and the console consumer will call **`toString`** on them to get the output. Typically you would insert them into the classpath by **setting the `CLASSPATH` environment variable** before executing the script."*
#### 💡 4.3 Consuming `__consumer_offsets` — a real diagnostic [#-43-consuming-__consumer_offsets--a-real-diagnostic]
**Why:** *"You may want to see **if a particular group is committing offsets AT ALL, or HOW OFTEN offsets are being committed.**"*
```bash
kafka-console-consumer.sh --bootstrap-server localhost:9092 \
--topic __consumer_offsets --from-beginning --max-messages 1 \
--formatter "kafka.coordinator.group.GroupMetadataManager\$OffsetsMessageFormatter" \
--consumer-property exclude.internal.topics=false
[my-group-name,my-topic,0]::[OffsetMetadata[1,NO_METADATA]
CommitTime 1623034799990 ExpirationTime 1623639599990]
```
**Three things this command requires that are easy to miss:** the special formatter class (with the `$` escaped), `exclude.internal.topics=false`, and `--max-messages` so you don't drown.
*(Note the `ExpirationTime` in the output — that's `offsets.retention.minutes` from Ch. 4 §6.8 made visible.)*
***
# 12.7 Replica verification — `kafka-replica-verification.sh` (/docs/kafka/administering-kafka/replica-verification-kafka-replica)
**Why it's needed — the gap it covers:**
> *"Partition replication works **similar to a regular Kafka consumer client**: the follower broker starts replicating at the oldest offset and **checkpoints the current offset to disk periodically.** When replication stops and restarts, it picks up from the last checkpoint. **IT IS POSSIBLE FOR PREVIOUSLY REPLICATED LOG SEGMENTS TO GET DELETED FROM A BROKER, AND THE FOLLOWER WILL NOT FILL IN THE GAPS IN THIS CASE.**"*
**What the tool does:** *"**fetch messages from ALL the replicas** for a given set of topic partitions, **check that all messages exist on all replicas**, and print out the **max lag** for given partitions. This process will **operate continuously in a loop until canceled.**"*
```bash
kafka-replica-verification.sh \
--broker-list kafka.host1.domain.com:9092,kafka.host2.domain.com:9092 \
--topic-white-list 'my.*'
2021-06-07 03:28:21,829: verification process is started.
2021-06-07 03:28:51,949: max lag is 0 for partition my-topic-0 at offset 4 among 1 partitions
```
> *"you must provide **an EXPLICIT comma-separated list of brokers**... By default, **all topics are validated**; however, you may also provide a regular expression."*
> ### ⚠️ CAUTION: CLUSTER IMPACT AHEAD [#️-caution-cluster-impact-ahead]
>
> *"The replica verification tool will have **an impact on your cluster SIMILAR TO REASSIGNING PARTITIONS**, as it must **READ ALL MESSAGES FROM THE OLDEST OFFSET** in order to verify the replica. In addition, **IT READS FROM ALL REPLICAS FOR A PARTITION IN PARALLEL**, so it should be **used with caution.**"*
***
# 12.12 Self-test (/docs/kafka/administering-kafka/self-test)
Why must you restrict access to the CLI tools even if you've configured ACLs?
Why should tool version match broker version, and what's the safest practice?
Give two topic-naming rules and the concrete failure each prevents.
Which of `--if-exists` / `--if-not-exists` is recommended, and why is the other dangerous?
Order these by severity and say what each means for producers and consumers: under-replicated, at-min-ISR, under-min-ISR, unavailable.
Why are URPs "not necessarily bad"? What should you actually alert on?
Interpret `Replicas: 0,1 Isr: 0` with min-ISR 1. What happens if broker 0 now fails?
Two reasons to add partitions. What breaks when you do, and what's the standing advice?
Why can't you reduce partitions? Give both workarounds.
Why is deleting an unused topic worth doing at all?
What must be true for `--delete` to work, and how do you confirm it succeeded?
Why should you delete only one or two topics at a time?
Define `CURRENT-OFFSET`, `LOG-END-OFFSET`, and `LAG` precisely.
What must be true to delete a consumer group? What's the alternative if you only want to clear one topic's offsets?
Give the safe six-step offset-reset workflow. What single flag separates export from destruction?
Why won't imported offsets take effect if consumers are running?
Explain the per-broker quota trap with the 5-broker / 10 MBps example.
Why should each consumer group have a unique `client.id`?
What does `--describe` on a topic config *not* show, and which API does show it?
Which topic config fixes Ch. 2's synchronized-segment-roll storm? Which two govern GDPR compaction timing?
What does `message.downconversion.enable` let you do, and what Ch. 6 problem does that address?
Name three notable broker-level dynamic configs. Why does one of them matter for incident response?
Why does leadership get skewed after a rolling restart, and what are the three remedies?
Distinguish `--property`, `--producer-property`, and `--consumer-property`.
How does the console producer decide the key? What happens with no separator?
Name the four message formatters and one use for each.
Write the command to inspect `__consumer_offsets`. Name the three non-obvious things it needs.
Walk the three steps of a partition reassignment. Which two files must you save, and why?
What does the controller actually do during a reassignment? What temporarily changes?
What are `--additional`, `--disable-rack-aware`, and `--throttle` for? Which combination is the emergency brake?
Why does moving leadership off a broker *before* reassignment speed things up? What must you disable temporarily?
How do you change a topic's replication factor? Why doesn't changing the cluster default suffice?
Give both caveats of `--cancel`.
What does `kafka-dump-log.sh --print-data-log` show you, and how do its fields map to the v2 batch header from Ch. 6?
Why does replica verification exist — what gap in replication does it detect? What does it cost to run?
When would you delete the `/admin/controller` znode, and what can't you control about the outcome?
Give both scenarios where topic deletion gets stuck, and the two-part fix.
List the four steps of manual topic deletion. What's the prerequisite, and which child-node gotcha will bite you?
**Previous:** [Chapter 11 — Securing Kafka](11-securing-kafka.md)
**Next:** [Chapter 13 — Monitoring Kafka](13-monitoring-kafka.md)
# 12.1 Topic operations — `kafka-topics.sh` (/docs/kafka/administering-kafka/topic-operations-kafka-topics)
> *"allows you to **create, modify, delete, and list** information about topics. **While some topic CONFIGURATIONS are possible through this command, they have been DEPRECATED, and it is recommended to use the more robust method of using the `kafka-config.sh` tool for configuration changes.**"*
#### 1.1 Creating a topic [#11-creating-a-topic]
**Three required arguments** — *"These arguments must be provided **even though some of them may have broker-level defaults configured already.**"*
```bash
kafka-topics.sh --bootstrap-server localhost:9092 --create \
--topic my-topic --replication-factor 2 --partitions 8
# Created topic "my-topic".
```
| Argument | Meaning |
| ---------------------- | ---------------------------------------------------------------------- |
| `--topic` | The name |
| `--replication-factor` | *"The number of replicas of the topic to maintain within the cluster"* |
| `--partitions` | *"The number of partitions to create"* |
**Rack awareness:** *"if the cluster is set up for **rack-aware replica assignment**, the replicas for each partition will be **in separate racks.** If rack-aware assignment is not desired, specify **`--disable-rack-aware`**."*
> ### ⚠️ GOOD TOPIC NAMING PRACTICES [#️-good-topic-naming-practices]
>
> *"Topic names may contain **alphanumeric characters, underscores, dashes, and periods**; however:*
>
> **① DO NOT USE PERIODS.** *"Internal metrics inside of Kafka **CONVERT PERIOD CHARACTERS TO UNDERSCORE CHARACTERS** (e.g., **'topic.1' becomes 'topic\_1' in metrics calculations**), which can result in **CONFLICTS in topic names.**"*
>
> **② DO NOT START WITH A DOUBLE UNDERSCORE.** *"By convention, **topics internal to Kafka operations are created with a double underscore** naming convention (like the `__consumer_offsets` topic)... it is not recommended to have topic names that begin with the double underscore naming convention **to prevent confusion.**"*
> ### `--if-exists` / `--if-not-exists` [#--if-exists----if-not-exists]
>
> * **`--if-not-exists`** with `--create`: *"you may want to use \[it] in automation... **that will not return an error if the topic already exists.**"* ✅
> * **`--if-exists`** with `--alter`: *"**using it is NOT RECOMMENDED. Using this argument will cause the command to not return an error if the topic being changed DOES NOT EXIST. THIS CAN MASK PROBLEMS where a topic does not exist that should have been created.**"* ❌
#### 1.2 Listing topics [#12-listing-topics]
```bash
kafka-topics.sh --bootstrap-server localhost:9092 --list
# __consumer_offsets
# my-topic
# other-topic
```
> *"formatted with **one topic per line, in NO PARTICULAR ORDER**."* Add **`--exclude-internal`** to *"remove all topics from the list that begin with the double underscore."*
#### 1.3 Describing topics [#13-describing-topics]
```bash
kafka-topics.sh --bootstrap-server localhost:9092 --describe --topic my-topic
Topic: my-topic PartitionCount: 8 ReplicationFactor: 2 Configs: segment.bytes=1073741824
Topic: my-topic Partition: 0 Leader: 1 Replicas: 1,0 Isr: 1,0
Topic: my-topic Partition: 1 Leader: 0 Replicas: 0,1 Isr: 0,1
...
```
> Output includes *"the **partition count, topic CONFIGURATION OVERRIDES**, and a listing of **each partition with its replica assignments**."*
*(Recall Ch. 6 §4.4: **the first replica in the `Replicas` list is the preferred leader** — here p0 prefers broker 1, p1 prefers broker 0. That alternation is what balanced leadership looks like.)*
#### 💡 1.4 The diagnostic filters — the most useful part of the whole chapter [#-14-the-diagnostic-filters--the-most-useful-part-of-the-whole-chapter]
> *"These can be helpful for **diagnosing cluster issues** more easily. For these commands **we generally do NOT specify the `--topic` argument** because the intention is to find **all** topics or partitions that match the criteria. **These options will NOT work with the `--list` command.**"*
**The escalating severity ladder — memorize this:**
**Worked example** — min-ISR 1, RF 2, host 0 up, host 1 down for maintenance:
```bash
kafka-topics.sh --bootstrap-server localhost:9092 --describe --at-min-isr-partitions
Topic: my-topic Partition: 0 Leader: 0 Replicas: 0,1 Isr: 0
...
```
Note `Replicas: 0,1` but `Isr: 0` — **exactly at min-ISR, zero redundancy remaining.**
#### 1.5 Adding partitions [#15-adding-partitions]
**Why:** *"the most common reason is to **horizontally scale a topic across more brokers by decreasing the throughput for a single partition.** Topics may also be increased **if a consumer needs to expand to run more copies in a single consumer group** since a partition can only be consumed by a single member in the group."*
```bash
kafka-topics.sh --bootstrap-server localhost:9092 \
--alter --topic my-topic --partitions 16
```
> ### ⚠️ ADJUSTING KEYED TOPICS [#️-adjusting-keyed-topics]
>
> *"Topics that are produced with **keyed messages can be VERY DIFFICULT to add partitions to from a consumer's point of view.** This is because **THE MAPPING OF KEYS TO PARTITIONS WILL CHANGE** when the number of partitions is changed. For this reason, **it is advisable to SET THE NUMBER OF PARTITIONS ONCE, WHEN THE TOPIC IS CREATED, AND AVOID RESIZING.**"*
*(Third appearance of this warning — Ch. 3 §9.4, Ch. 5 §7.2, and here. Kafka's authors really mean it.)*
#### 1.6 ⚠️ Reducing partitions is impossible [#16-️-reducing-partitions-is-impossible]
> *"**It is NOT POSSIBLE to reduce the number of partitions.** Deleting a partition would cause **part of the data in that topic to be deleted as well, which would be INCONSISTENT from a client point of view.** In addition, trying to redistribute the data to the remaining partitions would be **difficult and result in OUT-OF-ORDER messages.**"*
**The two workarounds:**
#### 1.7 Deleting a topic [#17-deleting-a-topic]
**Why bother deleting empty topics:**
> *"Even a topic with **no messages** uses cluster resources such as **disk space, open filehandles, and memory. THE CONTROLLER ALSO HAS JUNK METADATA that it must retain knowledge of, WHICH CAN HINDER PERFORMANCE AT LARGE SCALE.**"*
```bash
kafka-topics.sh --bootstrap-server localhost:9092 --delete --topic my-topic
# Note: This will have no impact if delete.topic.enable is not set to true.
```
**Prerequisite:** `delete.topic.enable=true` on the brokers. *"If it's set to `false`, then the request to delete the topic **will be IGNORED and will not succeed.**"*
**It's asynchronous, and there's a rate limit you should respect:**
> ### ⚠️ DATA LOSS AHEAD [#️-data-loss-ahead]
>
> *"Deleting a topic will also delete all its messages. **THIS IS NOT A REVERSIBLE OPERATION.** Make sure it is executed carefully."*
**And there is no success feedback:** *"You will notice **there is NO VISIBLE FEEDBACK** that the topic deletion was completed successfully or not. **Verify that deletion was successful by running the `--list` or `--describe` options.**"*
***
# 12.0 Two warnings before you run anything (/docs/kafka/administering-kafka/two-warnings-before-run)
> ### ⚠️ AUTHORIZING ADMIN OPERATIONS [#️-authorizing-admin-operations]
>
> *"While Apache Kafka implements authentication and authorization to control topic operations, **DEFAULT CONFIGURATIONS DO NOT RESTRICT THE USE OF THESE TOOLS.** This means that **these CLI tools can be used WITHOUT ANY AUTHENTICATION REQUIRED, which will allow operations such as topic changes to be executed WITH NO SECURITY CHECK OR AUDIT. ALWAYS ENSURE THAT ACCESS TO THIS TOOLING ON YOUR DEPLOYMENTS IS RESTRICTED TO ADMINISTRATORS ONLY** to prevent unauthorized changes."*
*(Cross-reference Ch. 11: ACLs govern the *protocol*, but tools that write directly to ZooKeeper bypass the authorizer entirely. Shell access to a broker host **is** admin access.)*
> ### ⚠️ CHECK THE VERSION [#️-check-the-version]
>
> \*"Many of the command-line tools have **a dependency on the version of Kafka running** to operate correctly. This includes some commands that may **store data in ZooKeeper rather than connecting to the brokers themselves.** For this reason, it is important to make sure **the version of the tools you are using MATCHES the version of the brokers in the cluster.**
>
> **The safest approach is to RUN THE TOOLS ON THE KAFKA BROKERS THEMSELVES, using the deployed version.**"\*
And later, more bluntly: *"**Older console consumers can potentially DAMAGE THE CLUSTER by interacting with the cluster or ZooKeeper in incorrect ways.**"*
***
# 12.9 ⚠️ Unsafe operations — "Here be dragons" (/docs/kafka/administering-kafka/unsafe-operations-here-dragons)
> *"There are some administrative tasks that are **technically possible to do but SHOULD NOT BE ATTEMPTED EXCEPT IN THE MOST EXTREME SITUATIONS.** Often this is when you are **diagnosing a problem and have RUN OUT OF OPTIONS**, or you have found a specific bug that you need to work around temporarily. **These tasks are usually UNDOCUMENTED, UNSUPPORTED, and POSE SOME AMOUNT OF RISK.**"*
> ### ⚠️ DANGER: HERE BE DRAGONS [#️-danger-here-be-dragons]
>
> *"The operations in this section often involve **working with the cluster metadata stored in ZooKeeper DIRECTLY. This can be a VERY DANGEROUS operation, so you must be very careful to NOT MODIFY the information in ZooKeeper directly, EXCEPT AS NOTED.**"*
#### 9.1 Moving the cluster controller [#91-moving-the-cluster-controller]
**When:** *"when troubleshooting a misbehaving cluster or broker, it may be useful to **forcibly move the controller to a different broker WITHOUT SHUTTING DOWN THE HOST.** One such example is **when the controller has suffered AN EXCEPTION or other problem that has left it RUNNING BUT NOT FUNCTIONAL.**"*
**How:** *"deleting the ZooKeeper znode at **`/admin/controller`** manually will cause **the current controller to RESIGN**, and the cluster will **randomly select a new controller.**"*
> ⚠️ *"**There is currently NO WAY to specify a SPECIFIC broker to be controller in Apache Kafka.**"*
> *"Moving the controller in these situations **does not normally have a high risk, but as it is not a normal task, it should not be performed regularly.**"*
*(Ch. 6 §2.2: the resigning controller's epoch is superseded, so its stale messages get fenced. That's why this is relatively safe.)*
#### 9.2 Unsticking topic deletion [#92-unsticking-topic-deletion]
**Two scenarios where deletion gets stuck:**
**The fix:**
#### 9.3 Deleting topics manually [#93-deleting-topics-manually]
**When:** *"If you are running a cluster with delete topics disabled, or if you find yourself needing to delete some topics outside of the normal flow."*
> ### ⚠️ SHUT DOWN BROKERS FIRST [#️-shut-down-brokers-first]
>
> *"Modifying the cluster metadata in ZooKeeper **when the cluster is ONLINE is a VERY DANGEROUS operation and can put the cluster into an UNSTABLE STATE. NEVER attempt to delete or modify topic metadata in ZooKeeper WHILE THE CLUSTER IS ONLINE.**"*
*(Requires **full cluster downtime**. This is the definition of a last resort.)*
***
# 15.1 A.0 The core principle (/docs/kafka/appendices-b/0-core-principle)
> \*"Apache Kafka is **primarily a Java application** and therefore should be able to run on **any system where you are able to install a JRE. IT HAS, HOWEVER, BEEN OPTIMIZED FOR LINUX-BASED OPERATING SYSTEMS, SO THAT IS WHERE IT WILL PERFORM BEST. RUNNING ON OTHER OPERATING SYSTEMS MAY RESULT IN BUGS SPECIFIC TO THE OS.**
>
> For this reason, **when using Kafka for development or test purposes on a common desktop OS, IT IS A GOOD IDEA TO CONSIDER RUNNING IN A VIRTUAL MACHINE THAT MATCHES YOUR EVENTUAL PRODUCTION ENVIRONMENT.**"\*
***
# 15.2 A.1 Windows (/docs/kafka/appendices-b/1-windows)
Two options as of Windows 10 — and the book is unambiguous about which to prefer:
#### Option 1 (preferred): Windows Subsystem for Linux (WSL) [#option-1-preferred-windows-subsystem-for-linux-wsl]
> *"you can install native Ubuntu support under Windows using **Windows Subsystem for Linux (WSL).** At the time of publication, **Microsoft still considers WSL to be an EXPERIMENTAL feature.** Though it acts similar to a virtual machine, **it DOES NOT REQUIRE THE RESOURCES OF A FULL VM and provides RICHER INTEGRATION with the Windows OS.**"*
> *"**The latter method is HIGHLY PREFERRED because it provides A MUCH SIMPLER SETUP THAT MORE CLOSELY MATCHES THE TYPICAL PRODUCTION ENVIRONMENT.**"*
```bash
# after installing WSL per Microsoft's "What Is the Windows Subsystem for
# Linux?" page, and installing the Ubuntu system package:
sudo apt install openjdk-16-jre-headless
```
> *"Once you have installed the JDK, you can proceed to install Apache Kafka **using the instructions in Chapter 2**."* — i.e. from here it is just Linux.
#### Option 2: native Java on Windows [#option-2-native-java-on-windows]
> ⚠️ *"For **older versions of Windows**, or if you prefer not to use WSL... **BE AWARE, HOWEVER, THAT THIS CAN INTRODUCE BUGS SPECIFIC TO THE WINDOWS ENVIRONMENT. THESE BUGS MAY NOT GET THE ATTENTION IN THE APACHE KAFKA DEVELOPMENT COMMUNITY AS SIMILAR PROBLEMS ON LINUX MIGHT.**"*
**Step 1 — Java.** *"install the latest version of **Oracle Java 16**... **Download a FULL JDK package** so that you have all the Java tools available."*
> ### ⚠️ BE CAREFUL WITH PATHS [#️-be-careful-with-paths]
>
> \*"it is **HIGHLY RECOMMENDED that you stick to installation paths that DO NOT CONTAIN SPACES. While Windows ALLOWS spaces in paths, APPLICATIONS THAT ARE DESIGNED TO RUN IN UNIX ENVIRONMENTS ARE NOT SET UP THIS WAY, and specifying paths will be difficult.**
>
> For example, if installing JDK 16.0.1, a good choice would be **`C:\Java\jdk-16.0.1`**."\*
**Step 2 — environment variables.** In Windows 10: *"System and Security" → System → "Advanced system settings" → Advanced tab → "Environment Variables"*
```txt
NEW USER VARIABLE: JAVA_HOME =
EDIT SYSTEM Path: add an entry %JAVA_HOME%\bin
```
**Step 3 — Kafka.** *"The installation **includes ZooKeeper**, so you do not have to install it separately."* At publication: **2.8.0 on Scala 2.13.0**. *"The downloaded file will be **gzip compressed and packaged with the tar utility**, so you will need to use a Windows application such as **8 Zip** to uncompress it."* Example target: `C:\kafka_2.13-2.8.0`.
**Step 4 — running it.** ⚠️ Two differences from Linux:
```powershell
# SHELL 1 — ZooKeeper
PS C:\> cd kafka_2.13-2.8.0
PS C:\kafka_2.13-2.8.0> bin\windows\zookeeper-server-start.bat `
C:\kafka_2.13-2.8.0\config\zookeeper.properties
# ... INFO PrepRequestProcessor (sid:0) started, reconfigEnabled=false
# SHELL 2 — Kafka
PS C:\> cd kafka_2.13-2.8.0
PS C:\kafka_2.13-2.8.0> .\bin\windows\kafka-server-start.bat `
C:\kafka_2.13-2.8.0\config\server.properties
# ... INFO [KafkaServer id=0] started (kafka.server.KafkaServer)
```
*(Note the last log line in the book's output: `Recorded new controller, from now on will use broker ...` — that's Ch. 6's controller registration, visible at startup.)*
***
# 15.3 A.2 macOS (/docs/kafka/appendices-b/2-macos)
> *"macOS runs on **Darwin, a Unix OS that is derived, in part, from FreeBSD.** This means that **many of the expectations of running on a Unix OS hold true**, and installing applications designed for Unix, like Apache Kafka, **is not too difficult.**"*
Two options: **Homebrew** (simple) or **manual** (*"for greater control over versions"*).
#### Option 1: Homebrew — one step [#option-1-homebrew--one-step]
```bash
brew install kafka
# ==> Installing dependencies for kafka: openjdk, openssl@1.1 and zookeeper
# ==> /usr/local/Cellar/kafka/2.8.0: 200 files, 68.2MB
```
> *"The Homebrew package manager will **ensure that you have all the dependencies installed first, INCLUDING JAVA.**"*
**Where things land** — *"Homebrew will install Kafka under `/usr/local/Cellar`, but the files will be **linked into other directories**"*:
```bash
/usr/local/bin/zkServer start
# Using config: /usr/local/etc/zookeeper/zoo.cfg
# Starting zookeeper ... STARTED
/usr/local/bin/kafka-server-start /usr/local/etc/kafka/server.properties
# ... INFO [KafkaServer id=0] started
```
*(This example starts Kafka **in the foreground**.)*
#### Option 2: manual [#option-2-manual]
Same shape as the Windows manual install: download the JDK from the Oracle Java SE page, download Kafka, expand it (example: `/usr/local/kafka_2.13-2.8.0`).
> *"Starting ZooKeeper and Kafka **looks just like starting them when using Linux**, though **you will need to make sure your `JAVA_HOME` directory is set first**:"*
```bash
export JAVA_HOME=`/usr/libexec/java_home -v 16.0.1`
echo $JAVA_HOME
# /Library/Java/JavaVirtualMachines/jdk-16.0.1.jdk/Contents/Home
/usr/local/kafka_2.13-2.8.0/bin/zookeeper-server-start.sh -daemon \
/usr/local/kafka_2.13-2.8.0/config/zookeeper.properties
/usr/local/kafka_2.13-2.8.0/bin/kafka-server-start.sh \
/usr/local/kafka_2.13-2.8.0/config/server.properties
```
💡 **`/usr/libexec/java_home -v `** is the macOS-idiomatic way to resolve a JDK path — worth knowing, since hardcoding the JDK path breaks on every Java update.
#### Summary table [#summary-table]
| Platform | Method | Verdict / notes |
| ------------------- | ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Windows 10+** | **WSL** | ✅ **Highly preferred** — *"much simpler setup that more closely matches the typical production environment"* |
| **Windows (older)** | Native Java | ⚠️ May hit Windows-specific bugs that *"may not get the attention"* upstream. No path spaces. `.bat` files. **One shell per process.** |
| **macOS** | **Homebrew** | Simplest; pulls Java in automatically; links into `/usr/local/{bin,etc,var}` |
| **macOS** | Manual | *"greater control over versions"*; set `JAVA_HOME` via `java_home -v` |
| **Any desktop OS** | **A VM matching production** | ✅ The appendix's own top-level recommendation |
***
## Appendix B — Additional Kafka Tools [#appendix-b--additional-kafka-tools]
> *"The Apache Kafka community has created **a robust ecosystem of tools and platforms** that make the task of running and using Kafka far easier."*
> ### CAVEAT EMPTOR [#caveat-emptor]
>
> *"While **the authors are affiliated with some of the companies and projects** that are included in this list, **neither they nor O'Reilly specifically endorse one tool over others. Please be sure to DO YOUR OWN RESEARCH on the suitability of these platforms and tools for the work that you need to do.**"*
***
# 15.4 B.1 Comprehensive platforms (managed Kafka) (/docs/kafka/appendices-b/b-1-comprehensive-platforms)
> *"This includes **managed deployments of ALL components**, such that **you can focus on USING Kafka and not on HOW TO RUN IT.** This can present an ideal solution for use cases where **resources are not available (or you do not want to dedicate them) for learning how to properly operate Kafka and the infrastructure required around it.** Several also provide tools, such as **schema management, REST interfaces, and in some cases client library support**, so that you can be assured **components interoperate correctly.**"*
| Platform | Clouds | Schema registry | REST proxy | Notes |
| ------------------- | ------------------------------------------ | ----------------------------- | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Confluent Cloud** | AWS, Azure, GCP | ✅ | ✅ | *"the company created by some of the original developers"*; *"combines a number of MUST-HAVE tools — including schema management, clients, a RESTful interface, and monitoring — into a single offering"*; *"backed with support provided by **a sizable portion of the core Apache Kafka contributors** employed by Confluent"*. ⚠️ Many components *"are available as standalone tools under the **Confluent Community License, WHICH DOES RESTRICT SOME USE CASES**"* |
| **Aiven** | AWS, Azure, GCP, **DigitalOcean, UpCloud** | ✅ **Karapace** | ✅ Karapace | Karapace is *"API-compatible with Confluent's components **but supported under the Apache 2.0 license, WHICH DOES NOT RESTRICT USE CASES**"* |
| **CloudKarafka** | AWS, GCP | ⚠️ Confluent's, **v5.0 only** | ⚠️ same | *"focuses on providing a managed Kafka solution with **integrations for popular infrastructure services (such as DataDog or Splunk)**"*; supports Confluent's components *"but **ONLY THE 5.0 VERSION PRIOR TO THE LICENSE CHANGES**"* |
| **Amazon MSK** | AWS only | ✅ via **AWS Glue** | ❌ *"not directly supported"* | *"Amazon promotes the use of **community tools (such as Cruise Control, Burrow, and Confluent's REST proxy) BUT DOES NOT DIRECTLY SUPPORT THEM.** As such, **MSK is somewhat LESS INTEGRATED than other offers, but can still provide a core Kafka cluster.**"* |
| **Azure HDInsight** | Azure | ❌ user-provided | ❌ user-provided | *"also supports **Hadoop, Spark, and other big data components**"*; *"focuses on the core Kafka cluster, leaving many of the other components... **up to the user to provide.** Some third parties have provided templates... **but they are NOT SUPPORTED by Microsoft.**"* |
| **Cloudera** | public cloud **and private** | (part of CDP) | (part of CDP) | *"a fixture in the Kafka community **since the early days**"*; Kafka as *"the stream data component of its overall **Customer Data Platform (CDP)**"*; *"**operates in the public cloud environments AS WELL AS providing PRIVATE options.**"* |
***
# 15.5 B.2 Cluster deployment and management (/docs/kafka/appendices-b/b-2-cluster-deployment)
> *"When running Kafka **outside of a managed platform**, you will need several things to assist you with running the cluster properly. This includes help with **provisioning and deployment, balancing data, and visualizing your clusters.**"*
| Tool | What it does |
| ---------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Strimzi** — `strimzi.io` | *"provides **KUBERNETES OPERATORS** for deploying Kafka clusters... **It does NOT provide managed services** but instead makes it easy for you to get up and running in a cloud, **whether public or private.** It also provides the **Strimzi Kafka Bridge**, a REST proxy implementation supported under the **Apache 2.0 license.** ⚠️ **At this time, Strimzi does NOT have support for a schema registry, DUE TO CONCERNS ABOUT LICENSES.**"* |
| **AKHQ** — `akhq.io` | *"a **GUI** for managing and interacting with Kafka clusters. It supports **configuration management, INCLUDING USERS AND ACLs**, and provides some support for components like the **Schema Registry and Kafka Connect.** It also provides tools for **working with data in the cluster AS AN ALTERNATIVE TO THE CONSOLE TOOLS.**"* |
| **JulieOps** (formerly Kafka Topology Builder) | *"automated management of **topics and ACLs using a GITOPS MODEL. More than viewing the state of the current configuration, JulieOps provides a means for DECLARATIVE CONFIGURATION AND CHANGE CONTROL of topics, schemas, ACLs, and more OVER TIME.**"* |
| **Cruise Control** — LinkedIn | ⭐ \*"LinkedIn's answer to how to manage **hundreds of clusters with thousands of brokers.** This tool began as a solution to **automated rebalancing of data** but has evolved to include **ANOMALY DETECTION and ADMINISTRATIVE OPERATIONS, such as adding and removing brokers.* **FOR ANYTHING MORE THAN A TESTING CLUSTER, IT IS A MUST-HAVE FOR ANY KAFKA OPERATOR.**"* |
| **Conduktor** — `conduktor.io` | *"**not open source**, \[but] a popular **DESKTOP** tool... supports many of the managed platforms (Confluent, Aiven, MSK) and many different components (Connect, kSQL, Streams). Also allows you to **interact with data in the clusters**, as opposed to using the console tools. **A FREE LICENSE is provided for development use that works with a SINGLE CLUSTER.**"* |
***
# 15.6 B.3 Monitoring and data exploration (/docs/kafka/appendices-b/b-3-monitoring-data)
> *"A critical part to running Kafka is to ensure that **your cluster AND your clients are healthy.** Like many applications, Kafka **exposes numerous metrics and other telemetry, BUT MAKING SENSE OF IT CAN BE CHALLENGING.** Many of the larger monitoring platforms (such as **Prometheus**) can easily fetch metrics from Kafka brokers and clients."*
| Tool | Purpose |
| ------------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Xinfra Monitor** (formerly Kafka Monitor) — LinkedIn | ⭐ *"monitor **AVAILABILITY** of Kafka clusters and brokers. It does this by using **a set of topics to generate SYNTHETIC DATA through the cluster** and measuring **latency, availability, and completeness. It's a valuable tool for measuring your Kafka deployment's health WITHOUT REQUIRING DIRECT INTERACTION WITH YOUR CLIENTS.**"* (Ch. 13 §9) |
| **Burrow** — LinkedIn | ⭐ *"**holistic monitoring of CONSUMER LAG**... provides **a view into the health of the consumers WITHOUT NEEDING TO DIRECTLY INTERACT WITH THEM.** Burrow is **actively supported by the community and has its OWN ECOSYSTEM OF TOOLS** to connect it with other components."* (Ch. 13 §8) |
| **Kafka Dashboard** — DataDog | *"an excellent Kafka Dashboard to give you **a head start** on integrating Kafka clusters into your monitoring stack. Designed to provide **a SINGLE-PANE view**, simplifying the view of many metrics."* |
| **Streams Explorer** — bakdata | *"visualizing the **flow of data through applications and connectors in a KUBERNETES deployment.** ⚠️ **It heavily relies on structuring your deployments using either Kafka Streams or Faust through bakdata's tools**, but it can then provide **an easily comprehensible view of those applications and their metrics.**"* |
| **kcat** (formerly kafkacat) | 💡 *"a **much-loved** alternative to the console producer and consumer... **It is SMALL, FAST, and WRITTEN IN C, SO IT DOES NOT HAVE JVM OVERHEAD.** It also supports **limited views into cluster status by showing METADATA OUTPUT** for the cluster."* |
*(**Xinfra Monitor** and **Burrow** are the two external monitors Ch. 13 says you cannot do without — because the broker fundamentally cannot report whether clients can use it, or how far behind consumers are.)*
***
# 15.7 B.4 Client libraries (/docs/kafka/appendices-b/b-4-client-libraries)
> *"The Apache Kafka project provides client libraries for **Java** applications, **but ONE LANGUAGE IS NEVER ENOUGH.** There are many implementations out there, with popular languages such as **Python, Go, and Ruby** having several options. In addition, **REST proxies (such as those from Confluent, Strimzi, or Karapace) can cover a variety of use cases.** Here are a few client implementations that **have stood the test of time.**"*
| Library | Language | License | Notes |
| -------------------- | ------------------- | ------------------ | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **librdkafka** | **C** | **two-clause BSD** | ⭐ *"regarded as **ONE OF THE BEST-PERFORMING LIBRARIES available. SO GOOD, IN FACT, THAT CONFLUENT SUPPORTS CLIENTS FOR GO, PYTHON, AND .NET THAT IT CREATED AS WRAPPERS AROUND librdkafka.** ... the license **makes it easy to use in any application.**"* |
| **Sarama** — Shopify | **Go** (native) | **MIT** | *"a native Golang implementation"* |
| **kafka-python** | **Python** (native) | **Apache 2.0** | *"another native client implementation"* |
***
# 15.8 B.5 Stream processing frameworks (/docs/kafka/appendices-b/b-5-stream-processing)
> *"While the Apache Kafka project includes Kafka Streams for building applications, **it's not the only choice out there.**"*
| Framework | Runs on | Character |
| ---------------- | --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Apache Samza** | **YARN** | *"specifically designed for Kafka. **While it PREDATES Kafka Streams, it was developed by MANY OF THE SAME PEOPLE, and as a result the two SHARE MANY CONCEPTS.** However, unlike Kafka Streams, **Samza runs on Yarn and provides A FULL FRAMEWORK for applications to run in.**"* |
| **Apache Spark** | — | *"**oriented toward BATCH processing.** It handles streams **by considering them to be FAST MICROBATCHES.** This means **the LATENCY IS A LITTLE HIGHER, BUT FAULT TOLERANCE IS SIMPLY HANDLED THROUGH REPROCESSING BATCHES, and LAMBDA ARCHITECTURE IS EASY.** It also has the benefit of **wide community support.**"* |
| **Apache Flink** | YARN, **Mesos, Kubernetes, standalone** | *"**specifically oriented toward stream processing and operates with VERY LOW LATENCY.** ... It also supports **Python and R** with provided high-level APIs."* |
| **Apache Beam** | Samza, Spark, Flink as **runners** | *"**doesn't provide stream processing DIRECTLY** but instead promotes itself as **A UNIFIED PROGRAMMING MODEL FOR BOTH BATCH AND STREAM PROCESSING.** It utilizes platforms like Samza, Spark, and Flink **as RUNNERS** for components in an overall processing pipeline."* |
**Mapping these onto Ch. 14 §7's selection criteria:**
***
# 15.9 Consolidated tool recommendations by job (/docs/kafka/appendices-b/consolidated-tool-recommendations-job)
***
# 15. Appendices A & B (/docs/kafka/appendices-b)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Appendix A & B*
***
## Appendix A — Installing Kafka on Other Operating Systems [#appendix-a--installing-kafka-on-other-operating-systems]
# 15.10 Self-test (/docs/kafka/appendices-b/self-test)
Kafka runs anywhere a JRE runs. So what's the actual recommendation for a desktop OS, and why?
Which Windows option does the book prefer, and give its two stated advantages.
Name three concrete gotchas of the native-Windows install.
Why does the appendix warn about spaces in install paths?
Where does Homebrew put Kafka's binaries, configs, and **data**?
What's the macOS-idiomatic way to set `JAVA_HOME`, and why is it better than hardcoding?
What's the licensing distinction between Confluent's Schema Registry / REST proxy and Aiven's Karapace? Why might it decide your platform choice?
Which managed platforms leave the schema registry and REST proxy up to you?
Which tool does the book call a "must-have for any Kafka operator," and name three jobs it does that Kafka won't do itself?
What does Strimzi provide, and what does it deliberately *not* provide — and why?
What's JulieOps' model, and what does it manage?
Which two external monitors does the book consider essential, and what does each measure that broker metrics cannot?
Why is kcat "much-loved" over the console tools?
Confluent's Go, Python, and .NET clients are wrappers around what? What's the trade-off versus a native client like Sarama?
Contrast Samza, Spark, Flink, and Beam on execution model and latency character.
Which framework would you pick for a low-millisecond fraud-detection action, and which would you specifically avoid? Why?
State the architectural divide between Kafka Streams and the other three frameworks, and the Ch. 1 design stance behind it.
**Previous:** [Chapter 14 — Stream Processing](14-stream-processing.md)
**Back to the index:** [README](README.md)
# 9.10 What actually breaks in production — Ch. 9 consolidated (/docs/kafka/building-data-pipelines-kafka/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ----------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------- |
| 1 | **A tangle of bespoke pipelines nobody can maintain; adopting a new system takes a quarter** | **Ad hoc pipelines** — one custom tool per pair of endpoints | Single integration substrate (Connect); think about the whole graph, not the immediate hop |
| 2 | **A DBA adds a column and every downstream app breaks (or all must deploy together)** | **Loss of metadata** — no schema propagation or evolution | Schema Registry + a converter that carries schemas |
| 3 | **Downstream team needs a field the pipeline dropped years ago; historical data unrecoverable** | **ETL over-processing**; *"historical data will require reprocessing (assuming it is available)"* | ELT-lean: only transform what benefits *every* consumer |
| 4 | **Pipeline requires constant changes as downstream needs shift** | **Extreme processing** coupling downstream systems to pipeline-time decisions | Preserve raw data; let Streams apps decide |
| 5 | **Credentials leaked from a Connect config file** | Connector configs contain credentials in plaintext | **External secret config providers** (Vault / AWS / Azure) |
| 6 | **Bad records discovered days later; can't reprocess** | Retention shorter than bug-discovery latency | *"Kafka can be configured to store all events for long periods"* — size retention to your detection latency |
| 7 | **Connect worker OOM / broker page-cache contention** | Connect running **on the broker machines** | *"run Connect on SEPARATE SERVERS from your Kafka brokers"* |
| 8 | **Connector plug-in not found** | Dependencies placed at the **top level** of `plugin.path` instead of in a per-connector subdirectory | One subdirectory per connector, containing the jar **and all its dependencies** |
| 9 | **Bizarre `NoSuchMethodError` / version conflicts** | Connectors added to the **Kafka Connect classpath**, bringing a dependency that conflicts with Kafka's | Use `plugin.path`, *"the recommended approach"* |
| 10 | **JDBC connector fails: "Access denied" / driver not found** | Driver missing (**doesn't ship with the connector for license reasons**), or table permissions | Download the MySQL driver into `/opt/connectors/jdbc`; **check the Connect worker log** |
| 11 | **Deletes never appear downstream; some updates missing** | **JDBC polling CDC** — scans by timestamp / incrementing PK; *"relatively inefficient and at times inaccurate"* | **Debezium** (reads the binlog/WAL directly) |
| 12 | **Database load spikes from the pipeline** | JDBC connector repeatedly scanning tables | Log-based CDC |
| 13 | **Duplicate documents in Elasticsearch after reprocessing** | Kafka records had **null keys** (JDBC doesn't populate them) and the sink generated new IDs | `key.ignore=true` → deterministic `topic+partition+offset` document ID |
| 14 | **File-based pipeline loses data** | **FileStream connectors** — *"many limitations and NO reliability guarantees"* | FilePulse / FileSystem Connector / SpoolDir |
| 15 | **Corrupt messages halt a sink connector** | No error tolerance configured | **`error.tolerance`** → silently drop, or route to a **dead letter queue** |
| 16 | **Syslog connector randomly stops receiving data** | Connector needs to listen on a **specific machine's port**, but distributed mode may schedule tasks **on any node** | **Standalone mode** for machine-pinned connectors |
| 17 | **Only one task runs despite `tasks.max=10`** | JDBC connector uses `MIN(tasks.max, number_of_tables)` | Understand the connector's own splitting logic |
| 18 | **Connector work distributed unevenly** | Connectors and tasks *"may start on any node"* | Worker rebalancing handles it; inspect via REST API |
| 19 | **Source connector reprocesses everything after a crash** | Logical offsets not stored / mis-designed | The framework stores offsets **after** broker ack — a connector authoring bug if it recurs |
| 20 | **A single-partition offset/config topic became a bottleneck or lost ordering** | Internal topics misconfigured | 1 partition + 3 replicas + compaction (Ch. 5 §4.2) |
| 21 | **Consumers can't parse Connect output** | JSON converter's `schemas.enable` mismatch, or Avro registry URL missing | Prefix converter params correctly (`key.converter.` / `value.converter.`) |
| 22 | **Headers not visible in console consumer** | Requires **Apache Kafka 2.7+** and `--property print.headers=true` | Upgrade / add the flag |
| 23 | **Sink writes overwhelm the target system** | No back pressure | Sink context provides back-pressure methods — a well-written connector uses them |
| 24 | **Pipeline can't guarantee exactly-once into a database** | Kafka transactions can't span systems (Ch. 8) | Use the sink context's **external offset storage** — commit data + offsets in the target's transaction |
***
# 9.9 Alternatives to Kafka Connect (/docs/kafka/building-data-pipelines-kafka/alternatives-kafka-connect)
| Alternative | When it makes sense | The drawback |
| ----------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Ingest frameworks for other datastores** — Flume (Hadoop), Logstash / Fluentd (Elasticsearch) | *"If you are actually building a **Hadoop-centric or Elastic-centric system** and Kafka is **just one of many inputs** into that system"* | *"We recommend Kafka's Connect API when **Kafka is an integral part of the architecture** and when the goal is to **connect large numbers of sources and sinks.**"* |
| **GUI-based ETL tools** — Informatica, Talend, Pentaho, Apache NiFi, StreamSets | *"if you are **already using these systems**... you may not be interested in adding another data integration system just for Kafka. They also make sense if you are using **a GUI-based approach**."* | *"usually built for **involved workflows** and will be **a somewhat heavy and involved solution if all you want to do is get data in and out of Kafka.** We believe that **data integration should focus on FAITHFUL DELIVERY OF MESSAGES UNDER ALL CONDITIONS, while most ETL tools add UNNECESSARY COMPLEXITY.**"* |
| **Stream processing frameworks** | *"If your destination system is supported and you **already intend to use that framework** to process events from Kafka"* — *"often saves a step (no need to store processed events in Kafka — just read them out and write them to another system)"* | *"it can be **more difficult to troubleshoot things like LOST AND CORRUPTED MESSAGES.**"* |
**The bigger framing:**
> *"We do encourage you to look at Kafka as a platform that can handle **data integration (with Connect), application integration (with producers and consumers), and stream processing.** **Kafka could be a viable REPLACEMENT for an ETL tool that only integrates data stores.**"*
***
# 9.3 Connect vs producer/consumer clients — the decision rule (/docs/kafka/building-data-pipelines-kafka/connect-vs-producer-consumer)
#### And if no connector exists yet? Still prefer Connect. [#and-if-no-connector-exists-yet-still-prefer-connect]
> *"**Connect is recommended** because it provides out-of-the-box features like:*
>
> * *configuration management*
> * *offset storage*
> * *parallelization*
> * *error handling*
> * *support for different data types*
> * *standard management REST APIs*
>
> ***Writing a small app that connects Kafka to a datastore SOUNDS SIMPLE, but there are MANY LITTLE DETAILS you will need to handle** concerning data types and configuration that make the task **nontrivial. What's more, you will need to MAINTAIN this pipeline app and DOCUMENT it, and your TEAMMATES WILL NEED TO LEARN HOW TO USE IT.**"*
**The most persuasive version of this argument appears later in the chapter:**
> *"Experienced developers know that **writing code that reads data from Kafka and inserts it into a database takes maybe A DAY OR TWO**, but if you need to handle **configuration, errors, REST APIs, monitoring, deployment, scaling up and down, and handling failures, IT CAN TAKE A FEW MONTHS to get everything right.** And most data integration pipelines involve **more than just the one source or target.** So now consider that effort spent on bespoke code for just a database integration, **repeated many times for other technologies.**"*
***
# 9.8 A deeper look at Connect internals (/docs/kafka/building-data-pipelines-kafka/deeper-look-connect-internals)
#### 8.1 Connectors vs tasks — a clean separation [#81-connectors-vs-tasks--a-clean-separation]
**The JDBC source example, concretely:**
> ⚠️ *"Note that when you start the connector via the REST API, **it may start on ANY node, and subsequently the tasks it starts may ALSO execute on ANY node.**"*
**"Storing offsets externally for exactly-once delivery"** is the concrete API behind §2.2's claim — the sink context is what lets a connector commit data + offsets in the *target system's* transaction.
#### 8.2 Workers — where all the hard operational work lives [#82-workers--where-all-the-hard-operational-work-lives]
**Worker responsibilities:**
> *"**The best way to understand workers is to realize that CONNECTORS AND TASKS are responsible for the 'MOVING DATA' part of data integration, while the WORKERS are responsible for the REST API, CONFIGURATION MANAGEMENT, RELIABILITY, HIGH AVAILABILITY, SCALING, AND LOAD BALANCING.**"*
> *"**This separation of concerns is the main benefit of using the Connect API versus the classic consumer/producer APIs.**"*
**Note the reuse:** worker failover uses **Kafka's consumer group protocol** (Ch. 4). Connect didn't invent a membership protocol; a Connect cluster *is* a consumer group.
#### 8.3 Converters and the Connect data model [#83-converters-and-the-connect-data-model]
> *"**This allows the Connect API to support different types of data stored in Kafka, INDEPENDENT of the connector implementation** (i.e., **ANY connector can be used with ANY record type, as long as a converter is available**)."*
Note: *"The JSON converter can be configured to **either include a schema in the result record or not** — so we can support **both structured and semistructured data.**"*
#### 8.4 Offset management — the killer feature [#84-offset-management--the-killer-feature]
> *"connectors need to know which data they have already processed, and they can use APIs provided by Kafka to maintain information on which events were already processed."*
##### Source connectors: *logical* partitions and offsets [#source-connectors-logical-partitions-and-offsets]
> *"the records the connector returns include **a LOGICAL PARTITION and a LOGICAL OFFSET. THOSE ARE NOT KAFKA PARTITIONS AND KAFKA OFFSETS** but rather partitions and offsets **as needed in the SOURCE SYSTEM.**"*
> 💡 *"**One of the most important design decisions involved in writing a source connector is deciding on a good way to PARTITION the data in the source system and to TRACK OFFSETS — this will impact THE LEVEL OF PARALLELISM the connector can achieve AND WHETHER IT CAN DELIVER AT-LEAST-ONCE OR EXACTLY-ONCE SEMANTICS.**"*
**The write ordering that makes it correct:**
Offsets are stored **after** the data is acked — the same "commit after processing" discipline as Ch. 7 §5.2, applied by the framework so connector authors can't get it wrong.
##### The three internal topics [#the-three-internal-topics]
*(Recall Ch. 5 §4.2: these are exactly the config topics that must be **1 partition** (strict ordering), **3 replicas** (availability), and **compacted** (indefinite retention).)*
##### Sink connectors: the inverse [#sink-connectors-the-inverse]
> *"they read Kafka records, which **already have a topic, partition, and offset**. Then they call the connector **`put()`** method that should store those records in the destination system. **If the connector reports success, they commit the offsets they've given to the connector back to Kafka, using the usual consumer commit methods.**"*
> *"**Offset tracking provided by the framework itself should make it easier for developers to write connectors and guarantee some level of CONSISTENT BEHAVIOR when using different connectors.**"*
***
# 9.11 Deploy / monitor / scale / recover (/docs/kafka/building-data-pipelines-kafka/deploy-monitor-scale-recover)
#### Deployment checklist [#deployment-checklist]
#### Scaling levers [#scaling-levers]
#### Monitoring [#monitoring]
| Signal | Why |
| ------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------- |
| **Connector / task status** (`status.storage.topic`, REST API) | Tasks fail individually and stay failed until restarted |
| **Connect worker logs** | *"Missing configurations or libraries are common causes for errors"* — the first place to look |
| **Source connector lag** (source system position vs stored offset) | Whether the source is keeping up |
| **Sink connector consumer lag** | Sinks are consumers — Ch. 7's lag monitoring applies directly |
| **Dead letter queue depth** | Silent `error.tolerance` drops are invisible otherwise |
| **Worker rebalance frequency** | Uses the consumer group protocol → same flapping concerns as Ch. 4 |
| **Target system write errors** | The sink's actual job |
#### Recovery / "backup" in pipeline terms [#recovery--backup-in-pipeline-terms]
#### The chapter's closing charge [#the-chapters-closing-charge]
> *"**Whatever data integration solution you eventually land on, THE MOST IMPORTANT FEATURE WILL ALWAYS BE ITS ABILITY TO DELIVER ALL MESSAGES UNDER ALL FAILURE CONDITIONS.** We believe that Kafka Connect is extremely reliable — based on its integration with Kafka's tried-and-true reliability features — **but it is important that you TEST the system of your choice, just like we do. Make sure your data integration system of choice can survive STOPPED PROCESSES, CRASHED MACHINES, NETWORK DELAYS, and HIGH LOADS without missing a message. After all, at their heart, data integration systems only have ONE JOB — delivering those messages.**"*
>
> *"...**It isn't enough that Kafka SUPPORTS at-least-once semantics; you must be sure you aren't ACCIDENTALLY CONFIGURING IT IN A WAY that may end up with less than complete reliability.**"*
*(That's the same charge as Ch. 7 §6 — validate, don't assume — now applied to the integration layer.)*
***
# 9.2 The eight considerations when building data pipelines (/docs/kafka/building-data-pipelines-kafka/eight-considerations-building-data)
> *"Those challenges are **not specific to Kafka** but are general data integration problems."*
#### 2.1 Timeliness [#21-timeliness]
> *"Some systems expect their data to arrive in **large bulks once a day**; others expect the data to arrive **a few milliseconds** after it is generated. Most data pipelines fit somewhere in between."*
**What a good system must do:** *"support **different timeliness requirements for different pipelines** and also **make the MIGRATION between different timetables easier as business requirements change.**"*
**And back pressure comes free:**
> *"This also makes it **trivial to apply back pressure — Kafka itself applies back pressure on producers (by delaying acks when needed) since consumption rate is driven ENTIRELY by the consumers.**"*
*(Note the mechanism: Kafka's pull model means a slow consumer can never stall a producer directly — Ch. 1's ActiveMQ lesson — but the broker can still throttle producers via delayed acks and quotas when it needs to.)*
#### 2.2 Reliability [#22-reliability]
> *"We want to **avoid single points of failure** and allow for **fast and automatic recovery** from all sorts of failure events. Data pipelines are often the way data arrives to **business-critical systems; failure for more than a few seconds can be hugely disruptive**, especially when the timeliness requirement is closer to the few-milliseconds end."*
**Delivery guarantees:**
**Where Kafka stands (per Ch. 7/8):**
> *"Kafka can provide **at-least-once on its own**, and **exactly-once when combined with an external data store that has a TRANSACTIONAL MODEL or UNIQUE KEYS.** Since **many of the end points ARE data stores that provide the right semantics** for exactly-once delivery, **a Kafka-based pipeline can often be implemented as exactly-once.**"*
> *"It is worth highlighting that **Kafka's Connect API makes it easier for connectors to build an end-to-end exactly-once pipeline by providing an API for integrating with the external systems WHEN HANDLING OFFSETS.** Indeed, **many of the available open source connectors support exactly-once delivery.**"*
This is exactly the workaround Ch. 8 §3.6 described (*"manage offsets in the database"*) — Connect gives connectors a first-class hook for it.
#### 2.3 High and varying throughput [#23-high-and-varying-throughput]
> *"they should be able to scale to very high throughputs... **Even more importantly, they should be able to ADAPT IF THROUGHPUT SUDDENLY INCREASES.**"*
**Kafka's own numbers:** *"capable of processing **hundreds of megabytes per second on even modest clusters**."*
**Connect's parallelism story:** *"the Kafka Connect API **focuses on parallelizing the work** and can do this **on a single node as well as by scaling out**, depending on system requirements... allows data sources and sinks to **split the work among multiple threads of execution and use the available CPU resources EVEN WHEN RUNNING ON A SINGLE MACHINE.**"*
Plus: *"Kafka also supports several types of **compression**, allowing users and admins to control the use of network and storage resources."*
#### 2.4 Data formats [#24-data-formats]
> *"One of the most important considerations... is **reconciling different data formats and data types.**"*
**The realistic mess the book describes:**
**Kafka's stance — deliberate agnosticism:**
> *"**Kafka itself and the Connect API are COMPLETELY AGNOSTIC when it comes to data formats.** ... Kafka Connect has its own **in-memory objects that include data types and schemas**, but... it allows for **pluggable CONVERTERS** to allow storing these records in any format. **This means that NO MATTER WHICH DATA FORMAT YOU USE FOR KAFKA, IT DOES NOT RESTRICT YOUR CHOICE OF CONNECTORS.**"*
**Schema propagation — the aspirational behavior:**
> *"Many sources and sinks have a schema; we can **read the schema from the source with the data, store it, and use it to validate compatibility or even UPDATE THE SCHEMA IN THE SINK DATABASE.** A classic example is a data pipeline from MySQL to Snowflake. **If someone added a column in MySQL, a great pipeline will make sure the column gets added to Snowflake too** as we are loading new data into it."*
**Sink-side format is the connector's job:** *"sink connectors are responsible for the format in which the data is written to the external system. **Some connectors choose to make this format pluggable** — for example, the S3 connector allows a choice between **Avro and Parquet**."*
**Behavioral differences, not just format differences:**
#### 2.5 Transformations — ETL vs ELT [#25-transformations--etl-vs-elt]
> *"Transformations are **more controversial** than other requirements."*
**The worked cost of ETL's drawback:**
> *"If the person who built the pipeline between MongoDB and MySQL decided to **filter certain events or remove fields**, all the users and applications who access the data in MySQL will only have access to **partial data.** If they require access to the missing fields, **the pipeline needs to be REBUILT, and historical data will require REPROCESSING (assuming it is available).**"*
That parenthetical — *"assuming it is available"* — is the real danger. A dropped field is often unrecoverable.
**Kafka's split:**
> ### ⚠️ WARNING — the one-to-many rule [#️-warning--the-one-to-many-rule]
>
> \*"When building an ETL system with Kafka, keep in mind that **Kafka allows you to build ONE-TO-MANY pipelines**, where the source data is written to Kafka **once** and then consumed by **multiple** applications and written to **multiple** target systems.
>
> **Some preprocessing and cleanup IS expected**, such as:
>
> * standardizing timestamps and data types
> * adding lineage
> * perhaps removing personal information
>
> *— transformations that will **benefit ALL consumers** of the data.*
>
> **But DON'T PREMATURELY CLEAN AND OPTIMIZE THE DATA ON INGEST BECAUSE IT MIGHT BE NEEDED LESS REFINED ELSEWHERE.**"\*
#### 2.6 Security [#26-security]
**The five questions:**
**What Kafka provides:** encryption on the wire (source→Kafka and Kafka→sink), **authentication via SASL**, authorization (*"so you can be sure that if a topic contains sensitive information, it can't be piped into less secured systems by someone unauthorized"*), and **an audit log to track access — unauthorized and authorized.** *"With some extra coding, it is also possible to track where the events in each topic came from and who modified them, so you can provide the entire lineage for each record."* (Ch. 11.)
##### ⚠️ Connector credentials — do not put them in config files [#️-connector-credentials--do-not-put-them-in-config-files]
> \*"Kafka Connect and its connectors need to be able to connect to, and authenticate with, external data systems, and **configuration of connectors will include CREDENTIALS.**
>
> **These days it is NOT RECOMMENDED to store credentials in configuration files**, since this means the configuration files have to be handled with extra care and have restricted access. **A common solution is to use an external secret management system such as HashiCorp Vault.**"\*
#### 2.7 Failure handling [#27-failure-handling]
> *"**Assuming that all data will be perfect all the time is DANGEROUS.** It is important to plan for failure handling **in advance.**"*
**The four questions:**
**Kafka's answer:** *"Because Kafka can be configured to **store all events for long periods of time**, it is possible to **go back in time and recover from errors** when needed. This also allows **replaying the events stored in Kafka to the target system if they were lost.**"*
That last question is the one that determines your retention setting: **your retention must exceed your bug-discovery latency.**
#### 2.8 Coupling and agility — three ways coupling sneaks in [#28-coupling-and-agility--three-ways-coupling-sneaks-in]
*(② is Ch. 1's schema-registry argument again — the lockstep-deploy problem. It shows up in *every* chapter because it's the same problem at every layer.)*
***
# 9.5 Example: file source → file sink (/docs/kafka/building-data-pipelines-kafka/example-file-source-file)
```bash
bin/connect-distributed.sh config/connect-distributed.properties &
# "In a real production environment, you'll want at least TWO OR THREE of
# these running to provide high availability."
```
**Create a source connector (pipe Kafka's own config file into a topic):**
```bash
echo '{"name":"load-kafka-config", "config":{"connector.class":
"FileStreamSource","file":"config/server.properties","topic":
"kafka-config-topic"}}' | curl -X POST -d @- http://localhost:8083/connectors \
-H "Content-Type: application/json"
```
Response includes `"tasks": [{"connector":"load-kafka-config","task":0}]` and `"type":"source"`.
**Verify:**
```bash
bin/kafka-console-consumer.sh --bootstrap-server=localhost:9092 \
--topic kafka-config-topic --from-beginning
{"schema":{"type":"string","optional":false},"payload":"# Licensed to the Apache..."}
{"schema":{"type":"string","optional":false},"payload":"broker.id=0"}
...
```
> *"Note that **by default, the JSON converter places a SCHEMA IN EACH RECORD.** In this specific case, the schema is very simple — there is only a single column, named `payload` of type string, containing a single line from the file for each record."*
**Create the sink:**
```bash
echo '{"name":"dump-kafka-config", "config":
{"connector.class":"FileStreamSink","file":"copy-of-server-properties",
"topics":"kafka-config-topic"}}' | curl -X POST -d @- \
http://localhost:8083/connectors --header "content-Type:application/json"
```
**The differences from the source config — note the plural:**
**Delete:**
```bash
curl -X DELETE http://localhost:8083/connectors/dump-kafka-config
```
> ### ⚠️ WARNING — FileStream connectors are demo-only [#️-warning--filestream-connectors-are-demo-only]
>
> *"This example uses FileStream connectors because they are simple and built into Kafka... **These should NOT be used for actual production pipelines, as they have MANY LIMITATIONS AND NO RELIABILITY GUARANTEES.** There are several alternatives if you want to ingest data from files: **FilePulse Connector, FileSystem Connector, or SpoolDir.**"*
***
# 9.6 Example: MySQL → Kafka → Elasticsearch (/docs/kafka/building-data-pipelines-kafka/example-mysql-kafka-elasticsearch)
#### 6.1 Getting connectors [#61-getting-connectors]
Three options: **Confluent Hub client**, **download from Confluent Hub** (or wherever the connector is hosted), or **build from source**:
```bash
git clone https://github.com/confluentinc/kafka-connect-elasticsearch
mvn install -DskipTests
# repeat for the JDBC connector
```
**Install into `plugin.path`:**
```bash
mkdir /opt/connectors/jdbc
mkdir /opt/connectors/elastic
cp .../kafka-connect-jdbc-10.3.x-SNAPSHOT.jar /opt/connectors/jdbc
cp .../kafka-connect-elasticsearch-11.1.0-SNAPSHOT.jar /opt/connectors/elastic
cp .../kafka-connect-elasticsearch-11.1.0-SNAPSHOT-package/share/java/\
kafka-connect-elasticsearch/* /opt/connectors/elastic
```
> ⚠️ *"since we need to connect not just to any database but specifically to MySQL, you'll need to **download and install a MySQL JDBC driver. THE DRIVER DOESN'T SHIP WITH THE CONNECTOR FOR LICENSE REASONS.**"* → place the jar in `/opt/connectors/jdbc`.
**Restart workers and confirm:**
```bash
curl http://localhost:8083/connector-plugins
# ElasticsearchSinkConnector (sink), JdbcSinkConnector (sink),
# JdbcSourceConnector (source)
```
#### 6.2 Source data [#62-source-data]
```sql
create database test;
use test;
create table login (username varchar(30), login_time datetime);
insert into login values ('gwenshap', now());
insert into login values ('tpalino', now());
```
#### 6.3 💡 Discovering configuration via the REST API [#63--discovering-configuration-via-the-rest-api]
**You don't have to read the docs — ask the API:**
```bash
curl -X PUT -d '{"connector.class":"JdbcSource"}' \
localhost:8083/connector-plugins/JdbcSourceConnector/config/validate/ \
--header "content-Type:application/json"
```
```json
{ "configs": [ { "definition": {
"default_value": "",
"display_name": "Timestamp Column Name",
"documentation": "The name of the timestamp column to use to detect
new or modified rows. This column may not be nullable.",
"group": "Mode",
"importance": "MEDIUM",
"name": "timestamp.column.name",
"order": 3, "required": false, "type": "STRING", "width": "MEDIUM" } },
... ] }
```
> *"We asked the REST API to **validate** configuration for a connector and sent it a configuration with **just the class name (this is the bare minimum configuration necessary).** As a response, we got **the JSON definition of ALL AVAILABLE CONFIGURATIONS.**"*
This is a genuinely useful trick: the validate endpoint doubles as **self-documenting config discovery**, including defaults, importance, and grouping.
#### 6.4 The JDBC source connector [#64-the-jdbc-source-connector]
```bash
echo '{"name":"mysql-login-connector", "config":{
"connector.class":"JdbcSourceConnector",
"connection.url":"jdbc:mysql://127.0.0.1:3306/test?user=root",
"mode":"timestamp",
"table.whitelist":"login",
"validate.non.null":false,
"timestamp.column.name":"login_time",
"topic.prefix":"mysql."}}' | curl -X POST -d @- \
http://localhost:8083/connectors --header "content-Type:application/json"
```
**Verify:**
```bash
bin/kafka-console-consumer.sh --bootstrap-server=localhost:9092 \
--topic mysql.login --from-beginning
```
**Troubleshooting — check the Connect worker logs:**
```txt
[2016-10-16 19:39:40,482] ERROR Error while starting connector
mysql-login-connector (org.apache.kafka.connect.runtime.WorkerConnector:108)
org.apache.kafka.connect.errors.ConnectException: java.sql.SQLException:
Access denied for user 'root;'@'localhost' (using password: NO)
```
> *"Other issues can involve **the existence of the driver in the classpath** or **permissions to read the table.**"*
Then: *"if you insert additional rows in the login table, you should **immediately** see them reflected in the `mysql.login` topic."*
> ### ⚠️ CHANGE DATA CAPTURE AND THE DEBEZIUM PROJECT [#️-change-data-capture-and-the-debezium-project]
>
> \*"The JDBC connector we are using **uses JDBC and SQL to SCAN database tables for new records.** It detects new records by using **timestamp fields or an incrementing primary key. THIS IS A RELATIVELY INEFFICIENT AND AT TIMES INACCURATE PROCESS.**
>
> **All relational databases have a TRANSACTION LOG (also called redo log, binlog, or write-ahead log)** as part of their implementation, and **many allow external systems to read data directly from their transaction log — a FAR MORE ACCURATE AND EFFICIENT process known as CHANGE DATA CAPTURE. Most modern ETL systems depend on change data capture as a data source.**
>
> \*"**The Debezium Project provides a collection of high-quality, open source, change capture connectors** for a variety of databases. **If you are planning on streaming data from a relational database to Kafka, we HIGHLY RECOMMEND using a Debezium change capture connector if one exists for your database.**
>
> In addition, **the Debezium documentation is one of the best we've seen** — it covers useful design patterns and use cases related to change data capture, **especially in the context of microservices.**"\*
#### 6.5 The Elasticsearch sink connector [#65-the-elasticsearch-sink-connector]
```bash
elasticsearch &
curl http://localhost:9200/ # verify it's up
```
```bash
echo '{"name":"elastic-login-connector", "config":{
"connector.class":"ElasticsearchSinkConnector",
"connection.url":"http://localhost:9200",
"type.name":"mysql-data",
"topics":"mysql.login",
"key.ignore":true}}' | curl -X POST -d @- \
http://localhost:8083/connectors --header "content-Type:application/json"
```
**The configs explained — and `key.ignore` is the interesting one:**
> * *"`connection.url` is simply the URL of the local Elasticsearch server."*
> * *"**Each topic in Kafka will become, by default, a separate Elasticsearch index, with the same name as the topic.**"*
> * *"**The JDBC connector DOES NOT POPULATE THE MESSAGE KEY. As a result, the events in Kafka have NULL KEYS.** Because the events in Kafka lack keys, we need to tell the Elasticsearch connector to **use the topic name, partition ID, and offset as the key for each event.** This is done by setting **`key.ignore` to true.**"*
**Verify the index and search it:**
```bash
$ curl 'localhost:9200/_cat/indices?v'
health status index ... docs.count ... store.size
yellow open mysql.login ... 2 ... 3.9kb
$ curl -s -X "GET" "http://localhost:9200/mysql.login/_search?pretty=true"
# hits: _id "mysql.login+0+0" → {"username":"gwenshap","login_time":1621699811000}
# _id "mysql.login+0+1" → {"username":"tpalino", "login_time":1621699816000}
```
> *"If the index isn't there, look for errors in the Connect worker log. **Missing configurations or libraries are common causes for errors.**"*
> ### BUILD YOUR OWN CONNECTORS [#build-your-own-connectors]
>
> *"**The Connector API is public and anyone can create a new connector.** So if the datastore you wish to integrate with does not have an existing connector, we encourage you to write your own. You can then **contribute it to Confluent Hub** so others can discover and use it."* Resources: multiple blog posts, talks from **Kafka Summit NY 2019**, **Kafka Summit London 2018**, **ApacheCon**, existing connectors as a starting point, and **an Apache Maven archetype** to jump-start. Ask on `users@kafka.apache.org`.
***
# 9. Building Data Pipelines (Kafka Connect) (/docs/kafka/building-data-pipelines-kafka)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 9*
> **Why Connect exists (added in Kafka 0.9):** *"We noticed that there were **specific challenges in integrating Kafka into data pipelines that EVERY ORGANIZATION had to solve**, and decided to add APIs to Kafka that solve some of those challenges **rather than force every organization to figure them out from scratch.**"*
***
# 9.4 Kafka Connect architecture (/docs/kafka/building-data-pipelines-kafka/kafka-connect-architecture)
**The vocabulary, precisely:**
| Term | Definition |
| --------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Connector plug-in** | *"libraries that Kafka Connect executes and that are responsible for moving the data"* |
| **Worker** | A process in the Connect cluster. You install plug-ins on workers |
| **Connector** | A *configured instance* — created/managed via REST API |
| **Task** | *"Connectors start additional tasks to move large amounts of data IN PARALLEL and use the available resources on the worker nodes more efficiently"* |
| **Converter** | *"support storing those data objects in Kafka in different formats"* |
**Converter availability:** *"**JSON format support is part of Apache Kafka**, and the **Confluent Schema Registry** provides **Avro, Protobuf, and JSON Schema** converters."* → *"This allows users to choose the format in which data is stored in Kafka **independent of the connectors they use.**"*
#### 4.1 Running Connect [#41-running-connect]
> *"Kafka Connect ships with Apache Kafka, so **there is no need to install it separately. For production use, especially if you are planning to use Connect to move large amounts of data or run many connectors, YOU SHOULD RUN CONNECT ON SEPARATE SERVERS FROM YOUR KAFKA BROKERS.**"*
*(Consistent with Ch. 2: never colocate a significant application with a broker — it competes for page cache.)*
```bash
bin/connect-distributed.sh config/connect-distributed.properties
```
#### Key worker configurations [#key-worker-configurations]
| Config | Notes |
| --------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`bootstrap.servers`** | *"You don't need to specify every broker... but it's **recommended to specify at least three.**"* |
| **`group.id`** | *"**All workers with the same group ID are part of the same Connect cluster.** A connector started on the cluster will run on **any** worker, and so will its tasks."* |
| **`plugin.path`** | Where connectors, converters, transformations, and secret providers live — see below |
| **`key.converter` / `value.converter`** | Default: **JSON** via `JSONConverter` (in Apache Kafka). Or `AvroConverter`, `ProtobufConverter`, `JsonSchemaConverter` (Confluent Schema Registry) |
| **`rest.host.name` / `rest.port`** | *"Connectors are typically configured and monitored through the REST API"* |
##### ⚠️ `plugin.path` layout — the classpath trap [#️-pluginpath-layout--the-classpath-trap]
##### Converter-specific configuration — the prefix rule [#converter-specific-configuration--the-prefix-rule]
#### Verifying the cluster [#verifying-the-cluster]
```bash
$ curl http://localhost:8083/
{"version":"3.0.0-SNAPSHOT","commit":"fae0784ce32a448a",
"kafka_cluster_id":"pfkYIGZQSXm8RylvACQHdg"}
$ curl http://localhost:8083/connector-plugins
[
{"class":"org.apache.kafka.connect.file.FileStreamSinkConnector", "type":"sink", "version":"3.0.0-SNAPSHOT"},
{"class":"org.apache.kafka.connect.file.FileStreamSourceConnector", "type":"source","version":"3.0.0-SNAPSHOT"},
{"class":"org.apache.kafka.connect.mirror.MirrorCheckpointConnector","type":"source","version":"1"},
{"class":"org.apache.kafka.connect.mirror.MirrorHeartbeatConnector", "type":"source","version":"1"},
{"class":"org.apache.kafka.connect.mirror.MirrorSourceConnector", "type":"source","version":"1"}
]
```
**Note what's in plain Apache Kafka:** file source, file sink, and **the three MirrorMaker 2.0 connectors** — MM2 is *built on Connect* (Ch. 10).
> ### STANDALONE MODE [#standalone-mode]
>
> \*"similar to distributed mode — you just run `bin/connect-standalone.sh`... You can also **pass in a connector configuration file on the command line instead of through the REST API.** In this mode, **all the connectors and tasks run on the one standalone worker.**
>
> **It is used in cases where connectors and tasks need to run on a SPECIFIC MACHINE** (e.g., **the syslog connector listens on a port, so you need to know which machines it is running on**)."\*
***
# 9.1 What problem does Kafka solve in a data pipeline? (/docs/kafka/building-data-pipelines-kafka/problem-does-kafka-solve)
#### Two shapes of pipeline [#two-shapes-of-pipeline]
#### The core value proposition, in one sentence [#the-core-value-proposition-in-one-sentence]
> *"The main value Kafka provides to data pipelines is its ability to serve as **a very large, reliable BUFFER between various stages in the pipeline.** This effectively **decouples producers and consumers** of data within the pipeline and **allows use of the SAME DATA from the source in MULTIPLE target applications and systems, all with DIFFERENT timeliness and availability requirements.**"*
> ### PUTTING DATA INTEGRATION IN CONTEXT [#putting-data-integration-in-context]
>
> \*"Some organizations think of Kafka as **an end point** of a pipeline. They look at questions such as 'How do I get data from Kafka to Elastic?' **This is a valid question to ask**... But we are going to start the discussion by looking at the use of Kafka **within a larger context that includes at least two (and possibly many more) end points that are not Kafka itself.**
>
> **We encourage anyone faced with a data-integration problem to consider the bigger picture and not focus only on the immediate end points. FOCUSING ON SHORT-TERM INTEGRATIONS IS HOW YOU END UP WITH A COMPLEX AND EXPENSIVE-TO-MAINTAIN DATA INTEGRATION MESS.**"\*
This is Ch. 1's N×M problem restated at the *integration tooling* layer. Each ad-hoc "get data from A to B" tool is a point-to-point connection in disguise.
***
# 9.12 Self-test (/docs/kafka/building-data-pipelines-kafka/self-test)
What are the two pipeline shapes involving Kafka, and why does the book insist you consider the bigger picture?
How does Kafka decouple timeliness requirements? Give the batch-consumer example.
Where does back pressure come from in a Kafka pipeline, and what mechanism does the broker use?
When can a Kafka-based pipeline be exactly-once end to end? What does Connect contribute?
Why does Kafka-as-buffer eliminate the need for a complex back-pressure mechanism?
Kafka and Connect are "completely agnostic" about data formats. What component makes that true?
Describe a "great pipeline"'s behavior when someone adds a column in MySQL.
Give one example each of push vs pull sources, and append-only vs updatable sinks.
Define ETL and ELT. Give the main drawback of each, with a concrete cost.
Which transformations *should* happen in the pipeline? State the test.
List the five security questions for a data pipeline. Which one is specific to crossing datacenters?
Why shouldn't connector credentials live in config files, and what does Connect provide instead?
Which failure-handling question determines your retention setting, and why?
Name the three ways coupling sneaks into a pipeline, with the concrete failure each produces.
State the decision rule: Connect vs producer/consumer clients.
No connector exists for your datastore. Why is Connect still recommended? Give the "day or two vs a few months" argument.
Draw a Connect cluster: workers, connectors, tasks, converters, internal topics.
Two workers have the same `group.id`. What does that mean? What happens when one crashes, and which protocol does that use?
Explain `plugin.path` layout. What's the one thing that "will not work," and why is the classpath alternative discouraged?
When would you use standalone mode instead of distributed?
What differs between the FileStreamSource and FileStreamSink configs? Why is one plural?
Why must you not use FileStream connectors in production?
How do you discover a connector's available configuration options without reading documentation?
Contrast JDBC polling with log-based CDC on four dimensions. What does the book recommend, and why?
Why is `key.ignore=true` needed for the Elasticsearch sink, and what does it make possible?
What are SMTs for, and where does the line fall between SMTs and Kafka Streams?
Give the `transforms.*` configuration naming convention.
Which SMT would you use to remove PII? To route by timestamp? To detect tombstones?
What is `error.tolerance`, and which connectors can use it?
State the three responsibilities of a *connector* (as opposed to a task). How does the JDBC source decide its task count?
What does the source task context provide? What does the sink task context provide, and why is one of those items critical for exactly-once?
State the one-sentence division of labor between connectors/tasks and workers.
What are logical partitions and offsets? Give the file and JDBC examples.
Why is the choice of source partitioning and offset tracking the most important design decision in a source connector?
In what order does the worker send data and store offsets, and why does the order matter?
Name the three internal Connect topics and what each holds. How should they be configured?
When would you choose Flume/Logstash over Connect? A GUI ETL tool? A stream processing framework — and what do you give up?
What single feature does the chapter say matters most in any integration system, and what does it tell you to do about it?
**Previous:** [Chapter 8 — Exactly-Once Semantics](08-exactly-once-semantics.md)
**Next:** [Chapter 10 — Cross-Cluster Data Mirroring](10-cross-cluster-data-mirroring.md)
# 9.7 Single Message Transformations (SMTs) (/docs/kafka/building-data-pipelines-kafka/single-message-transformations-smts)
**The division of labor, restated:**
#### The built-in SMTs in Apache Kafka [#the-built-in-smts-in-apache-kafka]
| SMT | What it does |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Cast** | Change data type of a field |
| **MaskField** | *"Replace the contents of a field with null. **Useful for removing sensitive or personally identifying data.**"* |
| **Filter** | *"Drop or include all messages matching a condition. Built-in conditions include matching on a **topic name**, a **particular header**, or **whether the message is a TOMBSTONE (that is, has a null value)**."* |
| **Flatten** | *"Transform a nested data structure to a flat one... by **concatenating all the names of all fields in the path** to a specific value."* |
| **HeaderFrom** | *"Move or copy fields from the message into the header."* |
| **InsertHeader** | *"Add a static string to the header of each message."* |
| **InsertField** | *"Add a new field... either using values from its **metadata such as offset**, or with a **static value**."* |
| **RegexRouter** | *"Change the destination topic using a **regular expression and a replacement string**."* |
| **ReplaceField** | *"Remove or rename a field."* |
| **TimestampConverter** | *"Modify the time format of a field — for example, from Unix Epoch to a String."* |
| **TimestampRouter** | *"Modify the topic based on the message timestamp. **Mostly useful in SINK connectors when we want to copy messages to specific TABLE PARTITIONS based on their timestamp** and the topic field is used to find an equivalent dataset in the destination system."* |
**Third-party collections:** GitHub — **Lenses.io**, **Aiven**, **Jeremy Custenborder** — or **Confluent Hub**. Learning: the **"Twelve Days of SMT"** blog series; tutorials for writing your own.
#### Worked example: add a lineage header [#worked-example-add-a-lineage-header]
**Motivation:** *"The header will indicate that the record was created by this MySQL connector, **which is useful in case auditors want to examine the LINEAGE of these records.**"*
```bash
echo '{
"name": "mysql-login-connector",
"config": {
"connector.class": "JdbcSourceConnector",
"connection.url": "jdbc:mysql://127.0.0.1:3306/test?user=root",
"mode": "timestamp",
"table.whitelist": "login",
"validate.non.null": "false",
"timestamp.column.name": "login_time",
"topic.prefix": "mysql.",
"name": "mysql-login-connector",
"transforms": "InsertHeader",
"transforms.InsertHeader.type":
"org.apache.kafka.connect.transforms.InsertHeader",
"transforms.InsertHeader.header": "MessageSource",
"transforms.InsertHeader.value.literal": "mysql-login-connector"
}}' | curl -X POST -d @- http://localhost:8083/connectors \
--header "content-Type:application/json"
```
**The naming convention to internalize:**
**Result** (needs **Apache Kafka 2.7+** to print headers in the console consumer):
```bash
bin/kafka-console-consumer.sh --bootstrap-server=localhost:9092 \
--topic mysql.login --from-beginning --property print.headers=true
NO_HEADERS {"schema":...,"payload":{"username":"tpalino",...}}
MessageSource:mysql-login-connector {"schema":...,"payload":{"username":"rajini",...}}
```
> *"the **old records show `NO_HEADERS`**, but the **new records show `MessageSource:mysql-login-connector`**."*
> ### ERROR HANDLING AND DEAD LETTER QUEUES [#error-handling-and-dead-letter-queues]
>
> *"**Transforms** is an example of a connector config that **isn't specific to one connector but can be used in the configuration of ANY connector.** Another very useful config that can be used in **any SINK connector** is **`error.tolerance`** — you can configure any connector to **silently drop corrupt messages**, or to **route them to a special topic called a 'DEAD LETTER QUEUE.'**"* → *"Kafka Connect Deep Dive — Error Handling and Dead Letter Queues"* blog post.
*(Compare Ch. 7 §5.2 Rule 5 Pattern B: the same DLQ pattern, but here it's a config flag instead of application code.)*
***
# 10.10 What actually breaks in production — Ch. 10 consolidated (/docs/kafka/cross-cluster-data-mirroring/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| 1 | **Brokers spread across DCs behave badly; timeouts everywhere** | Kafka was *"designed, developed, tested, and tuned, all within a single datacenter"* — defaults assume LAN | Don't do it, except as a deliberate **stretch cluster** with 3 DCs |
| 2 | **Data lost during a WAN partition** | MirrorMaker was **producing** remotely; consumed events couldn't be delivered | Run MirrorMaker at the **target** DC — remote consuming fails safe |
| 3 | **Enormous cross-DC bandwidth bill** | Multiple applications each consuming the same data across the WAN | One cluster per DC; **mirror once**, consume locally |
| 4 | **A user visits another city's branch and their data isn't there** | **Hub-and-spoke** gives no cross-regional data access | Only use it for data that *"can be completely separated between regional datacenters"*; otherwise active-active |
| 5 | **Events mirrored back and forth forever** | Active-active without cycle prevention | Per-DC topic namespaces (`SF.users` / `NYC.users`); MM2 does this by default via alias prefixing |
| 6 | **User writes to one DC, reads from another, and doesn't see their own write** | Asynchronous multi-master | *"Stick"* each user to a datacenter |
| 7 | **Two conflicting orders for the same user; downstream state diverges** | Concurrent writes in two DCs | *"You WILL have conflicts"* — define **consistent** resolution rules both DCs will compute identically |
| 8 | **N² mirroring processes to manage** | Active-active needs a flow per pair per direction | Tools that share processes per destination cluster; one shared config file |
| 9 | **Custom origin-DC headers break mirroring** | *"None of the existing mirroring tools will support your specific header format"* | Use the tool's own convention, or accept the extra work |
| 10 | **DR cluster too small to carry production during a real disaster** | Cost-driven undersizing | *"A risky decision because you can't be sure it will hold up"* |
| 11 | **Failover loses \~5,000 messages** | All mirroring is **asynchronous**: 1M msg/s × 5 ms = 5,000 messages | Expected. **Planned** failover avoids it (stop primary, let mirroring drain). Monitor DR lag continuously |
| 12 | **A line item arrives with no corresponding sale** | *"Mirroring solutions currently don't support transactions"* — related events across topics arrive independently | Applications must tolerate it (Ch. 8 §3.6) |
| 13 | **Consumers fail on the DR cluster: offset doesn't exist** | Mirrored `__consumer_offsets` but **offsets diverge** (retention skew, producer retries) | Time-based failover, or **offset translation** (MM2) |
| 14 | **DR consumer finds a committed offset with no matching record** | Offset commits and records mirror independently and race | Decide in advance: beginning or end? |
| 15 | **Failover happened but nobody can explain what was reprocessed** | Offset-based failover is inherently unexplainable | **Time-based failover** — *"'We failed back to 4:03 a.m.' sounds better"* |
| 16 | **Offset reset tool silently did nothing / conflicted** | Consumer group still running | *"The group should be STOPPED while running this type of tool and started immediately after"* |
| 17 | **After failing back, the two clusters are permanently inconsistent** | The old primary retained events the DR never got — reverse mirroring leaves **phantom history** | *"First SCRAPE the original cluster — delete all data and committed offsets"* |
| 18 | **Applications can't find the DR cluster** | Broker hostnames **hardcoded** in client configs | A DNS name (3 brokers is enough) repointed on failover; build failover logic into clients for low RTO |
| 19 | **Consumers on the DR cluster consume from the wrong position after DNS flip** | *"Most failover scenarios DO require BOUNCING consumer applications"* | Include the bounce in the runbook |
| 20 | **Two-DC "stretch cluster" goes fully down when one DC fails** | One DC always holds the **ZooKeeper majority** | **Three** datacenters (or 2.5 DC with a tiebreaker ZK node) |
| 21 | **Legacy MirrorMaker stalls for 5–10 minutes when a topic is added** | MM1's **consumer-group rebalances** | **MirrorMaker 2.0** — allocates partitions without the group protocol |
| 22 | **Test topics replicated across an expensive WAN link** | `topics = .*` | Use `prod.*` and a `test.*` exclusion list |
| 23 | **DR topics have weaker durability than production** | **`min.insync.replicas` is NOT migrated by default** | Configure it explicitly on the target; customize the exclusion list |
| 24 | **After failover, producers can't write to the DR cluster** | **`Topic:Write` ACLs are deliberately not migrated** (so only MM can write) | *"Appropriate access must be EXPLICITLY GRANTED at the time of failover"* — put it in the runbook |
| 25 | **Prefixed/wildcard ACLs missing on the DR cluster** | *"Only LITERAL topic ACLs that match topics being mirrored are migrated"* | Configure them on the target explicitly |
| 26 | **MirrorMaker overwrote a live consumer group's offsets** | It doesn't — there's an interlock | *"MirrorMaker does not overwrite offsets if consumers on the target cluster are actively using the target consumer group"* |
| 27 | **Mirroring throughput poor; only one task running** | `tasks.max` default is **1** | Minimum 2; benchmark 1→32 and set just below the tapering point |
| 28 | **MirrorMaker CPU saturated** | **Decompress + recompress** of compressed events | Expected; watch CPU while scaling tasks. Cluster Linking avoids it entirely |
| 29 | **A noisy topic starves your latency-critical mirror** | Shared MirrorMaker cluster | Separate MirrorMaker cluster for sensitive topics |
| 30 | **Cross-DC link never reaches available bandwidth** | Default TCP buffers/window; slow-start after idle | Tune client + broker socket buffers, `tcp_window_scaling=1`, `tcp_slow_start_after_idle=0` |
| 31 | **Consumer performance collapses after enabling SSL** | SSL **defeats zero-copy** on the broker's read path | Consider consume-locally/produce-remotely — **and then** set `acks=all`, retries, `errors.tolerance=none`. Re-measure on modern Java |
| 32 | **Cloud MirrorMaker can't connect to on-prem brokers** | Firewall blocks inbound connections to on-prem | Run MirrorMaker **on premises** (produce remotely) |
| 33 | **Reported lag jumps around by a minute's worth** | Method 1 reads **committed** offsets; MM commits every minute by default | Use **Burrow**, or Method 2, understanding both are approximations |
| 34 | **MirrorMaker silently dropped messages and no alert fired** | **Both** lag methods only track the *latest offset* | Message counts + checksums (Confluent Control Center), plus a **canary** |
| 35 | **Ordering broken in the target cluster during retries** | `max.in.flight > 1` with retries (Ch. 3 §6) | `max.in.flight=1` is *"currently the only way"* for MM to guarantee ordering — at a heavy WAN throughput cost |
| 36 | **Failover plan worked last year, fails now** | Untested plan; *"a plan that works today may stop working after an upgrade"* | Practice **at least quarterly**; Chaos-Monkey-style continuous testing |
***
# 10.6 Apache Kafka's MirrorMaker (/docs/kafka/cross-cluster-data-mirroring/apache-kafka-s-mirrormaker)
#### 6.1 The evolution — why MM1 was replaced [#61-the-evolution--why-mm1-was-replaced]
> ### MORE ABOUT MIRRORMAKER [#more-about-mirrormaker]
>
> *"MirrorMaker **sounds very simple**, but **because we were trying to be very efficient and get very close to exactly-once delivery, IT TURNED OUT TO BE TRICKY TO IMPLEMENT CORRECTLY. MirrorMaker has been REWRITTEN MULTIPLE TIMES.**"*
#### 6.2 Architecture [#62-architecture]
**A nice operational argument:** *"If we assume that most Kafka deployments include Kafka Connect for other reasons (**sending database change events into Kafka is a very popular use case**), then **by running MirrorMaker INSIDE Connect, we can CUT DOWN ON THE NUMBER OF CLUSTERS WE NEED TO MANAGE.**"*
#### 💡 The key MM2 design decisions [#-the-key-mm2-design-decisions]
**Vocabulary:** *"A **replication flow** defines the configuration of a **directional** flow from a source cluster to a target cluster. **Multiple replication flows can be defined to define COMPLEX TOPOLOGIES**, including hub-and-spoke, active-standby, and active-active."*
#### 6.3 Configuring MirrorMaker [#63-configuring-mirrormaker]
```bash
bin/connect-mirror-maker.sh etc/kafka/connect-mirror-maker.properties
```
##### Active-standby: New York → London [#active-standby-new-york--london]
```properties
clusters = NYC, LON # ① aliases
NYC.bootstrap.servers = kafka.nyc.example.com:9092 # ② prefix = alias
LON.bootstrap.servers = kafka.lon.example.com:9092
NYC->LON.enabled = true # ③ enable the flow
NYC->LON.topics = .* # ④ what to mirror
```
① *"Define **aliases** for the clusters used in replication flows."*
② *"Configure bootstrap for each cluster, **using the cluster alias as the prefix.**"*
③ *"Enable replication flow using the prefix **`source->target`**. **All configuration options for this flow use the same prefix.**"*
④ The topics to mirror.
##### Mirror topics — the naming strategy that prevents cycles [#mirror-topics--the-naming-strategy-that-prevents-cycles]
> *"for each replication flow, **a REGULAR EXPRESSION may be specified** for the topic names... In this example, we chose to replicate every topic, but **it is often good practice to use something like `prod.*` and AVOID REPLICATING TEST TOPICS.** A **separate topic EXCLUSION list** containing names or patterns like `test.*` may also be specified."*
**Auto-discovery:** *"MirrorMaker **periodically checks for new topics** in the source cluster and starts mirroring automatically if they match the configured patterns. **If more partitions are added to the source topic, the SAME NUMBER of partitions is automatically added to the target topic**, ensuring that events in the source topic **appear in the same partitions in the same order** in the target topic."*
##### Consumer offset migration [#consumer-offset-migration]
That interlock is the same principle as Ch. 5 §6.4 — never write offsets under a live group.
##### Topic configuration and ACL migration [#topic-configuration-and-acl-migration]
**Two failover landmines hidden in these defaults:**
1. `min.insync.replicas` is **not** migrated → your DR topics may have weaker durability than you assume.
2. `Topic:Write` ACLs are **deliberately not** migrated → **on failover, your producers cannot write until someone grants access.** Put this in the runbook.
##### `tasks.max` [#tasksmax]
> *"limits the maximum number of tasks... **The default is 1, but A MINIMUM OF 2 IS RECOMMENDED. When replicating a lot of topic partitions, HIGHER VALUES SHOULD BE USED if possible to increase parallelism.**"*
##### Configuration prefixes — the hierarchy [#configuration-prefixes--the-hierarchy]
#### 6.4 Multicluster topologies [#64-multicluster-topologies]
**Active-active NYC ↔ LON — just enable both directions:**
```properties
clusters = NYC, LON
NYC.bootstrap.servers = kafka.nyc.example.com:9092
LON.bootstrap.servers = kafka.lon.example.com:9092
NYC->LON.enabled = true
NYC->LON.topics = .*
LON->NYC.enabled = true
LON->NYC.topics = .*
```
> *"even though all topics from NYC are mirrored to LON and vice versa, **MirrorMaker ensures that the same event isn't constantly mirrored back and forth since REMOTE TOPICS USE THE CLUSTER ALIAS AS THE PREFIX.**"*
> 💡 **Operational best practice:** *"It is good practice to **use THE SAME CONFIGURATION FILE that contains the FULL replication topology for DIFFERENT MirrorMaker processes** since it **avoids conflicts when configs are shared using the internal configs topic in the target datacenter.**"* Start processes with the **`--clusters`** option to specify which target cluster this process serves.
**Fan out — add a third cluster:**
```properties
clusters = NYC, LON, SF
SF.bootstrap.servers = kafka.sf.example.com:9092
NYC->SF.enabled = true
NYC->SF.topics = .*
```
#### 6.5 Securing MirrorMaker [#65-securing-mirrormaker]
> *"For production clusters, **it is important to ensure that ALL cross-datacenter traffic is secure**... **SSL should be used to ENCRYPT ALL cross-datacenter traffic.**"*
```properties
NYC.security.protocol=SASL_SSL # ① match the broker listener
NYC.sasl.mechanism=PLAIN
NYC.sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule \
required username="MirrorMaker" password="MirrorMaker-password"; # ②
```
① *"Security protocol should **match that of the broker listener corresponding to the bootstrap servers** specified for the cluster. **SSL or SASL\_SSL is recommended.**"*
② *"For SSL, **keystores should be specified if mutual client authentication is enabled.**"*
##### The complete ACL list MirrorMaker needs [#the-complete-acl-list-mirrormaker-needs]
Note how the ACL list maps 1:1 to the feature list — each capability (data, configs, partitions, offsets, ACLs) needs its own grant pair. Miss one and that *one feature* silently stops working.
***
# 10.11 Decision guide (/docs/kafka/cross-cluster-data-mirroring/decision-guide)
#### The chapter's closing warning [#the-chapters-closing-warning]
> \*"remember that **multicluster configuration and mirroring pipelines SHOULD BE MONITORED AND TESTED JUST LIKE EVERYTHING ELSE you take into production. BECAUSE MULTICLUSTER MANAGEMENT IN KAFKA CAN BE EASIER THAN IT IS WITH RELATIONAL DATABASES, SOME ORGANIZATIONS TREAT IT AS AN AFTERTHOUGHT and neglect to apply proper design, planning, testing, deployment automation, monitoring, and maintenance.**
>
> By taking multicluster management seriously, **preferably as part of a HOLISTIC disaster or geodiversity plan for the ENTIRE ORGANIZATION that involves MULTIPLE APPLICATIONS AND DATA STORES**, you will greatly increase the chances of successfully managing multiple Kafka clusters."\*
***
# 10.7 Deploying MirrorMaker in production (/docs/kafka/cross-cluster-data-mirroring/deploying-mirrormaker-production)
#### 7.1 Deployment modes [#71-deployment-modes]
**All Connect deployment modes work:**
| Mode | Use |
| ---------------------------------------------------------------- | ----------------------------------------------------------------------------------- |
| **Standalone** | *"development and testing"* — one worker, one machine |
| **A connector in an existing distributed Connect cluster** | *"by explicitly configuring the connectors"* |
| **Distributed** — dedicated MM cluster or shared Connect cluster | ✅ *"For production use, we RECOMMEND running MirrorMaker in DISTRIBUTED MODE"* |
#### 7.2 ⚠️ Where to run MirrorMaker — the most important operational decision [#72-️-where-to-run-mirrormaker--the-most-important-operational-decision]
**The reasoning, in full — this is the chapter's best paragraph:**
> \*"long-distance networks can be a bit **less reliable** than those inside a datacenter. If there is a network partition and you lose connectivity between the datacenters, **having a CONSUMER that is unable to connect to a cluster is MUCH SAFER than a PRODUCER that can't connect.**
>
> **If the CONSUMER can't connect, it simply won't be able to read events, but THE EVENTS WILL STILL BE STORED IN THE SOURCE KAFKA CLUSTER and can remain there for a long time. THERE IS NO RISK OF LOSING EVENTS.**
>
> **On the other hand, if the events were ALREADY CONSUMED and MirrorMaker CAN'T PRODUCE them due to network partition, THERE IS ALWAYS A RISK THAT THESE EVENTS WILL ACCIDENTALLY GET LOST by MirrorMaker.**
>
> **So, REMOTE CONSUMING IS SAFER THAN REMOTE PRODUCING.**"\*
#### Two exceptions where you must produce remotely [#two-exceptions-where-you-must-produce-remotely]
##### Exception 1: encryption asymmetry — and the zero-copy reason [#exception-1-encryption-asymmetry--and-the-zero-copy-reason]
> \*"**When do you have to consume locally and produce remotely?** The answer is **when you need to encrypt the data while it is transferred between the datacenters but you don't need to encrypt the data inside the datacenter.**
>
> **Consumers take a SIGNIFICANT PERFORMANCE HIT when connecting to Kafka with SSL encryption — MUCH MORE SO THAN PRODUCERS. This is because use of SSL requires COPYING DATA FOR ENCRYPTION, which means CONSUMERS NO LONGER ENJOY THE PERFORMANCE BENEFITS OF THE USUAL ZERO-COPY OPTIMIZATION. And this performance hit ALSO AFFECTS THE KAFKA BROKERS THEMSELVES.**"\*
**⚠️ If you do this, you must compensate for the loss of the fail-safe:**
> 💡 *"Note that **newer versions of Java have SIGNIFICANTLY INCREASED SSL PERFORMANCE, so producing locally and consuming remotely may be a viable option EVEN WITH ENCRYPTION.**"* — i.e. re-measure before accepting the trade.
##### Exception 2: firewalls (hybrid cloud) [#exception-2-firewalls-hybrid-cloud]
> *"a hybrid scenario when mirroring from an **on-premises cluster to a cloud cluster. Secure on-premises clusters are likely to be BEHIND A FIREWALL THAT DOESN'T ALLOW INCOMING CONNECTIONS FROM THE CLOUD. Running MirrorMaker ON PREMISE allows ALL CONNECTIONS TO BE FROM ON PREMISES TO THE CLOUD.**"*
#### 7.3 Monitoring MirrorMaker [#73-monitoring-mirrormaker]
##### Kafka Connect metrics [#kafka-connect-metrics]
*"connector metrics to monitor **connector status**, source connector metrics to monitor **throughput**, and worker metrics to monitor **rebalance delays.** Connect also provides a **REST API** to view and manage connectors."*
##### MirrorMaker-specific metrics [#mirrormaker-specific-metrics]
| Metric | Meaning |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`replication-latency-ms`** | *"the time interval between the **record timestamp** and the time at which the record was **successfully produced to the target cluster.** ... useful to detect if the target is not keeping up."* |
| **`record-age-ms`** | *"the age of records at the time of replication"* |
| **`byte-rate`** | replication throughput |
| **`checkpoint-latency-ms`** | *"offset migration latency"* |
| **heartbeats** | *"MirrorMaker also emits **periodic heartbeats by default**, which can be used to monitor its health."* |
**How to interpret latency:** *"**Increased latency during peak hours may be OK if there is sufficient capacity to catch up later, but SUSTAINED INCREASE in latency may indicate INSUFFICIENT CAPACITY.**"*
##### ⚠️ Lag monitoring — two methods, neither perfect [#️-lag-monitoring--two-methods-neither-perfect]
##### Producer and consumer metrics for tuning [#producer-and-consumer-metrics-for-tuning]
##### Canary [#canary]
> *"If you monitor everything else, a canary isn't strictly necessary, **but we like to add it in for MULTIPLE LAYERS OF MONITORING.** It provides a process that, **every minute, sends an event to a special topic in the source cluster and tries to read the event from the destination cluster. It also alerts you if the event takes more than an acceptable amount of time to arrive.** This can mean that MirrorMaker is lagging or that **it isn't available at all.**"*
A canary is the only listed monitor that measures the thing you actually care about — **end-to-end delivery** — rather than a proxy for it.
***
# 10.4 Disaster recovery planning (/docs/kafka/cross-cluster-data-mirroring/disaster-recovery-planning)
#### 4.1 RTO and RPO [#41-rto-and-rpo]
**The implication chain:** `RPO=0` ⇒ synchronous replication ⇒ **you cannot use MirrorMaker** ⇒ you need a stretch cluster (§6) or a commercial synchronous solution (§9).
#### 4.2 ⚠️ Data loss and inconsistencies in unplanned failover [#42-️-data-loss-and-inconsistencies-in-unplanned-failover]
**The arithmetic — do this for your own throughput:**
**Planned failover is different:** *"you can **stop the primary cluster and WAIT for the mirroring process to mirror the remaining messages** before failing over applications to the DR cluster, thus **avoiding this data loss.**"*
**And the cross-topic consistency problem:**
> *"note that **mirroring solutions currently DON'T SUPPORT TRANSACTIONS**, which means that **if some events in multiple topics are related to each other (e.g., SALES and LINE ITEMS), you can have SOME events arrive to the DR site in time for the failover and OTHERS THAT DON'T. Your applications will need to be able to handle A LINE ITEM WITHOUT A CORRESPONDING SALE after you failover.**"*
*(This is exactly Ch. 8 §3.6's finding: MirrorMaker can be per-record exactly-once but cannot preserve transaction atomicity.)*
#### 4.3 The four failover-offset strategies [#43-the-four-failover-offset-strategies]
> *"**One of the challenging tasks in failing over to another cluster is making sure applications know where to start consuming data.**"*
##### Strategy A: Auto offset reset — simplest, lossiest [#strategy-a-auto-offset-reset--simplest-lossiest]
##### Strategy B: Replicate the `__consumer_offsets` topic — three serious caveats [#strategy-b-replicate-the-__consumer_offsets-topic--three-serious-caveats]
> *"If you mirror this topic to your DR cluster, when consumers start consuming from the DR cluster, **they will be able to pick up their old offsets and continue from where they left off. It is simple, but THERE IS A LONG LIST OF CAVEATS.**"*
**What you must decide in advance:**
> *"this approach has its limitations. **Still, this option lets you failover with a REDUCED NUMBER of duplicated or missing events compared to other approaches while still being simple to implement.**"*
##### Strategy C: Time-based failover — 💡 the good compromise [#strategy-c-time-based-failover---the-good-compromise]
**And the *human* argument, which is unusually candid and quite persuasive:**
> *"**the behavior is MUCH EASIER TO EXPLAIN TO EVERYONE IN THE COMPANY — 'We failed back to 4:03 a.m.' sounds better than 'We failed back to what may or may not be the latest committed offsets.'**"*
**Two ways to implement it:**
```bash
bin/kafka-consumer-groups.sh --bootstrap-server localhost:9092 \
--reset-offsets --all-topics --group my-group \
--to-datetime 2021-03-31T04:03:00.000 --execute
```
> *"This option is **recommended in deployments that need to guarantee a level of certainty in their failover.**"*
*(Note the "stop the group first" requirement — same constraint as Ch. 5 §6.4's `alterConsumerGroupOffsets`.)*
##### Strategy D: Offset translation — most precise [#strategy-d-offset-translation--most-precise]
**Summary of the four:**
| Strategy | Precision | Complexity | Residual problem |
| ------------------------------------- | ----------- | ----------------------- | -------------------------------------- |
| **A. auto.offset.reset** | Worst | Trivial | Mass duplicates *or* unknown data loss |
| **B. Replicate `__consumer_offsets`** | Poor–medium | Low | Offsets diverge; commits ≠ records |
| **C. Time-based** | Good | Low–medium | Bounded duplicates; *explainable* |
| **D. Offset translation** | Best | Medium (built into MM2) | Commits still race the records |
#### 4.4 ⚠️ After the failover — you probably have to scrape the old primary [#44-️-after-the-failover--you-probably-have-to-scrape-the-old-primary]
> *"It is tempting to simply **modify the mirroring processes to reverse their direction** and start mirroring from the new primary to the old one. **However, this leads to two important questions:**"*
> **The remedy:** *"for scenarios where **consistency and ordering guarantees are critical**, the simplest solution is to **FIRST SCRAPE THE ORIGINAL CLUSTER — DELETE ALL THE DATA AND COMMITTED OFFSETS — and then start mirroring from the new primary back to what is now the new DR cluster. This gives you A CLEAN SLATE that is identical to the new primary.**"*
That is a genuinely expensive operation to discover mid-incident. **Plan and rehearse it.**
#### 4.5 Cluster discovery [#45-cluster-discovery]
> *"in the event of failover, **your applications will need to know how to start communicating with the failover cluster. If you HARDCODED THE HOSTNAMES of your primary cluster brokers in the producer and consumer properties, THIS WILL BE CHALLENGING.**"*
*(Note the `client.dns.lookup` interaction from Ch. 5 §3.1 — a DNS alias plus SASL needs `resolve_canonical_bootstrap_servers_only`.)*
***
# 10.1 Five use cases for cross-cluster mirroring (/docs/kafka/cross-cluster-data-mirroring/five-use-cases-cross)
#### ① Regional and central clusters [#-regional-and-central-clusters]
> *"the classic example is a company that **modifies prices based on supply and demand.** This company can have a datacenter in each city in which it has a presence, **collects information about local supply and demand, and adjusts prices accordingly.** All this information will then be **mirrored to a central cluster where business analysts can run company-wide reports on its revenue.**"*
#### ② High availability (HA) and disaster recovery (DR) [#-high-availability-ha-and-disaster-recovery-dr]
> *"The applications run on just one Kafka cluster and don't need data from other locations, **but you are concerned about the possibility of the ENTIRE CLUSTER becoming unavailable.** For redundancy, you'd like a second Kafka cluster with all the data... so in case of emergency you can direct your applications to the second cluster."*
#### ③ Regulatory compliance [#-regulatory-compliance]
> *"Companies operating in different countries may need **different configurations and policies to conform to legal and regulatory requirements in each country.** For instance, **some datasets may be stored in separate clusters with STRICT ACCESS CONTROL, with SUBSETS of data replicated to other clusters with WIDER ACCESS.** To comply with regulatory policies that govern **retention period** in each region, datasets may be stored in clusters in different regions with different configurations."*
#### ④ Cloud migrations [#-cloud-migrations]
> *"if a new application is deployed **in the cloud** but requires some data that is updated by applications running **on premises** and stored in an on-premises database, you can **use Kafka Connect to capture database changes to the LOCAL Kafka cluster and then MIRROR these changes to the CLOUD Kafka cluster** where the new application can use them. **This helps CONTROL THE COSTS of cross-datacenter traffic as well as improve GOVERNANCE AND SECURITY of the traffic.**"*
**The pattern:** *Connect for the source hop, MirrorMaker for the WAN hop.* Never let N cloud applications each pull across the WAN.
#### ⑤ Aggregation of data from edge clusters [#-aggregation-of-data-from-edge-clusters]
> \*"Several industries, including **retail, telecommunications, transportation, and healthcare**, generate data from **small devices with limited connectivity.** An aggregate cluster with high availability can support analytics and other use cases for data from a large number of edge clusters.
>
> **This REDUCES connectivity, availability, and durability requirements on low-footprint edge clusters**, for example, in **IoT** use cases. A highly available aggregate cluster **provides business continuity EVEN WHEN EDGE CLUSTERS ARE OFFLINE** and **simplifies the development of applications that don't have to directly deal with a large number of edge clusters with unstable networks.**"\*
***
# 10. Cross-Cluster Data Mirroring (/docs/kafka/cross-cluster-data-mirroring)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 10*
> **Terminology, established up front:** *"In most databases, continuously copying data between database servers is called **replication**. Since we've used replication to describe movement of data between Kafka nodes that are part of the **same** cluster, we'll call copying of data **between Kafka clusters MIRRORING.** Apache Kafka's built-in cross-cluster replicator is called **MirrorMaker.**"*
**The easy case, dismissed immediately:** *"In some cases, the clusters are **completely separated**... different departments, different use cases, different SLAs/workloads, different security requirements. **Those use cases are fairly easy — managing multiple distinct clusters is the same as running a single cluster multiple times.**"* This chapter is about the **interdependent** case.
***
# 10.3 Multicluster architectures (/docs/kafka/cross-cluster-data-mirroring/multicluster-architectures)
#### 3.1 Hub-and-spoke [#31-hub-and-spoke]
**When to use it:** *"data is produced in multiple datacenters and **some consumers need access to the entire dataset.** The architecture also allows for applications in each datacenter to **only process data local to that specific datacenter. But it does NOT give access to the entire dataset from every datacenter.**"*
**Benefits:**
**The drawback, illustrated with the bank example:**
> \*"**Processors in one regional datacenter CAN'T ACCESS DATA IN ANOTHER.**
>
> Suppose we are a large bank with branches in multiple cities. We store user profiles and account history in a Kafka cluster in each city, replicated to a central cluster for business analytics. When users connect to the website or visit their local branch, they are routed to their local cluster.
>
> **However, suppose that a user visits a branch in a DIFFERENT CITY. Because the user information doesn't exist in the city they are visiting, the branch will be forced to interact with a remote cluster (NOT RECOMMENDED) or HAVE NO WAY TO ACCESS THE USER'S INFORMATION (REALLY EMBARRASSING).**"\*
> *"For this reason, use of this pattern is usually **limited to only parts of the dataset that can be COMPLETELY SEPARATED between regional datacenters.**"*
**Implementation note:** *"for each regional datacenter you need **at least one mirroring process ON THE CENTRAL DATACENTER**"* (consistent with principle 3 — consume remotely). *"If the same topic exists in multiple datacenters, you can write all the events **to one topic with the same name** in the central cluster, **or** write events from each datacenter to a **separate topic.**"*
#### 3.2 Active-active [#32-active-active]
**Benefits:**
**The drawback: conflicts from asynchronous multi-master writes.**
##### Conflict type 1 — read-your-own-write violation [#conflict-type-1--read-your-own-write-violation]
> *"If a user sends an event to one datacenter and reads events from another, **it is possible that the event they wrote HASN'T ARRIVED at the second datacenter yet.** To the user, it will look like they **just added a book to their wish list and clicked on the wish list, but the book isn't there.**"*
**Mitigation:** *"developers usually find a way to **'STICK' each user to a specific datacenter** and make sure they use the same cluster most of the time (unless they connect from a remote location or the datacenter becomes unavailable)."*
##### Conflict type 2 — genuinely conflicting writes [#conflict-type-2--genuinely-conflicting-writes]
> *"An event from one datacenter says the user ordered **book A**, and an event from more or less the same time at a second datacenter says the same user ordered **book B.** After mirroring, **both datacenters have both events** and thus **each datacenter has two CONFLICTING events.**"*
**The design questions you must answer:**
> *"**It is important to keep in mind that if you use this architecture, YOU WILL HAVE CONFLICTS AND WILL NEED TO DEAL WITH THEM.**"*
##### ⚠️ Avoiding infinite mirroring loops — the namespace trick [#️-avoiding-infinite-mirroring-loops--the-namespace-trick]
**The alternative: record headers.**
> \*"**Record headers introduced in Apache Kafka in version 0.11.0 enable events to be TAGGED WITH THEIR ORIGINATING DATACENTER.** Header information may also be used to **avoid endless mirroring loops** and to allow **processing events from different datacenters separately.** You can also implement this feature by using a **structured data format for the record values (Avro is our favorite example)**...
>
> ⚠️ *"**However, this does require EXTRA EFFORT when mirroring, since NONE OF THE EXISTING MIRRORING TOOLS WILL SUPPORT YOUR SPECIFIC HEADER FORMAT.**"*
**The N² problem:**
> *"part of the challenge of active-active mirroring, **especially with more than two datacenters, is that you will need mirroring tasks for EACH PAIR of datacenters AND EACH DIRECTION.** Many mirroring tools these days can **share processes**, for example, using the same process for all mirroring to a destination cluster."*
**The verdict:**
> *"If you find ways to handle the challenges... then this architecture is **HIGHLY RECOMMENDED. It is the MOST SCALABLE, RESILIENT, FLEXIBLE, AND COST-EFFECTIVE option we are aware of.** So, it is well worth the effort to figure out solutions for **avoiding replication cycles, keeping users mostly in the same datacenter, and handling conflicts.**"*
#### 3.3 Active-standby [#33-active-standby]
> *"This is often **a LEGAL REQUIREMENT rather than something that the business is actually planning on doing — but you still need to be ready.**"*
**Benefits:** *"**simplicity in setup** and the fact that it **can be used in pretty much any use case.** You simply install a second cluster and set up a mirroring process... **No need to worry about access to data, handling conflicts, and other architectural complexities.**"*
#### ⚠️ The two disadvantages — and the second one is brutal [#️-the-two-disadvantages--and-the-second-one-is-brutal]
**And the practice requirement:**
> *"it should go without saying that **whichever failover method you choose, YOUR SRE TEAM MUST PRACTICE IT ON A REGULAR BASIS. A plan that works today may stop working after an upgrade**, or perhaps new use cases make the existing tooling obsolete. **Once a quarter is usually the BARE MINIMUM** for failover practices. Strong SRE teams practice far more frequently. **Netflix's famous Chaos Monkey**, a service that randomly causes disasters, is the extreme — **any day may become failover practice day.**"*
***
# 10.9 Other cross-cluster mirroring solutions (/docs/kafka/cross-cluster-data-mirroring/other-cross-cluster-mirroring)
> *"MirrorMaker **also has some limitations when used in practice.** It is worthwhile to look at some of the alternatives and the ways they address MirrorMaker limitations and complexities."*
#### 9.1 Uber uReplicator [#91-uber-ureplicator]
**The problem Uber hit with legacy MirrorMaker:**
**The solution:** *"Uber decided to use **Apache Helix** as a **central (but highly available) controller** to manage the topic list and the partitions assigned to each uReplicator instance. Administrators use a **REST API** to add new topics to the list in Helix... Uber replaced the Kafka consumers with a **'Helix consumer'** — **this consumer takes its partition assignment FROM THE APACHE HELIX CONTROLLER rather than as a result of an agreement between the consumers.** As a result, the Helix consumer **can AVOID REBALANCES and instead LISTEN TO CHANGES in the assigned partitions that arrive from Helix.**"*
**The verdict:** *"uReplicator's **dependency on Apache Helix introduces A NEW COMPONENT TO LEARN AND MANAGE, adding complexity** to any deployment. As we saw earlier, **MirrorMaker 2.0 solves many of these scalability and fault-tolerance issues WITHOUT ANY EXTERNAL DEPENDENCIES.**"*
*(MM2's "allocate partitions without the consumer group protocol" is the same idea, internalized.)*
#### 9.2 LinkedIn Brooklin [#92-linkedin-brooklin]
> *"LinkedIn built a mirroring solution on top of its data streaming system called **Brooklin**. Brooklin is a **distributed service that can stream data between different HETEROGENEOUS data source and target systems**, including Kafka."*
**Three use cases:**
> *"designed for high reliability and has been **tested with Kafka at scale. It is used to mirror TRILLIONS OF MESSAGES A DAY** and has been optimized for **stability, performance, and operability.** Brooklin comes with a **REST API** for management operations. It is **a SHARED SERVICE that can process a large number of data pipelines, enabling the same service to mirror data across multiple Kafka clusters.**"*
#### 9.3 Confluent's three solutions [#93-confluents-three-solutions]
> *"At the same time that Uber developed its uReplicator, Confluent independently developed **Confluent Replicator. Despite the similarities in names, the projects have ALMOST NOTHING IN COMMON — they are different solutions to two different sets of MirrorMaker problems.**"*
##### Confluent Replicator [#confluent-replicator]
| | MirrorMaker 2.0 | Confluent Replicator |
| ------------------------------ | ------------------------- | -------------------------------------------------------------------------------------------- |
| Basis | Kafka Connect | Kafka Connect *"can run on existing Connect clusters"* |
| Data replication + topologies | ✅ | ✅ |
| Consumer offset migration | ✅ | ✅ |
| Topic config migration | ✅ | ✅ |
| **ACL migration** | ✅ | ❌ *"Replicator doesn't migrate ACLs"* |
| **Offset translation** | ✅ *"for any client"* | ⚠️ *"(using timestamp interceptor) **only for Java clients**"* |
| **Local/remote topic concept** | ✅ | ❌ *"but it supports **aggregate topics**"* |
| **Cycle prevention** | topic-name prefixing | *"using **provenance headers**"* |
| Monitoring | Connect + MM metrics | *"replication lag... REST API or **Control Center UI**"* |
| **Schema migration** | ❌ | ✅ *"supports schema migration between clusters and **can perform schema translation**"* |
##### Multi-Region Clusters (MRC) — a smarter stretch cluster [#multi-region-clusters-mrc--a-smarter-stretch-cluster]
That last mechanism is elegant: it **automatically trades throughput for availability exactly when needed**, and trades back when the emergency ends — precisely the manual decision that `unclean.leader.election.enable` forces on you in open-source Kafka (Ch. 7 §3.2).
##### Cluster Linking — offset-preserving mirroring [#cluster-linking--offset-preserving-mirroring]
> *"introduced as a preview feature in Confluent Platform 6.0, **builds inter-cluster replication DIRECTLY INTO the Confluent Server. By using THE SAME PROTOCOL AS INTER-BROKER REPLICATION within a cluster, Cluster Linking performs OFFSET-PRESERVING REPLICATION across clusters, enabling SEAMLESS MIGRATION OF CLIENTS WITHOUT ANY NEED FOR OFFSET TRANSLATION.**"*
#### Solution selection matrix [#solution-selection-matrix]
***
# 10.2 The realities of cross-datacenter communication (/docs/kafka/cross-cluster-data-mirroring/realities-cross-datacenter-communication)
> *"The solutions we'll discuss **may seem overly complicated without understanding that they represent TRADE-OFFS IN THE FACE OF SPECIFIC NETWORK CONDITIONS.**"*
#### ⚠️ Why you can't just stretch a normal cluster over a WAN [#️-why-you-cant-just-stretch-a-normal-cluster-over-a-wan]
> *"Apache Kafka's brokers and clients were **designed, developed, tested, and tuned, ALL WITHIN A SINGLE DATACENTER. We assumed low latency and high bandwidth between brokers and clients. This is apparent in THE DEFAULT TIMEOUTS AND SIZING OF VARIOUS BUFFERS.** For this reason, **it is NOT RECOMMENDED (except in specific cases) to install some Kafka brokers in one datacenter and others in another datacenter.**"*
#### 💡 The central design insight: consume remotely, don't produce remotely [#-the-central-design-insight-consume-remotely-dont-produce-remotely]
**Why ③ is safe — this is the argument to internalize:**
> *"in the event of a **network partition that prevents a consumer from reading data, THE RECORDS REMAIN SAFE INSIDE THE KAFKA BROKERS until communications resume and consumers can read them. THERE IS NO RISK OF ACCIDENTAL DATA LOSS due to network partitions.**"*
**If you do have to produce remotely:** *"you need to account for higher latency and the potential for more network errors. You can handle the errors by **increasing the number of producer retries**, and handle the higher latency by **increasing the size of the buffers that hold records between attempts to send them.**"*
**And one more efficiency argument:**
> *"because bandwidth is limited, **if there are MULTIPLE applications in one datacenter that need to read data from Kafka brokers in another datacenter, we prefer to install a Kafka cluster in each datacenter and MIRROR THE NECESSARY DATA BETWEEN THEM ONCE rather than have MULTIPLE applications consume the same data across the WAN.**"*
#### The three guiding principles [#the-three-guiding-principles]
***
# 10.12 Self-test (/docs/kafka/cross-cluster-data-mirroring/self-test)
Distinguish "replication" from "mirroring" in Kafka's vocabulary.
Name the five use cases for cross-cluster mirroring. Which one reduces requirements on the *source* clusters, and how?
Give the three realities of cross-datacenter communication. Why does high latency make bandwidth harder to use?
Why is it not recommended to spread one ordinary cluster's brokers across datacenters?
Explain precisely why remote *consuming* is safer than remote *producing*.
State the three guiding principles for multicluster architecture.
Draw hub-and-spoke. What is its single defining limitation, and what does the bank example illustrate?
In hub-and-spoke, where do the mirroring processes run, and why?
What are the two benefits of active-active? Which type of failover does the second one enable?
Describe both conflict types in active-active and the standard mitigation for the first.
Explain the topic-namespace trick that prevents mirroring loops. What must consumers subscribe to?
What's the alternative to namespaces, and what's the catch?
Give both disadvantages of active-standby. Quote the "bottom line" about Kafka failover.
Why is a smaller DR cluster a risky economy?
Define RTO and RPO. What does each one force on your architecture at its extreme?
Do the unplanned-failover data-loss arithmetic for 1M msg/s and 5 ms lag. How does planned failover differ?
Why can a line item arrive without its sale after failover?
List the four failover-offset strategies with their trade-offs.
Give all three caveats of mirroring `__consumer_offsets`, including the retention-skew example.
Why is time-based failover recommended despite being imprecise? Give the *social* argument.
How does offset translation store mappings efficiently? Walk the 495/500 → 596/600 example.
After a successful failover, why can't you just reverse the mirroring direction? What's the remedy?
Why do most failover scenarios require bouncing consumers, and what would avoid it?
Why does a stretch cluster need three datacenters rather than two? What is 2.5 DC?
What does a stretch cluster protect against — and what does it *not* protect against?
What was MM1's central flaw? What are MM2's three key design decisions?
What does MM2 migrate besides data records?
What happens to a topic's name when it's mirrored, and what two problems does that solve?
Which topic config is *not* migrated by default, and which ACL is deliberately not migrated? What's the failover consequence of each?
Give MirrorMaker's full ACL requirements, source and target.
Where should MirrorMaker run by default? Name the two exceptions and the reason for each.
Why do consumers suffer more from SSL than producers? What must you configure if you produce remotely because of it?
Give both lag-monitoring methods, why each is inaccurate, and the gap neither one covers.
Describe the procedure for finding the right `tasks.max`. What resource must you watch, and why?
How do you determine whether MirrorMaker's producer or consumer is the bottleneck? Give two methods.
Why is `max.in.flight=1` recommended for MirrorMaker when Ch. 3 said idempotence solves this? What does it cost over a WAN?
What problem did uReplicator solve, how, and what did MM2 do about it?
Compare MM2 and Confluent Replicator on ACL migration, offset translation, cycle prevention, and schema handling.
What is an "observer" in MRC, and what does automatic observer promotion accomplish?
How does Cluster Linking preserve offsets? Why must mirror topics be read-only, and what performance cost does it avoid?
**Previous:** [Chapter 9 — Building Data Pipelines](09-building-data-pipelines.md)
**Next:** [Chapter 11 — Securing Kafka](11-securing-kafka.md)
# 10.5 Stretch clusters — the synchronous option (/docs/kafka/cross-cluster-data-mirroring/stretch-clusters-synchronous-option)
> *"Stretch clusters are **fundamentally different** from other multidatacenter scenarios. **To start with, THEY ARE NOT MULTICLUSTER — IT IS JUST ONE CLUSTER.** As a result, **we don't need a mirroring process** to keep two clusters in sync. **Kafka's NORMAL REPLICATION mechanism is used, as usual.**"*
**How it achieves synchronous cross-DC durability:**
> *"we can configure things so **the acknowledgment will be sent AFTER the message is written successfully to Kafka brokers in TWO DATACENTERS.** This involves using **rack definitions** to make sure each partition has replicas in multiple datacenters, and the use of **`min.insync.replicas` and `acks=all`.**"*
**Plus follower fetching (2.4.0+):** *"brokers can also be configured to enable **consumers to fetch from the CLOSEST replica** using rack definitions. Brokers match their rack with that of the consumer to find the local replica that is most up-to-date, **falling back to the leader if a suitable local replica is not available. Consumers fetching from followers in their LOCAL datacenter achieve HIGHER THROUGHPUT, LOWER LATENCY, and LOWER COST by reducing cross-datacenter traffic.**"* (Ch. 4 §6.6, Ch. 6 §4.2.)
**Advantages:**
**Limitations:**
#### ⚠️ Why THREE datacenters, not two — the ZooKeeper quorum argument [#️-why-three-datacenters-not-two--the-zookeeper-quorum-argument]
**Feasibility:** *"if you can install Kafka (and ZooKeeper) in **at least three datacenters with HIGH BANDWIDTH and LOW LATENCY** between them. This can be done if your company **owns three buildings on the same street**, or — **more commonly — by using THREE AVAILABILITY ZONES inside ONE REGION of your cloud provider.**"*
> ### 2.5 DC ARCHITECTURE [#25-dc-architecture]
>
> *"A popular model for stretch clusters is a **2.5 DC** architecture with **both Kafka and ZooKeeper running in TWO datacenters**, and a third **'0.5' datacenter with ONE ZooKeeper node to provide quorum if a datacenter fails.**"*
>
> *(Also possible: ZooKeeper and Kafka in two datacenters using a **ZooKeeper group configuration that allows for MANUAL failover** — *"However, this setup is uncommon."*)*
***
# 10.8 Tuning MirrorMaker (/docs/kafka/cross-cluster-data-mirroring/tuning-mirrormaker)
#### 8.1 Sizing the cluster [#81-sizing-the-cluster]
#### 8.2 Finding the right `tasks.max` — an actual procedure [#82-finding-the-right-tasksmax--an-actual-procedure]
*(The decompress/recompress cost is the same broker-side cost from Ch. 6 §5.5 — and it's why Cluster Linking (§9) is faster: it avoids the round trip entirely.)*
#### 8.3 Isolate your sensitive topics [#83-isolate-your-sensitive-topics]
> *"you may want to **separate sensitive topics — those that absolutely require LOW LATENCY and where the mirror must be as close to the source as possible — to a SEPARATE MIRRORMAKER CLUSTER. This will PREVENT A BLOATED TOPIC OR AN OUT-OF-CONTROL PRODUCER FROM SLOWING DOWN YOUR MOST SENSITIVE DATA PIPELINE.**"*
#### 8.4 TCP stack tuning (cross-datacenter) [#84-tcp-stack-tuning-cross-datacenter]
> *"tuning the Linux network is **a large and complex topic**"* — recommended reading: **Performance Tuning for Linux Servers** by Sandra K. Johnson et al. (IBM Press).
*(`tcp_slow_start_after_idle=0` is the WAN-specific one: it stops TCP from resetting its congestion window after idle periods, which matters enormously on high-latency links.)*
#### 8.5 💡 Diagnosing the bottleneck: producer or consumer? [#85--diagnosing-the-bottleneck-producer-or-consumer]
#### 8.6 Tuning the producer [#86-tuning-the-producer]
| Config | When to change it — **based on a metric** |
| ------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`linger.ms` / `batch.size`** | *"If monitoring shows the producer consistently sends **partially empty batches** (i.e., `batch-size-avg` and `batch-size-max` are **lower than configured `batch.size`**), you can increase throughput by **introducing a bit of latency** — increase `linger.ms`."* Conversely: *"If you are sending **full batches** and have memory to spare, **increase `batch.size`.**"* |
| **`max.in.flight.requests.per.connection`** | See below — a real correctness/throughput trade |
##### ⚠️ The MirrorMaker ordering caveat [#️-the-mirrormaker-ordering-caveat]
> \*"**Limiting the number of in-flight requests to 1 is currently THE ONLY WAY FOR MIRRORMAKER TO GUARANTEE THAT MESSAGE ORDERING IS PRESERVED if some messages require multiple retries** before they are successfully acknowledged. But this means **every request sent by the producer has to be acknowledged by the target cluster before the next message is sent. THIS CAN LIMIT THROUGHPUT, ESPECIALLY IF THERE IS SIGNIFICANT LATENCY** before the brokers acknowledge.
>
> **If message order is NOT critical for your use case, using the default value of 5 can SIGNIFICANTLY INCREASE YOUR THROUGHPUT.**"\*
#### 8.7 Tuning the consumer [#87-tuning-the-consumer]
| Config | The metric that tells you to change it |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`fetch.max.bytes`** | *"If `fetch-size-avg` and `fetch-size-max` are **close to `fetch.max.bytes`**, the consumer is reading **as much as it is allowed.** If you have memory, **increase it.**"* |
| **`fetch.min.bytes` + `fetch.max.wait.ms`** | *"If **`fetch-rate` is HIGH**, the consumer is **sending too many requests and not receiving enough data in each.** Increase **both** so the consumer receives more per request and the broker waits until enough data is available."* |
*(Every tuning recommendation in this section is metric-driven. That's the model to copy: never tune a config without the metric that justifies it.)*
***
# 8.4 3.6 ⚠️ What transactions do NOT solve (/docs/kafka/exactly-once-semantics/3-6-transactions-not)
> *"transactions were added to Kafka to provide **multipartition atomic WRITES (but not READS)** and to **fence zombie producers** in stream processing applications. As a result, they provide exactly-once guarantees **when used within chains of consume-process-produce stream processing tasks. In other contexts, transactions will either straight-out NOT WORK or will require ADDITIONAL EFFORT.**"*
> **"The two main mistakes are assuming that exactly-once guarantees apply on ACTIONS OTHER THAN PRODUCING TO KAFKA, and that consumers ALWAYS READ ENTIRE TRANSACTIONS and have information about TRANSACTION BOUNDARIES."**
#### ① Side effects while stream processing [#-side-effects-while-stream-processing]
> *"Let's say that the record processing step includes **sending email to users.** Enabling exactly-once semantics **will NOT guarantee that the email will only be sent once. The guarantee only applies to records written to Kafka.** Using sequence numbers to deduplicate records or using markers to abort or cancel a transaction works within Kafka, **but IT WILL NOT UN-SEND AN EMAIL.** The same is true for any action with external effects: **calling a REST API, writing to a file, etc.**"*
#### ② Reading from Kafka, writing to a database [#-reading-from-kafka-writing-to-a-database]
> *"the application is writing to an **external database** rather than to Kafka. In this scenario, **THERE IS NO PRODUCER INVOLVED** — records are written to the database using a database driver (likely JDBC) and offsets are committed to Kafka within the consumer. **There is NO MECHANISM that allows writing results to an external database and committing offsets to Kafka within a single transaction.**"*
**The workaround — flip which system owns the transaction:**
> *"Instead, we could **manage offsets IN THE DATABASE** (as explained in Chapter 4) and **commit both data and offsets to the database in a single transaction** — this would rely on **the DATABASE's transactional guarantees rather than Kafka's.**"*
> ### NOTE — the OUTBOX PATTERN [#note--the-outbox-pattern]
>
> *"Microservices often need to **update the database AND publish a message to Kafka within a single atomic transaction**, so either both will happen or neither will. **Kafka transactions will NOT do this.**"*
>
> **Direction 1 — Kafka is the outbox:**
>
> ```
> microservice ──► publishes ONLY to a Kafka topic (the "outbox")
> │
> ▼
> a separate MESSAGE RELAY SERVICE
> │
> ▼
> updates the database
>
> ⚠ "Because Kafka won't guarantee an exactly-once update to the database,
> IT IS IMPORTANT TO MAKE SURE THE UPDATE IS IDEMPOTENT."
> ► Guarantees: "the message will EVENTUALLY make it to Kafka, the topic
> consumers, AND the database — OR TO NONE OF THOSE."
> ```
>
> **Direction 2 — a database table is the outbox:**
>
> ```
> microservice ──► writes data AND an outbox row in ONE DB transaction
> │
> ▼
> a relay service reads the outbox table
> │
> ▼
> produces to Kafka
>
> "This pattern is PREFERRED when built-in RDBMS constraints, such as
> UNIQUENESS and FOREIGN KEYS, are useful."
> ```
>
> Reference: *"The **Debezium** project published an in-depth blog post on the outbox pattern with detailed examples."*
#### ③ Database → Kafka → another database (end-to-end DB transactions) [#-database--kafka--another-database-end-to-end-db-transactions]
> *"**It is very tempting to believe** that we can build an app that will read data from a database, **identify database transactions**, write the records to Kafka, and from there write records to another database, **still maintaining the original transactions from the source database. Unfortunately, Kafka transactions don't have the necessary functionality to support these kinds of end-to-end guarantees.**"*
**Two independent reasons:**
#### ④ Copying data from one Kafka cluster to another [#-copying-data-from-one-kafka-cluster-to-another]
> *"This one is **more subtle** — **it IS possible to support exactly-once guarantees when copying data from one Kafka cluster to another.** There is a description of how this is done in the KIP for adding exactly-once capabilities in **MirrorMaker 2.0**. At the time of this writing, the proposal is **still in draft**, but the algorithm is clearly described. This proposal includes the guarantee that **each record in the source cluster will be copied to the destination cluster exactly once.**"*
> **"However, this does NOT guarantee that TRANSACTIONS WILL BE ATOMIC.** If an app produces several records and offsets transactionally, and then MirrorMaker 2.0 copies them to another cluster, **the transactional properties and guarantees WILL BE LOST during the copy process.**"\*
**Same root cause:** *"the consumer reading data from Kafka **can't know or guarantee that it is getting ALL the events in a transaction.** For example, **it can replicate PART of a transaction if it is only subscribed to a subset of the topics.**"*
```txt
exactly-once RECORD delivery ✓ possible (MM2 KIP)
preserving TRANSACTION atomicity across clusters ✗ not possible
```
#### ⑤ Publish/subscribe pattern [#-publishsubscribe-pattern]
> *"We've discussed exactly-once in the context of the consume-process-produce pattern, but **the publish/subscribe pattern is a very common use case.** Using transactions in a pub/sub use case provides **some** guarantees: consumers configured with `read_committed` will not see records that were published as part of an **aborted** transaction. **But those guarantees FALL SHORT of exactly-once. CONSUMERS MAY PROCESS A MESSAGE MORE THAN ONCE, DEPENDING ON THEIR OWN OFFSET COMMIT LOGIC.**"*
**The JMS comparison — and the important difference:**
> *"The guarantees Kafka provides in this case are **similar to those provided by JMS transactions but DEPEND ON CONSUMERS in `read_committed` mode** to guarantee that uncommitted transactions will remain invisible. **JMS brokers withhold uncommitted transactions from ALL consumers.**"*
> ### ⚠️ WARNING — the transaction deadlock [#️-warning--the-transaction-deadlock]
>
> *"**An important pattern to AVOID is publishing a message and then WAITING FOR ANOTHER APPLICATION TO RESPOND before committing the transaction. The other application WILL NOT RECEIVE THE MESSAGE UNTIL AFTER THE TRANSACTION WAS COMMITTED, resulting in a DEADLOCK.**"*
**Summary table:**
| Scenario | Exactly-once? |
| -------------------------------------------------------- | ------------------------------------------------------------------------- |
| Consume → process → produce **to Kafka** | ✅ **Yes** — this is what transactions are for |
| Any **external side effect** (email, REST, file) | ❌ No — cannot be rolled back |
| Kafka → **external database** | ❌ No producer involved → use the **DB's** transaction (offsets in the DB) |
| DB write + Kafka publish atomically | ❌ No → **outbox pattern** (either direction) |
| DB → Kafka → DB preserving source transactions | ❌ No — plus consumers have no transaction boundaries |
| Kafka cluster → Kafka cluster, **per record** | ⚠️ Yes (MM2 KIP, draft) |
| Kafka cluster → Kafka cluster, **transaction atomicity** | ❌ No |
| Plain **publish/subscribe** | ⚠️ Partial — aborted records hidden, but duplicates still possible |
***
# 8.5 3.7 Transactional IDs and fencing (/docs/kafka/exactly-once-semantics/3-7-transactional-ids)
> *"Choosing the transactional ID for producers is important and **a bit more challenging than it seems.** Assigning the transactional ID incorrectly can lead to **either application errors or LOSS OF EXACTLY-ONCE GUARANTEES.**"*
**The two key requirements:**
#### The pre-2.5 world: static partition mapping [#the-pre-25-world-static-partition-mapping]
> *"Until release 2.5, **the only way to guarantee fencing was to STATICALLY MAP THE TRANSACTIONAL ID TO PARTITIONS.** This guaranteed that **each partition will always be consumed with the same transactional ID.**"*
**Why anything else breaks:**
> **The book flags its own example:** *"In those releases, the previous example would be **INCORRECT** — transactional IDs are assigned randomly to threads without making sure the same transactional ID is always used to write to the same partition."*
#### KIP-447 (Kafka 2.5): fencing by consumer group metadata [#kip-447-kafka-25-fencing-by-consumer-group-metadata]
> *"In Apache Kafka 2.5, **KIP-447** introduced a **second method of fencing based on CONSUMER GROUP METADATA**, in addition to transactional IDs. We use the producer offset commit method and **pass as an argument the CONSUMER GROUP METADATA rather than just the consumer group ID.**"*
**The scenario — and why the old way was wasteful:**
**This is what makes the code example in §4 legal.** With `subscribe()`, *"partitions assigned to this instance can change at any point as a result of rebalance"* — impossible to reconcile with static transactional-ID↔partition mapping. Group-generation fencing removes that constraint.
***
# 8.7 What actually breaks in production — Ch. 8 consolidated (/docs/kafka/exactly-once-semantics/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 1 | **Duplicates even though nothing failed** | Leader wrote **and replicated**, then crashed before acking; the producer retried to the new leader | `enable.idempotence=true` |
| 2 | **Aggregates (averages, counts, balances) are subtly wrong and nobody can prove why** | A duplicate was folded into an aggregate — *"impossible to correct the result without reprocessing the input"* | Transactions / Kafka Streams `processing.guarantee` |
| 3 | **"Out of order sequence number" in the logs; team ignores it** | The producer *can* continue, **but the gap means messages 3–26 were LOST** | Review reliability config (Ch. 7); **check whether unclean leader election occurred** |
| 4 | **Duplicates after a producer restart, with idempotence on** | Each init gets a **brand-new PID** — the broker cannot correlate old and new producers | Use **`transactional.id`** for a stable identity across restarts |
| 5 | **A revived frozen producer wrote duplicates; not detected as a zombie** | *"we have two totally different producers with different IDs"* | Transactions + epoch fencing |
| 6 | **Duplicates from two app instances reading the same source file** | Idempotence is **per producer instance**, not a global dedup service | Partition the work; dedupe by business key upstream |
| 7 | **Duplicates from your own retry loop** | `producer.send()` called twice — *"the producer has no way of knowing that the two records are in fact the same"* | Rely on the built-in retry mechanism only |
| 8 | **Fatal `UNKNOWN_PRODUCER_ID`** | Pre-2.5 producer state not kept long enough; known **partition-reassignment edge case** where a new leader had no state | Upgrade to **2.5+** (KIP-360) |
| 9 | **Exactly-once configured, consumers still see aborted records** | **`isolation.level` defaults to `read_uncommitted`** — aborted records are physically in the log | `isolation.level=read_committed` on **every** downstream consumer |
| 10 | **`read_committed` consumers stall for up to 15 minutes** | A long-running open transaction holds back the **LSO**; everything after it is withheld until commit/abort or `transaction.timeout.ms` | Commit transactions frequently; tune `transaction.timeout.ms` |
| 11 | **End-to-end latency worse after enabling transactions** | Same LSO mechanism — read\_committed consumers always lag | Treat transaction duration as a latency budget |
| 12 | **Duplicate emails / double API charges despite exactly-once** | *"The guarantee only applies to records written to Kafka... it will not un-send an email"* | Make external effects **idempotent** (idempotency keys), or move them outside the transaction |
| 13 | **Can't atomically write to a DB and commit Kafka offsets** | *"There is no mechanism"* — no producer is involved | Manage **offsets in the database**; use the DB's transaction |
| 14 | **Microservice updated the DB but the Kafka message was lost (or vice versa)** | Expecting Kafka transactions to span systems | **Outbox pattern** (Kafka-as-outbox with an idempotent DB update, or table-as-outbox when you need RDBMS constraints) |
| 15 | **Source database transactions not preserved through Kafka** | Consumers have **no transaction boundary information** and may be lagging on some topics | Not solvable with Kafka transactions; redesign |
| 16 | **MirrorMaker copied records exactly once but transactions lost atomicity** | Cross-cluster copy can't guarantee it sees **all** events in a transaction — *"it can replicate part of a transaction if it is only subscribed to a subset of the topics"* | Accept per-record exactly-once only |
| 17 | **Pub/sub consumers still process messages twice** | Transactions don't govern **consumer offset commit logic** | You still need consumer-side discipline (Ch. 7 §5) |
| 18 | **Deadlock: producer waits for a reply that can never arrive** | Published inside a transaction, then waited for a `read_committed` consumer to respond **before committing** | Never block on a response before committing the transaction |
| 19 | **Zombie not fenced; duplicates in the output** | Transactional ID **changed** between the failed instance and its replacement (`A` → `B`) | Pre-2.5: statically map transactional ID → partitions. 2.5+: pass **`consumer.groupMetadata()`** to `sendOffsetsToTransaction()` |
| 20 | **`ProducerFencedException` / `InvalidProducerEpochException`** | **You are the zombie.** A newer instance holds your transactional ID | *"Nothing to do but die gracefully"* — don't retry |
| 21 | **Offsets committed outside the transaction; exactly-once silently broken** | Called `consumer.commitSync()`, or left autocommit on | `enable.auto.commit=false` **and never call consumer commit APIs**; offsets only via `sendOffsetsToTransaction()` |
| 22 | **Broker OOM / severe GC after weeks of normal operation** | **Producer-state accumulation** — new transactional/producer IDs created at a high rate, retained `transactional.id.expiration.ms` (**7 days**). *3/sec ⇒ 1.8M entries ⇒ \~5 GB* | Few **long-lived** producers; if FaaS makes that impossible, **lower `transactional.id.expiration.ms`** |
| 23 | **Transaction throughput poor** | Very small transactions — overhead is **per transaction**, and init/commit are **synchronous stops** | Batch more messages per transaction |
| 24 | **In-flight transactions left hanging after a crash** | Normal — resolved by design | `initTransactions()` **aborts older in-flight transactions**; the coordinator auto-aborts after `transaction.timeout.ms`; a new coordinator picks up logged intent |
| 25 | **Rebalance broke transactional processing (pre-2.5)** | Transactional producers needed **static** partition assignment; `subscribe()` reassigns freely | Kafka 2.5+ with group-metadata fencing; *"commit transactions whenever the related partitions are revoked"* |
| 26 | **Kafka Streams app doesn't scale with many partitions** | One transactional producer per partition-identity was required | `processing.guarantee=exactly_once_beta` (brokers 2.5+, Streams 2.6+) |
***
# 8.8 Deploy / monitor / configure (/docs/kafka/exactly-once-semantics/deploy-monitor-configure)
#### The configuration recipes [#the-configuration-recipes]
#### Monitoring [#monitoring]
| Signal | Meaning |
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------ |
| **`record-error-rate`** (producer) | Includes benign **duplicate rejections** — *"should not cause any alarm"*, but a rising trend is worth understanding |
| **`ErrorsPerSec`** of `RequestMetrics` (broker) | *"includes a separate count for each type of error"* — this is how you separate duplicate rejections from real errors |
| **"out of order sequence number" in producer logs** | ⚠️ **Message loss** between producer and broker — investigate config + unclean elections |
| **`UNKNOWN_PRODUCER_ID`** | Producer state expired or missing (pre-2.5 bug, or expiration too low) |
| **`ProducerFencedException` rate** | Zombie instances being fenced — expected during failover, alarming if constant |
| **read\_committed consumer lag vs read\_uncommitted** | Measures your **LSO stall** — i.e. how long your transactions stay open |
| **Broker heap / GC** | The **producer-state accumulation** leak (#22) shows up here first |
| **Transaction commit latency** | The synchronous stop in the produce path |
#### The relationship to Ch. 7 [#the-relationship-to-ch-7]
***
# 8.5 How transactions work internally (/docs/kafka/exactly-once-semantics/how-transactions-work-internally)
> *"We can use transactions by calling the APIs without understanding how they work. **But having some mental model of what is going on under the hood will help us troubleshoot applications that do not behave as expected.**"*
#### 5.1 The algorithm [#51-the-algorithm]
> *"The basic algorithm for transactions in Kafka was **inspired by Chandy-Lamport snapshots**, in which **'marker' control messages are sent into communication channels, and consistent state is determined based on the arrival of the marker.**"*
**The problem markers alone don't solve:** *"what happens if the producer crashes after only writing commit messages to a **subset** of the partitions?"*
**The answer: two-phase commit + a transaction log.**
**The transaction log:** an internal topic called **`__transaction_state`**.
#### 5.2 Walking the API calls [#52-walking-the-api-calls]
**Two things worth noticing:**
1. **`beginTransaction()` is purely client-side.** The broker learns about the transaction lazily, per-partition, via `AddPartitionsToTxnRequest`. So a transaction that never sends anything never exists on the broker.
2. **Once the intent is logged, completion is guaranteed by the coordinator, not by your process.** That's what makes it a real two-phase commit rather than best-effort marker writing.
#### 5.3 ⚠️ The producer-state memory leak — a genuine production hazard [#53-️-the-producer-state-memory-leak--a-genuine-production-hazard]
> *"Each broker that receives records from transactional or idempotent producers will **store the producer/transactional IDs IN MEMORY**, together with related state for each of the **last five batches** sent by the producer: sequence numbers, offsets, and such. **This state is stored for `transactional.id.expiration.ms` milliseconds AFTER the producer stopped being active (SEVEN DAYS BY DEFAULT).** This allows the producer to resume activity without running into `UNKNOWN_PRODUCER_ID` errors."*
> **"It is possible to cause something similar to a MEMORY LEAK in the broker by creating new idempotent producers or new transactional IDs at a very high rate but NEVER REUSING THEM."**
**The arithmetic — do this calculation for your own workload:**
Three new producers per second is *nothing* — that's a modest serverless workload or a badly-written client that constructs a producer per request.
**The recommendations:**
**Note the FaaS callout explicitly.** Lambda-style architectures that spin up a producer per invocation are the canonical way to hit this.
***
# 8.4 How to use transactions (/docs/kafka/exactly-once-semantics/how-use-transactions)
#### 4.1 The recommended way — don't use them directly [#41-the-recommended-way--dont-use-them-directly]
> *"The **most common and most recommended** way to use transactions is to **enable exactly-once guarantees in Kafka Streams.** This way, **we will not use transactions directly at all**, but rather Kafka Streams will use them for us behind the scenes... **Transactions were designed with this use case in mind, so using them via Kafka Streams is the easiest and most likely to work as expected.**"*
```properties
processing.guarantee=exactly_once
# or
processing.guarantee=exactly_once_beta
```
> *"That's it."*
> ### NOTE — `exactly_once_beta` [#note--exactly_once_beta]
>
> *"a slightly different method of handling application instances that **crash or hang with in-flight transactions.** Introduced in **release 2.5 to Kafka brokers, and release 2.6 to Kafka Streams.** **The main benefit of this method is the ability to handle MANY PARTITIONS WITH A SINGLE TRANSACTIONAL PRODUCER and therefore create MORE SCALABLE Kafka Streams applications.**"*
(This is the practical payoff of KIP-447's group-metadata fencing.)
#### 4.2 Using the transactional API directly [#42-using-the-transactional-api-directly]
```java
Properties producerProps = new Properties();
producerProps.put(ProducerConfig.BOOTSTRAP_SERVERS_CONFIG, "localhost:9092");
producerProps.put(ProducerConfig.CLIENT_ID_CONFIG, "DemoProducer");
producerProps.put(ProducerConfig.TRANSACTIONAL_ID_CONFIG, transactionalId); // ①
producer = new KafkaProducer<>(producerProps);
Properties consumerProps = new Properties();
consumerProps.put(ConsumerConfig.BOOTSTRAP_SERVERS_CONFIG, "localhost:9092");
consumerProps.put(ConsumerConfig.GROUP_ID_CONFIG, groupId);
props.put(ConsumerConfig.ENABLE_AUTO_COMMIT_CONFIG, "false"); // ②
consumerProps.put(ConsumerConfig.ISOLATION_LEVEL_CONFIG, "read_committed"); // ③
consumer = new KafkaConsumer<>(consumerProps);
producer.initTransactions(); // ④
consumer.subscribe(Collections.singleton(inputTopic)); // ⑤
while (true) {
try {
ConsumerRecords records =
consumer.poll(Duration.ofMillis(200));
if (records.count() > 0) {
producer.beginTransaction(); // ⑥
for (ConsumerRecord record : records) {
ProducerRecord customizedRecord = transform(record); // ⑦
producer.send(customizedRecord);
}
Map offsets = consumerOffsets();
producer.sendOffsetsToTransaction(offsets, consumer.groupMetadata()); // ⑧
producer.commitTransaction(); // ⑨
}
} catch (ProducerFencedException|InvalidProducerEpochException e) { // ⑩
throw new KafkaException(String.format(
"The transactional.id %s is used by another process", transactionalId));
} catch (KafkaException e) { // ⑪
producer.abortTransaction();
resetToLastCommittedPositions(consumer);
}}
```
**Annotations:**
**①** *"Configuring a producer with `transactional.id` makes it a transactional producer capable of producing atomic multipartition writes. **The transactional ID must be UNIQUE and LONG-LIVED. Essentially it defines an INSTANCE OF THE APPLICATION.**"*
**②** *"Consumers that are part of the transactions **DON'T COMMIT THEIR OWN OFFSETS — the PRODUCER writes offsets as part of the transaction. So offset commit should be disabled.**"*
**③** *"To read transactions cleanly (i.e., ignore in-flight and aborted transactions), we set the consumer isolation level to `read_committed`. **Note that the consumer will still read NONTRANSACTIONAL WRITES, in addition to reading committed transactions.**"* (The input needn't be transactional — *"there is no such requirement for the input."*)
**④** *"The first thing a transactional producer must do is initialize. This:*
**⑤** *"Here we are using the `subscribe` consumer API, which means that **partitions assigned to this instance can change at any point as a result of rebalance.** Prior to release 2.5 (KIP-447), this was much more challenging..."* (see §3.7). *"When using this method, **it also makes sense to commit transactions whenever the related partitions are revoked.**"*
**⑥** *"This method guarantees that **everything that is produced from the time it was called, until the transaction is either committed or aborted, is part of a single atomic transaction.**"*
**⑦** *"This is where we process the records — all our business logic goes here."*
**⑧** ⚠️ *"it is important to commit the offsets **as part of the transaction.** This guarantees that **if we fail to produce results, we won't commit the offsets for records that were not, in fact, processed.** **Note that it is important NOT to commit offsets in ANY OTHER WAY — disable offset auto-commit, and DON'T CALL ANY OF THE CONSUMER COMMIT APIs. Committing offsets by any other method does not provide transactional guarantees.**"*
**⑨** *"Once this method returns successfully, **the entire transaction has made it through**, and we can continue to read and process the next batch."*
**⑩** *"**If we got this exception, it means WE ARE THE ZOMBIE.** Somehow our application froze or disconnected, and there is a newer instance of the app with our transactional ID running. **Most likely the transaction we started has already been aborted and someone else is processing those records. NOTHING TO DO BUT DIE GRACEFULLY.**"*
**⑪** *"If we got an error while writing a transaction, we can **abort the transaction, set the consumer position back, and try again.**"*
*(Full example: Apache Kafka GitHub, *"which includes a demo driver and a simple exactly-once processor that runs in separate threads."*)*
***
# 8.2 The idempotent producer (/docs/kafka/exactly-once-semantics/idempotent-producer)
#### 2.1 What "idempotent" means [#21-what-idempotent-means]
> *"A service is called idempotent if **performing the same operation multiple times has the same result as performing it a single time.**"*
**The database illustration:**
#### 2.2 The classic duplicate — nothing actually failed [#22-the-classic-duplicate--nothing-actually-failed]
**The consequences, in the book's own concrete terms:** *"in others they can lead to **inventory miscounts, bad financial statements, or sending someone two umbrellas instead of the one they ordered.**"*
**Why retry logic can never solve this alone:** the producer cannot distinguish *"the write failed"* from *"the write succeeded and the ack was lost."* Only the **broker** can, because only the broker knows what it already has. Hence the fix lives at the protocol level.
#### 2.3 How it works — PID + sequence number [#23-how-it-works--pid--sequence-number]
**That's why the 5-in-flight limit exists.** It isn't arbitrary — it's the size of the broker's dedup window.
**Duplicate detection:**
> *"When a broker receives a message that it already accepted before, **it will reject the duplicate with an appropriate error.** This error is **logged by the producer and is reflected in its metrics but DOES NOT CAUSE ANY EXCEPTION and SHOULD NOT CAUSE ANY ALARM.**"*
| Where | Metric |
| ------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Producer client** | added to **`record-error-rate`** |
| **Broker** | part of **`ErrorsPerSec`** of the **`RequestMetrics`** type — *"which includes a separate count for each type of error"* |
#### 2.4 ⚠️ "Out of order sequence number" — an error you should investigate, not ignore [#24-️-out-of-order-sequence-number--an-error-you-should-investigate-not-ignore]
> *"The broker expects message number 2 to be followed by message number 3; what happens if the broker receives **message number 27** instead? In such cases the broker will respond with an **'out of order sequence' error**, but **if we use an idempotent producer without using transactions, this error CAN BE IGNORED.**"*
> ### ⚠️ WARNING — but it's telling you something [#️-warning--but-its-telling-you-something]
>
> \*"While the producer will continue normally after encountering an 'out of order sequence number' exception, **this error typically indicates that MESSAGES WERE LOST between the producer and the broker** — if the broker received message number 2 followed by message number 27, **something must have happened to messages 3 to 26.**
>
> When encountering such an error in the logs, it is worth:
>
> * **revisiting the producer and topic configuration** and making sure the producer is configured with **recommended values for high reliability**, and
> * **checking whether UNCLEAN LEADER ELECTION has occurred.**"\*
#### 2.5 Behavior under failure — the two cases [#25-behavior-under-failure--the-two-cases]
##### Case A: Producer restart — ⚠️ idempotence does NOT survive it [#case-a-producer-restart--️-idempotence-does-not-survive-it]
> *"when the producer starts, if the idempotent producer is enabled, the producer will initialize and reach out to a Kafka broker to generate a producer ID. **EACH INITIALIZATION OF A PRODUCER WILL RESULT IN A COMPLETELY NEW ID** (assuming that we did not enable transactions)."*
**This is precisely the gap `transactional.id` fills** (§3.4): a *stable* identity across restarts, which a randomly-assigned PID can never provide.
##### Case B: Broker failure — idempotence *does* survive it [#case-b-broker-failure--idempotence-does-survive-it]
The setup: producer writes to topic A partition 0; leader on broker 5, follower on broker 3. Broker 5 fails, broker 3 becomes leader. *"But how will broker 3 know which sequences were already produced in order to reject duplicates?"*
**Four layers of producer-state durability:**
*(④ is why Ch. 6's batch header carries producer ID, producer epoch, and first sequence — the log is the ultimate source of truth for dedup state.)*
**The edge case with a satisfying answer:**
> *"what happens if **there are no messages**? Imagine that a topic has two hours of retention, but no new messages arrived in the last two hours — there will be **no messages to use to recover the state** if a broker crashed. **Luckily, NO MESSAGES ALSO MEANS NO DUPLICATES.** We will start accepting messages immediately (**while logging a warning about the lack of state**), and create the producer state from the new messages that arrive."*
#### 2.6 ⚠️ Limitations — exactly what idempotence does NOT cover [#26-️-limitations--exactly-what-idempotence-does-not-cover]
##### Limitation 1: Calling `send()` twice yourself [#limitation-1-calling-send-twice-yourself]
> *"Calling `producer.send()` twice with the same message **will create a duplicate**, and the idempotent producer won't prevent it. This is because **the producer has no way of knowing that the two records that were sent are in fact the same record.**"*
> *"It is always a good idea to use the built-in retry mechanism of the producer rather than catching producer exceptions and retrying from the application itself; **the idempotent producer makes this pattern EVEN MORE APPEALING — it is the easiest way to avoid duplicates when retrying.**"*
*(Same conclusion as Ch. 7 §4.4, now with a second reason: your retry loop is invisible to dedup; the producer's is not.)*
##### Limitation 2: Multiple producer instances [#limitation-2-multiple-producer-instances]
> *"It is also rather common to have applications that have **multiple instances or even one instance with multiple producers.** If two of these producers attempt to send identical messages, **the idempotent producer will not detect the duplication.**"*
**The concrete example:**
#### 2.7 How to use it [#27-how-to-use-it]
```properties
enable.idempotence=true
```
> *"This is the easy part... **If the producer is already configured with `acks=all`, there will be NO DIFFERENCE IN PERFORMANCE.**"*
**Four things change:**
| Change | Detail |
| ----------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **One extra API call at startup** | to retrieve a producer ID |
| **96 bits added per record batch** | producer ID (**a long**) + sequence ID for the **first** message in the batch (*"sequence IDs for each message in the batch are derived from the sequence ID of the first message plus a delta"*) — *"barely any overhead for most workloads"* |
| **Brokers validate sequence numbers** | *"from any single producer instance and guarantee the lack of duplicate messages"* |
| **Ordering guaranteed through all failure scenarios** | *"even if `max.in.flight.requests.per.connection` is set to more than 1 (**5 is the default and also the highest value supported by the idempotent producer**)"* |
**That last row is the one people undervalue.** Without idempotence, `retries>0 + max.in.flight>1` silently reorders (Ch. 3 §6). Idempotence gives you **ordering AND dedup AND 5 in-flight requests** simultaneously — there is no reason not to enable it.
> ### NOTE — KIP-360 improvements in 2.5 [#note--kip-360-improvements-in-25]
>
> *"Idempotent producer logic and error handling **improved significantly in version 2.5** (both producer and broker side) as a result of **KIP-360**."*
>
> **Before 2.5:**
>
> * *"the producer state was **not always maintained for long enough**, which resulted in **fatal `UNKNOWN_PRODUCER_ID` errors** in various scenarios."*
> * A **known edge case with partition reassignment**: *"the new replica became the leader **before any writes happened from a specific producer**, meaning that the new leader had **no state for that partition.**"*
> * *"previous versions attempted to **rewrite the sequence IDs** in some error scenarios, **which could lead to duplicates.**"*
>
> **In newer versions:** *"if we encounter a fatal error for a record batch, **this batch and all the batches that are in flight will be rejected. The user who writes the application can handle the exception and decide whether to skip those records or retry and risk duplicates and reordering.**"*
***
# 8. Exactly-Once Semantics (/docs/kafka/exactly-once-semantics)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 8*
> **The chapter's own summary, which is the best one-liner in the book:** *"**Exactly-once semantics in Kafka is the opposite of chess: it is challenging to understand but easy to use.**"*
**Two mechanisms, two different problems:**
***
# 8.1 Why at-least-once isn't always enough (/docs/kafka/exactly-once-semantics/least-once-isn-t)
Ch. 7 delivered **at-least-once** — *"the guarantee that Kafka will not lose messages that it acknowledged as committed. **This still leaves open the possibility of duplicate messages.**"*
**When duplicates are fine:**
> *"In simple systems where messages are produced and then consumed by various applications, duplicates are **an annoyance that is fairly easy to handle. Most real-world applications contain unique identifiers that consuming applications can use to deduplicate.**"*
**When they are not — and this is the whole argument for this chapter:**
> *"Things become more complicated when we look at **stream processing applications that AGGREGATE events.** When inspecting an application that consumes events, computes an average, and produces the results, **it is often IMPOSSIBLE for those who check the results to DETECT that the average is incorrect because an event was processed twice.**"*
This is the key asymmetry: **aggregation destroys the evidence.** A duplicated record leaves a fingerprint you can dedupe; a duplicated *contribution to a sum* leaves nothing.
***
# 8.6 Performance of transactions (/docs/kafka/exactly-once-semantics/performance-transactions)
#### Producer side — overhead is per-transaction, not per-message [#producer-side--overhead-is-per-transaction-not-per-message]
> ### 💡 The key optimization [#-the-key-optimization]
>
> *"the overhead of transactions on the producer is **INDEPENDENT OF THE NUMBER OF MESSAGES IN A TRANSACTION.** So **a larger number of messages per transaction will BOTH reduce the relative overhead AND reduce the number of synchronous stops, resulting in HIGHER THROUGHPUT overall.**"*
#### Consumer side — no throughput cost, but a latency cost [#consumer-side--no-throughput-cost-but-a-latency-cost]
#### ⚠️ The transaction-size tension [#️-the-transaction-size-tension]
Note that the *producer's* natural unit is "one transaction per `poll()` batch," which conveniently ties transaction size to `max.poll.records` — giving you one knob that moves both.
***
# 8.9 Self-test (/docs/kafka/exactly-once-semantics/self-test)
Why are duplicates "an annoyance" in simple pipelines but a correctness disaster in aggregations?
Give the classic duplicate scenario. What is the crucial detail about whether the original write succeeded?
Why can no amount of producer-side retry logic solve that scenario? What must know the answer?
What four values uniquely identify a message for idempotence? How many messages does a broker track per partition?
Why does idempotence require `max.in.flight.requests ≤ 5`?
Duplicate rejections show up in which producer metric and which broker metric?
The broker reports "out of order sequence number." Why can you ignore it, and why should you absolutely not?
Does idempotence survive a producer restart? Explain precisely why or why not.
List the four layers of producer-state durability that let idempotence survive broker failure and restart.
A broker crashes and there are no messages in the retention window to rebuild state from. Why is this fine?
Give the two hard limitations of the idempotent producer, with a concrete example of each.
What exactly changes when you set `enable.idempotence=true`? What's the performance cost if `acks=all` is already set?
Which ordering problem from Ch. 3 does idempotence solve, and what does that let you keep that the folklore fix takes away?
Distinguish "transactions" from "exactly-once semantics." Which do Flink and Spark Streaming use?
Describe both problems transactions solve — crashes and zombies — and note why they need different mechanisms.
What's the insight that makes atomic offset-commit-plus-produce possible?
Contrast `producer.id` and `transactional.id` on origin, lifetime, and purpose.
How does epoch-based zombie fencing work here, and where else in Kafka have you seen the identical pattern?
`isolation.level` defaults to what, and why is that default fatal to exactly-once expectations?
What is the LSO? Draw what a `read_committed` consumer sees when a transaction is left open, and state the worst-case stall.
Name three things `read_committed` does *not* tell a consumer.
Your processing step sends an email. What does exactly-once guarantee about that email?
You need to write to a database and commit Kafka offsets atomically. What can't you do, and what should you do?
Describe both directions of the outbox pattern, and when you'd prefer the table-as-outbox variant.
Why can't Kafka preserve source-database transaction boundaries through a Kafka hop? Give both reasons.
MirrorMaker 2.0 can copy records exactly once. What can it *not* preserve, and why?
Compare Kafka's `read_committed` to JMS transaction visibility. Who enforces invisibility in each?
Describe the transaction deadlock pattern and why it's unavoidable once you enter it.
Producer A fails and is replaced by producer B; A returns as a zombie. Why isn't A fenced, and what are the two eras' fixes?
What does KIP-447 add, and what wasteful workaround does it eliminate?
What are the three configuration requirements for exactly-once with the raw API? Which one is easiest to get wrong?
What should you do on `ProducerFencedException`?
Give the four steps of the transaction algorithm. At which step are you "doomed to commit or abort eventually," and why does that matter?
What does `beginTransaction()` actually send to the broker? How does the coordinator learn which partitions are involved?
The transaction coordinator crashes after logging commit intent. What happens?
Do the producer-state memory arithmetic for 3 new producers/second over a week. What are the two mitigations?
Why does a larger transaction improve producer throughput, and what does it cost on the consumer side?
Transactions add latency but not throughput cost on the consumer. Explain both halves.
**Previous:** [Chapter 7 — Reliable Data Delivery](07-reliable-data-delivery.md)
**Next:** [Chapter 9 — Building Data Pipelines](09-building-data-pipelines.md)
# 8.3 Transactions (/docs/kafka/exactly-once-semantics/transactions)
#### 3.1 What they were built for — and the naming distinction [#31-what-they-were-built-for--and-the-naming-distinction]
> *"transactions were added to Kafka to **guarantee the correctness of applications developed using Kafka Streams.** In order for a stream processing application to generate correct results, **each input record must be processed exactly one time, and its processing result will be reflected exactly one time, even in case of failure.**"*
> ### NOTE — vocabulary that matters [#note--vocabulary-that-matters]
>
> *"**Transactions** is the name of the **underlying MECHANISM. Exactly-once semantics** or **exactly-once guarantees** is the **BEHAVIOR of a stream processing application.** Kafka Streams **uses** transactions to implement its exactly-once guarantees. **Other stream processing frameworks, such as Spark Streaming or Flink, use DIFFERENT MECHANISMS** to provide their users with exactly-once semantics."*
**The scope, stated up front:**
> *"transactions in Kafka were **developed specifically for stream processing applications.** And therefore they were built to work with the **'consume-process-produce'** pattern... the processing of each input record will be considered complete **after the application's internal state has been updated AND the results were successfully produced to output topics.**"*
#### 3.2 Use cases [#32-use-cases]
> *"Transactions are useful for any stream processing application where **accuracy is important, and especially where stream processing includes AGGREGATION and/or JOINS.**"*
**Why filtering/mapping doesn't need them:**
> *"If the stream processing application only performs **single record transformation and filtering, there is no internal state to update**, and even if duplicates were introduced, **it is fairly straightforward to filter them out of the output stream.** When the stream processing application **aggregates several records into one**, it is much more difficult to check whether a result record is wrong because some input records were counted more than once; **it is impossible to correct the result without reprocessing the input.**"*
> *"**Financial applications** are typical examples... However, **because it is rather trivial to configure any Kafka Streams application to provide exactly-once guarantees, we've seen it enabled in more mundane use cases, including, for instance, CHATBOTS.**"*
#### 3.3 The two problems transactions solve [#33-the-two-problems-transactions-solve]
**Note that these need *different* solutions:**
* Problem 1 is an **atomicity** problem → solved by **atomic multipartition writes**.
* Problem 2 is an **identity/authority** problem → solved by **zombie fencing with epochs**.
Transactions provide both.
#### 3.4 The mechanism: atomic multipartition writes [#34-the-mechanism-atomic-multipartition-writes]
**The insight that makes it possible:**
> *"Exactly-once processing means that **consuming, processing, and producing are done ATOMICALLY.** Either the offset of the original message is committed **and** the result is successfully produced, **or neither of these things happen.**"*
>
> *"**committing offsets AND producing results BOTH involve writing messages to partitions.** However, the results are written to an output topic, and offsets are written to the **`__consumer_offsets`** topic. **If we can open a transaction, write both messages, and commit if both were written successfully — or abort to retry if they were not — we will get the exactly-once semantics we are after.**"*
#### The transactional producer [#the-transactional-producer]
> *"A transactional producer is simply a Kafka producer that is **configured with a `transactional.id`** and has been **initialized using `initTransactions()`**."*
| | `producer.id` | `transactional.id` |
| -------- | -------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| Origin | *"generated automatically by Kafka brokers"* | *"part of the **producer configuration**"* |
| Lifetime | New on every init | *"expected to **persist between restarts**"* |
| Purpose | dedup within one producer's life | **"the main role of the `transactional.id` is to IDENTIFY THE SAME PRODUCER ACROSS RESTARTS"** |
> *"Kafka brokers maintain **`transactional.id` → `producer.id` mapping**, so if `initTransactions()` is called again with an existing `transactional.id`, **the producer will also be assigned the SAME `producer.id` instead of a new random number.**"*
**This is exactly the hole from §2.5 Case A, now plugged.** A stable `transactional.id` yields a stable `producer.id`, which makes cross-restart zombie detection possible.
#### Zombie fencing via epochs [#zombie-fencing-via-epochs]
> *"The usual way of fencing zombies — **using an epoch** — is used here. Kafka **increments the epoch number associated with a `transactional.id` when `initTransaction()` is invoked.** **Send, commit, and abort requests from producers with the same `transactional.id` but LOWER EPOCHS will be rejected with the `FencedProducer` error. The older producer will not be able to write to the output stream and will be forced to `close()`, preventing the zombie from introducing duplicate records.**"*
> *"In **Apache Kafka 2.5 and later**, there is also an option to add **consumer group metadata** to the transaction metadata. This metadata will also be used for fencing, **which will allow producers with DIFFERENT transactional IDs to write to the same partitions while still fencing against zombie instances.**"* (→ §3.7)
#### 3.5 The consumer side — `isolation.level` [#35-the-consumer-side--isolationlevel]
> *"Transactions are a **producer feature for the most part**... **However, this isn't quite enough — records written transactionally, EVEN ONES THAT ARE PART OF TRANSACTIONS THAT WERE EVENTUALLY ABORTED, ARE WRITTEN TO PARTITIONS JUST LIKE ANY OTHER RECORDS.** Consumers need to be configured with the right isolation guarantees, **otherwise we won't have the exactly-once guarantees we expected.**"*
**⚠️ Aborted records are physically in the log.** Transactions don't prevent the write; they mark it. Filtering is the *consumer's* job, and it's off by default.
##### What `read_committed` does NOT give you [#what-read_committed-does-not-give-you]
> *"Configuring `read_committed` mode **does NOT guarantee that the application will get ALL messages that are part of a specific transaction.** It is possible to **subscribe to only a SUBSET of topics** that were part of the transaction and therefore get a subset of the messages. **In addition, the application CAN'T KNOW WHEN TRANSACTIONS BEGIN OR END, or WHICH MESSAGES ARE PART OF WHICH TRANSACTION.**"*
This limitation is the root cause of most of §3.6's failure cases. Consumers see *committed* records, but have **no transaction boundary information whatsoever.**
##### The Last Stable Offset (LSO) — and the latency it costs [#the-last-stable-offset-lso--and-the-latency-it-costs]
> *"To guarantee that messages will be read in order, `read_committed` mode **will not return messages that were produced AFTER the point when the FIRST STILL-OPEN TRANSACTION BEGAN (known as the Last Stable Offset, or LSO).** Those messages will be **withheld** until that transaction is **committed or aborted by the producer**, or until they reach **`transaction.timeout.ms` (default of 15 minutes)** and are **aborted by the broker.**"*
>
> **"Holding a transaction open for a long duration will introduce higher end-to-end latency by delaying consumers."**
#### The guarantee you actually get [#the-guarantee-you-actually-get]
> *"Our simple stream processing job will have exactly-once guarantees on its output **even if the input was written NONTRANSACTIONALLY.** The atomic multipartition produce guarantees that **if the output records were committed to the output topic, the offset of the input records was ALSO committed for that consumer, and as a result the input records will NOT BE PROCESSED AGAIN.**"*
***
# 2.9 What actually breaks in production — Ch. 2 consolidated (/docs/kafka/installing-kafka/actually-breaks-production-ch)
| # | Failure | Root cause | Fix / detection |
| -- | -------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------------------------------------------------- |
| 1 | **Broker won't start, logs an ID error** | Duplicate `broker.id` | Unique IDs derived from hostname |
| 2 | **One disk fills while others are empty** | `log.dirs` placement uses **fewest partitions, not least space** | Monitor per-mount usage, not just total; rebalance manually; prefer equal-size disks |
| 3 | **Broker restart takes hours after a crash** | `num.recovery.threads.per.data.dir` = 1 default; unclean shutdown must **check and truncate** every segment | Raise it (remember: × number of log dirs) |
| 4 | **Topics appear from nowhere; typo'd topic names in prod** | `auto.create.topics.enable=true`, and **metadata requests alone create topics** | Set `false`; provision topics explicitly |
| 5 | **All leadership piles onto one broker after a restart** | No automatic preferred-leader rebalance | `auto.leader.rebalance.enable=true`; tune imbalance thresholds |
| 6 | **"1 week retention" actually keeps 17 days** | Low-volume topic never fills a 1 GB segment; **retention only applies to closed segments** | Lower `log.segment.bytes` and/or set `log.roll.ms` for low-volume topics |
| 7 | **Disk usage jumps after a partition reassignment and never falls** | Time retention reads **mtime**; partition moves **reset mtime** → excess retention | Expect it; verify after rebalances; consider size-based retention for moved partitions |
| 8 | **Data deleted earlier than expected** | Both `log.retention.bytes` and `log.retention.ms` set — **either** triggers deletion | Pick one policy; the book explicitly recommends this "to prevent surprises and unwanted data loss" |
| 9 | **Retention silently doubles after adding partitions** | `log.retention.bytes` is **per partition** | Recompute total retention whenever partition count changes |
| 10 | **I/O latency spikes at regular intervals** | Time-based segment roll: the clock **starts at broker start**, so **all low-volume partitions roll simultaneously** | Stagger, or use size-based rolling where possible |
| 11 | **Consumer permanently stuck on one partition** | `fetch.message.max.bytes` \< broker `message.max.bytes` | Raise consumer fetch size **and** `replica.fetch.max.bytes` before raising broker max |
| 12 | **Replication stalls on one partition** | `replica.fetch.max.bytes` \< `message.max.bytes` | same as above |
| 13 | **Producer got an ack, message is gone** | `min.insync.replicas=1` (default) + leader failed before replicating; leadership moved to a replica lacking the write | `min.insync.replicas=2` **with** producer `acks=all`; RF ≥ 3 (Ch. 7) |
| 14 | **A rolling upgrade causes an outage** | RF only 1 above `min.insync.replicas` → planned outage consumes the entire redundancy budget, then a disk dies | **RF++** (2 above `min.insync.replicas`) |
| 15 | **Consumer performance degrades after colocating another service** | Page cache contention — Kafka's read path *is* the page cache | Don't colocate significant applications with brokers |
| 16 | **Replication falls behind under peak load; under-replicated partitions climb** | **Saturated NIC** — outbound = consumers × inbound + replication + mirroring | 10 Gb NICs minimum; count replication as a consumer when sizing |
| 17 | **Whole-broker performance collapse, everything slow** | **Swapping** — pages evicted, page cache starved | `vm.swappiness=1` (**not 0** — semantics changed in kernel 3.5-rc1) |
| 18 | **Periodic multi-second I/O stalls** | `vm.dirty_ratio` too high → **forced synchronous flushes** | Tune with `/proc/vmstat` under real load; require replication if running high |
| 19 | **Kernel flushes constantly, throughput poor** | `vm.dirty_background_ratio` set to 0 | Use \~5; never 0 |
| 20 | **"Too many open files"** | FD/map limits below `partitions × (partition_size/segment_size) + connections` | `vm.max_map_count` 400k–600k; raise ulimits |
| 21 | **Filesystem corruption / data loss after a host crash** | Ext4 **delayed allocation** + unsafe tuning (long commit interval) | Prefer **XFS**; if Ext4, understand you accepted this risk |
| 22 | **Unexplained write amplification** | `atime` updates on every read | mount `noatime` (safe — Kafka doesn't use atime; `mtime` is preserved) |
| 23 | **Connections dropped under burst / poor large-transfer throughput** | Default kernel socket + backlog sizes | `net.core.*mem_*`, `net.ipv4.tcp_*mem`, `tcp_window_scaling=1`, raise `tcp_max_syn_backlog` and `netdev_max_backlog` |
| 24 | **Long GC pauses → broker drops out of the ISR / ZK session expires** | CMS default (pre-G1GC era compatibility choice) | G1GC with `MaxGCPauseMillis=20`, `InitiatingHeapOccupancyPercent=35`; fixed heap `-Xms == -Xmx` |
| 25 | **A rack loses power and a partition goes offline despite RF 3** | Rack awareness applies to **newly created partitions only**; **reassignments silently break it and nothing monitors it** | Cruise Control or equivalent; audit replica placement after every reassignment |
| 26 | **You lose all broker data when a cloud VM moves** | Ephemeral disks | Azure **Managed Disks**; AWS: understand local-SSD vs EBS tradeoff |
| 27 | **Several brokers go offline simultaneously; controller acts strangely for hours afterward** | Shared ZooKeeper ensemble with a noisy application → ZK latency/timeout → brokers lose ZK together → offline partitions + controller stress + **subtle errors long afterward** | Dedicated ensemble for Kafka; chroot per cluster; never colocate a busy app |
| 28 | **Cluster hits a wall at high partition counts; produce/consume/controller queues back up** | Exceeded replica-per-broker limits | ≤ **14,000 replicas/broker**, ≤ **1,000,000 replicas/cluster** (replicas = partitions × RF) |
| 29 | **`/tmp/kafka-logs` disappears on reboot** | The install example's default path | Always set `log.dirs` to real, persistent, monitored mounts |
| 30 | **ZooKeeper maintenance causes an outage** | 3-node ensemble tolerates 1 failure; a rolling node swap uses it all | 5-node ensemble; never exceed 7; add **observers** for read-only load |
***
# 2.4 Configuring the broker — every knob, and its tradeoff (/docs/kafka/installing-kafka/configuring-broker-every-knob)
#### 4.1 General broker parameters [#41-general-broker-parameters]
##### `broker.id` [#brokerid]
Integer identifier, default `0`, **must be unique per broker within a cluster**.
The selection is *technically* arbitrary and **can be moved between brokers for maintenance**. But:
> **Set it to something intrinsic to the host** so that during maintenance it is not onerous to map broker IDs to hosts. If hostnames are `host1.example.com`, `host2.example.com`, use `1` and `2`.
*Why this matters at 3am:* every metric, log line, and `--describe` output refers to brokers by ID. If ID↔host requires a lookup table, you're slower during an incident.
##### `listeners` [#listeners]
Comma-separated list of URIs with listener names. Replaces the deprecated simple `port` config.
```properties
listeners=PLAINTEXT://localhost:9092,SSL://:9091
```
Format: `://:`
| Value | Effect |
| -------------------------- | --------------------------------------------------------------------- |
| hostname `0.0.0.0` | bind **all interfaces** |
| hostname **empty** | bind the **default interface** |
| non-standard listener name | you **must also** configure `listener.security.protocol.map` |
| port **\< 1024** | **Kafka must be started as root** — *not a recommended configuration* |
##### `zookeeper.connect` [#zookeeperconnect]
Format: **semicolon**-separated list of `hostname:port/path`.
```properties
zookeeper.connect=zoo1:2181;zoo2:2181;zoo3:2181/kafka-cluster-a
```
* `hostname` / `port` — ZooKeeper server and its client port.
* `/path` — **optional chroot** to use as the root for this Kafka cluster. Omitted → root path. **If the chroot path doesn't exist, the broker creates it at startup.**
> ### Why use a chroot path? [#why-use-a-chroot-path]
>
> * It lets the **ZooKeeper ensemble be shared with other applications, including other Kafka clusters, without conflict.**
> * Also: **specify multiple ZooKeeper servers** (all in the same ensemble) so the broker can connect to another member **on server failure**.
*Practical consequence:* if you ever want a second Kafka cluster (and Ch. 1 says you eventually will — for data segregation, security isolation, or DR), a chroot from day one means you don't have to migrate ZooKeeper state later.
##### `log.dirs` (prefer over `log.dir`) [#logdirs-prefer-over-logdir]
Where log segments live. `log.dirs` is a **comma-separated list of local paths**; if unset it falls back to `log.dir`.
**Placement algorithm — and its flaw:**
> The broker stores partitions across directories in a **"least-used" fashion**, with **one partition's log segments stored within the same path**.
> **The broker places a new partition in the path that has the *least number of partitions* currently stored in it — NOT the least amount of disk space used.** So **an even distribution of data across multiple directories is not guaranteed.**
##### `num.recovery.threads.per.data.dir` [#numrecoverythreadsperdatadir]
A configurable thread pool used in exactly three situations:
1. **Normal startup** — open each partition's log segments
2. **Startup after a failure** — **check and truncate** each partition's log segments
3. **Shutdown** — cleanly close log segments
Default: **one thread per log directory.**
> Since these threads are **only used during startup and shutdown**, it is reasonable to set a larger number to parallelize. **Specifically, when recovering from an unclean shutdown, this can mean the difference of several hours when restarting a broker with a large number of partitions!**
**Multiplication trap:** the value is **per log directory**. `num.recovery.threads.per.data.dir=8` with 3 paths in `log.dirs` = **24 threads total**.
##### `auto.create.topics.enable` [#autocreatetopicsenable]
Default: broker auto-creates a topic when:
1. A producer starts writing to it
2. A consumer starts reading from it
3. **Any client requests metadata for it**
> This can be undesirable, **especially as there is no way to validate the existence of a topic through the Kafka protocol without causing it to be created.**
That's a genuinely nasty property: *checking* creates. Set to `false` if you manage topic creation explicitly (manually or via provisioning).
##### `auto.leader.rebalance.enable` [#autoleaderrebalanceenable]
Prevents the cluster becoming unbalanced with all topic leadership on one broker.
##### `delete.topic.enable` [#deletetopicenable]
Set `false` to **prevent arbitrary topic deletion** — for environments with data retention guidelines or where an accidental `--delete` would be catastrophic.
***
#### 4.2 Topic defaults [#42-topic-defaults]
These are **cluster-wide defaults for newly created topics**. Set them to *"baseline values appropriate for the majority of the topics in the cluster."* Per-topic overrides go through the admin tools (Ch. 12).
> **Removed in newer versions:** `log.retention.hours.per.topic`, `log.retention.bytes.per.topic`, `log.segment.bytes.per.topic` — these old broker-side per-topic overrides are **no longer supported**. Use admin tools.
##### `num.partitions` [#numpartitions]
Partitions for a new topic, **primarily when auto-creation is enabled**. Default: **1**.
> **Partition count can only be increased, never decreased.** If a topic needs *fewer* partitions than `num.partitions`, you must create it manually.
Many users set partition count **equal to, or a multiple of, the number of brokers** so partitions distribute evenly → message load distributes evenly. A 10-partition topic on a 10-broker cluster with balanced leadership has optimal throughput. **Not a requirement** — you can balance load other ways, e.g. multiple topics.
###### How to choose the number of partitions — the full checklist [#how-to-choose-the-number-of-partitions--the-full-checklist]
1. **What throughput do you expect for the topic?** (100 KBps? 1 GBps?)
2. **What is your maximum throughput consuming from a *single* partition?** *A partition is always consumed completely by a single consumer* — even without consumer groups, the consumer must read all messages in the partition. **If your slow consumer writes to a database that never handles more than 50 MBps per writing thread, you are limited to 50 MBps per partition.**
3. You can do the same exercise for **per-producer per-partition throughput**, but **producers are typically much faster than consumers, so it's usually safe to skip.**
4. **If you send messages to partitions based on keys, adding partitions later can be very challenging** → **calculate throughput on expected *future* usage, not current.**
5. Consider **partitions per broker**, and **available disk space and network bandwidth per broker**.
6. **Avoid overestimating** — each partition uses **memory and other resources** on the broker and **increases the time for metadata updates and leadership transfers**.
7. **Are you mirroring data?** Factor in mirroring throughput. **Large partitions can become a bottleneck in many mirroring configurations.**
8. **Cloud IOPS limits.** There may be hard IOPS caps per VM/disk. **Too many partitions increases IOPS due to the parallelism involved** → you hit quotas.
**The formula:**
**Heuristic when you have no numbers:**
> Limit the size of the partition on disk to **less than 6 GB per day of retention**. *Starting small and expanding as needed is easier than starting too large.*
**Summary of the tension:** *"you want many partitions, but not too many."*
##### `default.replication.factor` [#defaultreplicationfactor]
Replication factor for auto-created topics.
**The recommendation, and the reasoning — this is one of the best passages in the chapter:**
> Set replication factor to **at least 1 above `min.insync.replicas`**.
> For more fault resistance, if you have large enough clusters and enough hardware, set it **2 above `min.insync.replicas`** — abbreviated **RF++**.
> **RF++ allows easier maintenance and prevents outages.**
**Why RF++:** *"to allow for one **planned** outage within the replica set and one **unplanned** outage to occur simultaneously."*
For a typical cluster this means **a minimum of three replicas of every partition.** The scenario to hold in your head: *a network switch outage, a disk failure, or some other unplanned problem **during a rolling deployment or upgrade** of Kafka or the OS.* Rolling upgrades are exactly when your redundancy budget is already partly spent.
##### Retention: `log.retention.ms` / `.minutes` / `.hours` [#retention-logretentionms--minutes--hours]
Default in the config file is `log.retention.hours=168` (**one week**).
**All three control the same thing; the smaller unit takes precedence if more than one is specified.**
> **Recommended: use `log.retention.ms`** — because the smaller unit wins, setting `ms` guarantees the value you set is the one used.
> ### Retention by time uses **mtime**, and that's a trap [#retention-by-time-uses-mtime-and-thats-a-trap]
>
> Time-based retention examines the **last modified time (mtime) of each log segment file on disk.** Under normal operations that's when the segment was closed — i.e. the timestamp of the last message in the file.
> **However, when using administrative tools to move partitions between brokers, this time is NOT accurate and will result in excess retention for these partitions.** (Ch. 12 covers partition moves.)
*Why:* a partition move copies files, which resets mtime to "now" — so a 6-day-old segment looks brand new and survives another full retention period. Disk usage jumps after a rebalance and nobody knows why.
##### `log.retention.bytes` [#logretentionbytes]
Total bytes retained — **applied per partition**.
> **All retention is performed for individual partitions, not the topic.** So **if you expand the partition count of a topic, retention also increases** when using `log.retention.bytes`.
> `-1` = **infinite retention**.
> ### Configuring retention by size AND time [#configuring-retention-by-size-and-time]
>
> If both `log.retention.bytes` and a time parameter are set, **messages may be removed when *either* criterion is met.**
>
> * `retention.ms = 1 day`, `retention.bytes = 1 GB`: messages **less than 1 day old can be deleted** if the day's volume exceeds 1 GB.
> * Conversely, if volume is under 1 GB, messages are **deleted after 1 day anyway**.
>
> **Recommendation: for simplicity choose either size- or time-based retention — not both — to prevent surprises and unwanted data loss.** Both can be used for advanced configurations.
##### `log.segment.bytes` [#logsegmentbytes]
**Retention operates on log *segments*, not individual messages.** Messages append to the **current** (active) segment. When it reaches `log.segment.bytes` (**default 1 GB**), it is **closed** and a new one opened. **Only a closed segment can be considered for expiration.**
**Smaller segments** → files closed and allocated more often → **reduces overall efficiency of disk writes**.
###### The low-volume-topic retention bug — work through this one [#the-low-volume-topic-retention-bug--work-through-this-one]
> A topic receiving only **100 MB/day** with default `log.segment.bytes` (1 GB) takes **10 days to fill one segment.** Messages cannot expire until the segment is closed. If `log.retention.ms` = 1 week, there will actually be **up to 17 days of messages retained**: once the segment closes with 10 days of messages inside, that segment must be retained **7 more days** before it expires — because **the segment can't be removed until its *last* message can be expired.**
**Fix:** for low-volume topics, lower `log.segment.bytes` **or** set `log.roll.ms`. Otherwise your "1 week retention" is a lie and your GDPR/compliance deletion window is wrong.
> ### Segment size affects offset-by-timestamp lookups [#segment-size-affects-offset-by-timestamp-lookups]
>
> When you request the offset for a partition **at a specific timestamp**, Kafka finds the log segment that **was being written at that time** — using the file's creation and last-modified times, looking for a file **created before** the timestamp and **last modified after** it. **The offset at the beginning of that segment (which is also the filename) is returned.**
>
> **Consequence: coarser segments = coarser timestamp seeks.** With 1 GB segments on a slow topic, "seek to 10:15am" can land you days earlier. This directly affects `--from-timestamp` resets and time-based reprocessing.
##### `log.roll.ms` [#logrollms]
Time after which a segment should be closed. **No default** — so **by default segments close by size only.**
`log.segment.bytes` and `log.roll.ms` are **not mutually exclusive**: Kafka closes a segment **when either the size limit or the time limit is reached, whichever comes first.**
> ### Disk performance with time-based segments — the thundering herd [#disk-performance-with-time-based-segments--the-thundering-herd]
>
> When using a time-based segment limit, consider the impact of **many log segments being closed simultaneously.** This happens when **many partitions never reach the size limit**: the time-limit clock **starts when the broker starts**, so it **always fires at the same moment for all those low-volume partitions.**
##### `min.insync.replicas` [#mininsyncreplicas]
> Setting `min.insync.replicas = 2` ensures **at least two replicas are caught up and "in sync" with the producer.** Used **in tandem with the producer config `acks=all`.** This ensures **at least two replicas (leader + one other) acknowledge a write** for it to be successful.
**The failure it prevents:**
> **Trade-off, stated plainly:** configuring for higher durability **is less efficient due to the extra overhead**. Clusters with **high throughput that can tolerate occasional message loss aren't recommended to change this from the default of 1.**
That's an honest statement rarely made in docs: `min.insync.replicas=1` is a *deliberate* data-loss-tolerant setting, and it's the default. (Full treatment in Ch. 7.)
##### `message.max.bytes` [#messagemaxbytes]
Max size of a producible message. **Default 1000000 (≈1 MB).** A producer exceeding it **gets an error back and the message is not accepted.**
> As with all byte sizes on the broker, this deals with **compressed** message size — producers can send messages **much larger uncompressed**, provided they compress to under the limit.
**Performance impacts of raising it:**
* Broker threads handling network connections/requests **work longer on each request**
* **Larger disk writes** → impacts I/O throughput
* Alternatives mentioned: **blob stores and/or tiered storage** (not covered in the chapter)
> ### Coordinating message size across three configs — the stuck-consumer bug [#coordinating-message-size-across-three-configs--the-stuck-consumer-bug]
>
> ```
> broker : message.max.bytes
> consumer : fetch.message.max.bytes ← must be ≥ broker's
> brokers : replica.fetch.max.bytes ← must be ≥ broker's (for replication)
> ```
>
> If `fetch.message.max.bytes` \< `message.max.bytes`, consumers hitting a larger message **fail to fetch it**, resulting in **a consumer that gets stuck and cannot proceed.** Same rule for `replica.fetch.max.bytes` on brokers — otherwise **replication stalls** on that partition.
This is a "raise one number, break the cluster silently" trap. Raise all three together, in the right order (consumers/replicas first, broker last).
***
# 2.6 Configuring Kafka clusters (/docs/kafka/installing-kafka/configuring-kafka-clusters)
#### 6.1 How many brokers? [#61-how-many-brokers]
Cluster size is bound by **four** things:
##### (1) Disk capacity [#1-disk-capacity]
> Increasing the replication factor increases storage requirements **by at least 100%**, depending on the factor chosen. ("Replicas" = the number of *different brokers* a single partition is copied to.)
##### (2) Replica capacity per broker — the numbers you should memorize [#2-replica-capacity-per-broker--the-numbers-you-should-memorize]
| | Old official recommendation | **Current recommendation** (well-configured env) |
| ------------------------ | --------------------------- | ------------------------------------------------ |
| Replicas **per broker** | no more than **4,000** | **no more than 14,000** |
| Replicas **per cluster** | no more than **200,000** | **no more than 1,000,000** |
*"Advances in cluster efficiency have allowed Kafka to scale much larger."* Note these are **replicas**, not partitions — `partitions × RF`. People routinely mis-plan by a factor of 3 here.
##### (3) CPU capacity [#3-cpu-capacity]
Usually not a major bottleneck, **but it can be with an excessive amount of client connections and requests on a broker.** Watch **overall CPU usage relative to how many unique clients and consumer groups there are**, and expand to meet those needs.
##### (4) Network capacity — the worked example [#4-network-capacity--the-worked-example]
> If the network interface on a single broker is used to **80% capacity at peak**, and there are **two consumers** of that data, **the consumers will not be able to keep up with peak traffic unless there are two brokers.**
> **If replication is being used, that is an additional consumer of the data that must be accounted for.**
Also account for **traffic not being consistent over the retention period** (bursts at peak times). And you may scale out to more brokers to address **lesser disk throughput or available system memory**.
#### 6.2 Broker configuration [#62-broker-configuration]
Only two requirements to join a cluster (see §0): identical `zookeeper.connect`, unique `broker.id`. Other cluster-relevant params — **specifically those controlling replication** — are covered in later chapters.
***
# 2.10 Deploy / monitor / scale / back up — Chapter 2's contribution (/docs/kafka/installing-kafka/deploy-monitor-scale-back)
#### Deploy checklist (production) [#deploy-checklist-production]
#### Monitoring implied by this chapter [#monitoring-implied-by-this-chapter]
| Signal | Why Ch. 2 makes it matter |
| -------------------------------------------- | -------------------------------------------------------- |
| **Per-log-dir disk usage** (not aggregate) | placement is by partition count, not bytes |
| **Under-replicated partitions** | the symptom of NIC saturation and of GC/ZK trouble |
| **Offline partitions** | what a ZK interruption produces |
| **Active controller count (== 1)** | ZK stress destabilizes the controller |
| **GC pause time** | long pauses → ISR drops → ZK session expiry |
| **Dirty pages** (`/proc/vmstat`) | to tune `vm.dirty_*` empirically under load |
| **Swap usage (should be \~0)** | swapping starves the page cache |
| **Open file descriptors** | segments + connections grow with partitions |
| **NIC utilization** | outbound = consumers × inbound + replication + mirroring |
| **Replicas per broker / per cluster** | hard ceilings: 14k / 1M |
| **ZooKeeper latency + outstanding requests** | Kafka is *sensitive* to ZK latency |
| **Rack-awareness audit after reassignments** | Kafka will not tell you it broke |
#### Scaling levers from this chapter [#scaling-levers-from-this-chapter]
#### "Backup" — what Ch. 2 actually gives you [#backup--what-ch-2-actually-gives-you]
Chapter 2 has no backup section, and that absence is informative. What it *does* give you:
1. **Replication factor** — intra-cluster redundancy. **Not a backup**: it won't protect you from a bad `--delete`, which is exactly why `delete.topic.enable=false` exists as a config.
2. **Retention** — a bounded replay window, and Ch. 2 shows how easily you can *miscalculate* it (segment rolling, mtime resets, dual policies). Your "backup window" is only as accurate as your understanding of segment mechanics.
3. **`delete.topic.enable=false`** — the closest thing to a "protect me from myself" control.
4. **Managed disks over ephemeral** — cloud-specific durability floor. *"If a VM is moved, you run the risk of losing all the data on your Kafka broker."*
5. **ZooKeeper's `dataDir`** — genuinely worth backing up separately; it holds cluster metadata, and ZK has its own snapshot/txn-log durability story.
Real archival remains a **sink** concern (Connect → S3/HDFS, Ch. 9) and DR remains a **cross-cluster** concern (MirrorMaker, Ch. 10).
***
# 2.2 Environment setup (/docs/kafka/installing-kafka/environment-setup)
#### 2.1 Operating system [#21-operating-system]
Kafka is a **Java application** and runs on Windows, macOS, Linux, and others. **Linux is the recommendation for the general use case.** (Windows/macOS install → Appendix A.)
#### 2.2 Java [#22-java]
* Needed **before** ZooKeeper or Kafka.
* Works with **all OpenJDK-based implementations, including Oracle JDK**.
* Latest Kafka versions support **Java 8 and Java 11**.
* Runtime (JRE) is enough to *run*, but **the full JDK is recommended when developing tools and applications**.
* **Install the latest released patch version** — older versions may carry security vulnerabilities.
* Book's examples assume `/usr/java/jdk-11.0.10`.
#### 2.3 ZooKeeper — what it's for and why it's shaped that way [#23-zookeeper--what-its-for-and-why-its-shaped-that-way]
**What Kafka stores in ZooKeeper:** metadata about the **Kafka cluster** (brokers, topics, partitions), and (historically) **consumer client details**.
ZooKeeper is *"a centralized service for maintaining configuration information, naming, providing distributed synchronization, and providing group services."*
Tested extensively with the **stable 3.5** release; the book uses **3.5.9**.
> **Note:** you *can* run a ZooKeeper server from scripts inside the Kafka distribution, but *"it is trivial to install a full version of ZooKeeper from the distribution"* — do that instead for anything real.
##### Standalone server (dev only) [#standalone-server-dev-only]
```bash
tar -zxf apache-zookeeper-3.5.9-bin.tar.gz
mv apache-zookeeper-3.5.9-bin /usr/local/zookeeper
mkdir -p /var/lib/zookeeper
cat > /usr/local/zookeeper/conf/zoo.cfg << EOF
tickTime=2000
dataDir=/var/lib/zookeeper
clientPort=2181
EOF
export JAVA_HOME=/usr/java/jdk-11.0.10
/usr/local/zookeeper/bin/zkServer.sh start
```
**Verify** with a four-letter command over the client port:
```bash
telnet localhost 2181
srvr
# Zookeeper version: 3.5.9-...
# Latency min/avg/max: 0/0/0
# Received: 1
# Sent: 0
# Connections: 1
# Outstanding: 0
# Zxid: 0x0
# Mode: standalone ← this is what you're checking
# Node count: 5
```
##### Ensemble (production) [#ensemble-production]
ZooKeeper is designed to run as a cluster — an **ensemble** — for high availability.
**Why an odd number of servers?** Because of the balancing/consensus algorithm, **a majority of members (a quorum) must be working for ZooKeeper to respond to requests.**
**Sizing guidance — and the reasoning matters more than the number:**
> **Consider a five-node ensemble.** To make configuration changes to the ensemble, including swapping a node, **you reload nodes one at a time.** If your ensemble cannot tolerate more than one node being down, **doing maintenance work introduces additional risk.**
That is the whole argument for 5 over 3: a 3-node ensemble tolerates 1 failure, and *routine maintenance consumes that entire budget* — so a hardware failure during a rolling restart takes you down.
> **Do not run more than seven nodes** — performance degrades due to the nature of the consensus protocol.
> If 5–7 nodes can't support the load due to too many client connections, **add observer nodes** to balance read-only traffic (observers participate in reads but not in the quorum vote).
**Ensemble configuration** — shared config on all nodes:
```properties
tickTime=2000
dataDir=/var/lib/zookeeper
clientPort=2181
initLimit=20
syncLimit=5
server.1=zoo1.example.com:2888:3888
server.2=zoo2.example.com:2888:3888
server.3=zoo3.example.com:2888:3888
```
| Param | Meaning |
| ----------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `tickTime` | The base time unit (ms). Everything else is a multiple of this. |
| `initLimit` | How long to allow **followers to connect with a leader**. `20 × 2000ms = 40s` |
| `syncLimit` | How long **out-of-sync followers** may lag the leader. `5 × 2000ms = 10s` |
| `server.X=host:peerPort:leaderPort` | `X` = server ID (**integer; need not be zero-based or sequential**), `peerPort` (2888) = **inter-ensemble communication**, `leaderPort` (3888) = **leader election** |
**Plus, per server:** a file named **`myid`** in `dataDir` containing that server's ID number, **matching the config**.
**Port reachability rules — a classic firewall bug:**
> **Single-machine ensemble for testing:** set all hostnames to `localhost`, give each instance unique `peerPort`/`leaderPort`, and a **separate `zoo.cfg` per instance** with unique `dataDir` and `clientPort`. **Testing purposes only — not recommended for production.** (It shares a single failure domain, which defeats the point.)
***
# 2. Installing Kafka (/docs/kafka/installing-kafka)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 2*
> **Learning goal:** this is the *operations* chapter disguised as an install guide. The install is 3 commands; the value is in **why each config exists, what it trades away, and which defaults will hurt you in production.**
***
# 2.3 Installing a Kafka broker (/docs/kafka/installing-kafka/installing-kafka-broker)
```bash
tar -zxf kafka_2.13-2.7.0.tgz
mv kafka_2.13-2.7.0 /usr/local/kafka
mkdir /tmp/kafka-logs # ← never do this in production
export JAVA_HOME=/usr/java/jdk-11.0.10
/usr/local/kafka/bin/kafka-server-start.sh -daemon \
/usr/local/kafka/config/server.properties
```
*(At press time the current release was 2.8.0 on Scala 2.13.0; examples use 2.7.0.)*
**Smoke test — create, produce, consume:**
```bash
# create
kafka-topics.sh --bootstrap-server localhost:9092 --create \
--replication-factor 1 --partitions 1 --topic test
# describe
kafka-topics.sh --bootstrap-server localhost:9092 --describe --topic test
# Topic:test PartitionCount:1 ReplicationFactor:1 Configs:
# Topic: test Partition: 0 Leader: 0 Replicas: 0 Isr: 0
# produce (Ctrl-C to stop)
kafka-console-producer.sh --bootstrap-server localhost:9092 --topic test
> Test Message 1
> Test Message 2
# consume
kafka-console-consumer.sh --bootstrap-server localhost:9092 --topic test --from-beginning
# Test Message 1
# Test Message 2
```
> ### `--zookeeper` is deprecated [#--zookeeper-is-deprecated]
>
> Older tooling used a `--zookeeper` connection string. **Deprecated in almost all cases.** Current best practice: **`--bootstrap-server`**, connecting directly to a broker. In a cluster you can give the `host:port` of **any** broker.
***
# 2.7 OS tuning (/docs/kafka/installing-kafka/os-tuning)
> Most Linux distributions ship kernel-tuning defaults that work fairly well for most applications, but a few changes **improve performance for a Kafka broker**. These revolve around **virtual memory**, **networking**, and **the disk mount point used for log segments**. Typically configured in **`/etc/sysctl.conf`** (check your distro docs).
#### 7.1 Virtual memory [#71-virtual-memory]
##### Swap — avoid at (almost) all costs [#swap--avoid-at-almost-all-costs]
**Two reasons:**
1. The cost of swapped-out pages *"will show up as a noticeable impact on all aspects of performance in Kafka."*
2. **Kafka makes heavy use of the page cache** — *"if the VM system is swapping to disk, there is not enough memory being allocated to page cache."*
**Should you disable swap entirely?** You *can* — swap isn't a requirement — but:
> It **does provide a safety net if something catastrophic happens.** Having swap can **prevent the OS from abruptly killing a process due to an out-of-memory condition.**
**Therefore:**
```properties
vm.swappiness = 1
```
> The parameter is **a percentage of how likely the VM subsystem is to use swap space rather than dropping pages from the page cache.** *It is preferable to reduce the memory available for page cache rather than utilize any amount of swap memory.*
> ### Why not swappiness = 0? — the changed-semantics trap [#why-not-swappiness--0--the-changed-semantics-trap]
>
> The old recommendation was `0`, which used to mean *"do not swap unless there is an out-of-memory condition."*
> **The meaning changed as of Linux kernel 3.5-rc1**, backported widely (**RHEL kernels as of 2.6.32-303**). `0` now means **"never swap under any circumstances"** — which removes the OOM safety net.
> **Hence `1` is now the recommendation.**
This is a great example of a config whose *semantics* changed underneath a widely-copied best practice. If you inherited a runbook that says `vm.swappiness=0`, it's wrong for the reason you think it's right.
##### Dirty pages [#dirty-pages]
Kafka relies on **disk I/O performance to give producers good response times** — which is also why log segments go on a fast disk (SSD, or a subsystem with significant **NVRAM for caching**, e.g. RAID).
```properties
vm.dirty_background_ratio = 5 # default 10 — % of total system memory
```
> Lower it below the default of 10. **5 is appropriate in many situations.**
> **Do NOT set it to zero** — that causes the kernel to **continually flush pages**, which **eliminates the kernel's ability to buffer disk writes against temporary spikes in the underlying device performance.**
```properties
vm.dirty_ratio = 60..80 # default 20 — % of total system memory
```
> This is the total dirty pages allowed **before the kernel forces synchronous operations to flush them.** Raise above the default of 20; **between 60 and 80 is reasonable.**
>
> **This introduces risk** in two ways: the **amount of unflushed disk activity** (data in memory, not on disk), and **the potential for long I/O pauses if synchronous flushes are forced.**
> **If a higher `vm.dirty_ratio` is chosen, it is highly recommended that replication be used in the cluster to guard against system failures.**
The tradeoff, stated plainly:
**Observe the actual behavior instead of guessing:**
```bash
cat /proc/vmstat | egrep "dirty|writeback"
# nr_dirty 21845
# nr_writeback 0
# nr_writeback_temp 0
# nr_dirty_threshold 32715981
# nr_dirty_background_threshold 2726331
```
> Review the number of dirty pages over time while the cluster is **under load**, in production or simulated.
##### File descriptors / memory maps [#file-descriptors--memory-maps]
Kafka uses file descriptors for **log segments and open connections**.
```txt
minimum needed ≈ (number_of_partitions) × (partition_size / segment_size)
+ number of connections the broker makes
```
```properties
vm.max_map_count = 400000 # or 600000 — based on the calculation above
vm.overcommit_memory = 0
```
> Set `vm.max_map_count` to **a very large number based on the above calculation**; 400,000 or 600,000 has generally been successful.
> `vm.overcommit_memory = 0` (the default) means **the kernel determines the amount of free memory from an application.** A non-zero value *"could lead the operating system to grab too much memory, depriving memory for Kafka to operate optimally. This is common for applications with high ingestion rates."*
#### 7.2 Disk / filesystem [#72-disk--filesystem]
Outside of hardware and RAID choice, **the filesystem has the next largest impact on performance.**
| | **XFS** (recommended) | Ext4 |
| ------------- | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Status | **Default FS for many Linux distros** | — |
| Performance | **Outperforms Ext4 for most workloads with minimal tuning** | Can perform well |
| Tuning needed | **Beyond the FS's own automatic tuning, none** | **Requires parameters considered less safe** |
| Specific risk | Also uses delayed allocation, but **generally safer than Ext4's** | Longer `commit` interval to force less frequent flushes; **delayed allocation of blocks → greater chance of data loss and filesystem corruption on system failure** |
| Batching | **More efficient when batching disk writes** → better overall I/O throughput | — |
**Mount options, regardless of filesystem:**
```txt
mount -o noatime,largeio ...
```
**`noatime` — why it's safe here specifically.** File metadata has three timestamps: **ctime** (creation), **mtime** (last modified), **atime** (last access). By default **atime is updated every time a file is read** → **a large number of disk writes** (writes generated *by reads* — the worst kind).
> The `atime` attribute is **generally of little use** unless an app needs to know if a file was accessed since last modified (use `relatime` then). **`atime` is not used by Kafka at all, so disabling it is safe.** `noatime` prevents these updates **without affecting proper handling of `ctime` and `mtime`.**
*(Important: Kafka **does** rely on `mtime` — that's how time-based retention works. `noatime` leaves `mtime` intact, which is why it's safe.)*
**`largeio`** — helps improve efficiency **when there are larger disk writes**.
#### 7.3 Networking [#73-networking]
> The kernel is **not tuned by default for large, high-speed data transfers.** The recommended changes for Kafka are **the same as those suggested for most web servers and other networking applications.**
**Socket buffers (all sockets):**
```properties
net.core.wmem_default = 131072 # 128 KiB
net.core.rmem_default = 131072 # 128 KiB
net.core.wmem_max = 2097152 # 2 MiB
net.core.rmem_max = 2097152 # 2 MiB
```
> *"Keep in mind that the maximum size does not indicate that every socket will have this much buffer space allocated; it only allows up to that much if needed."*
**TCP-specific buffers — set separately**, as three space-separated integers `min default max`:
```properties
net.ipv4.tcp_wmem = 4096 65536 2048000
net.ipv4.tcp_rmem = 4096 65536 2048000
# 4KiB 64KiB 2MiB
```
> **The maximum cannot be larger than `net.core.wmem_max` / `net.core.rmem_max`.** Based on actual workload, you may want to increase the maximums to allow greater buffering.
**Other useful parameters:**
| Parameter | Set to | Why |
| ------------------------------ | -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| `net.ipv4.tcp_window_scaling` | `1` | Clients **transfer data more efficiently**, and data can be **buffered on the broker side** |
| `net.ipv4.tcp_max_syn_backlog` | **> 1024** (default) | Allows **a greater number of simultaneous connections to be accepted** |
| `net.core.netdev_max_backlog` | **> 1000** (default) | Assists with **bursts of network traffic**, specifically at **multigigabit speeds**, by allowing more packets to be queued for the kernel to process |
***
# 2.1 What problem does this chapter solve? (/docs/kafka/installing-kafka/problem-does-chapter-solve)
The default `server.properties` shipped with Kafka is *"sufficient to run a standalone server as a proof of concept, but most likely will not be sufficient for large installations."*
That sentence is the chapter's thesis. The defaults are optimized for **"it starts on my laptop"**, not for **"it survives a rack losing power at 3am."** The gap between those two is what this chapter closes.
***
# 2.8 Production concerns (/docs/kafka/installing-kafka/production-concerns)
#### 8.1 Garbage collector — use G1GC [#81-garbage-collector--use-g1gc]
> Tuning Java GC *"has always been something of an art."* **Thankfully this changed with Java 7 and G1GC** (Garbage-First). Initially considered unstable, **it saw marked improvement in JDK8 and JDK11.** **G1GC is now recommended as the default collector for Kafka.**
**Why G1GC fits Kafka:**
* **Automatically adjusts to different workloads**
* Provides **consistent pause times over the application's lifetime**
* **Handles large heaps with ease by segmenting the heap into smaller zones and not collecting over the entire heap in each pause**
**Two knobs:**
| Option | Default | Meaning |
| -------------------------------- | ---------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `MaxGCPauseMillis` | **200 ms** | **Preferred** pause time per GC cycle — **not a fixed maximum; G1GC can and will exceed it if required.** G1GC schedules cycle frequency and number of zones collected so each cycle takes \~this long |
| `InitiatingHeapOccupancyPercent` | **45** | Percentage of **total heap** in use before G1GC **starts a collection cycle**. Includes **both new (Eden) and old zone usage in total** |
> The Kafka broker is **fairly efficient in how it uses heap and creates garbage objects**, so **it is possible to set these options lower.**
**Reference tuning — for a server with 64 GB memory running Kafka in a 5 GB heap:**
> **Kafka was originally released before G1GC was available and considered stable.** Therefore **Kafka defaults to concurrent mark and sweep (CMS)** for compatibility with all JVMs. **New best practice is to use G1GC for anything Java 1.8 and later.**
**Change it via environment variable:**
```bash
export KAFKA_JVM_PERFORMANCE_OPTS="-server -Xmx6g -Xms6g \
-XX:MetaspaceSize=96m -XX:+UseG1GC \
-XX:MaxGCPauseMillis=20 -XX:InitiatingHeapOccupancyPercent=35 \
-XX:G1HeapRegionSize=16M -XX:MinMetaspaceFreeRatio=50 \
-XX:MaxMetaspaceFreeRatio=80 -XX:+ExplicitGCInvokesConcurrent"
/usr/local/kafka/bin/kafka-server-start.sh -daemon \
/usr/local/kafka/config/server.properties
```
Note `-Xmx6g -Xms6g` — **min and max heap set equal**, the standard practice to avoid heap resizing pauses.
#### 8.2 Datacenter layout [#82-datacenter-layout]
**Why it matters:** in test/dev, broker physical location doesn't matter much. **In production, downtime means dollars lost** — through loss of service to users **or loss of telemetry on what users are doing.**
> **If not addressed prior to deploying Kafka, expensive maintenance to move servers around may be needed.** A datacenter environment **that has a concept of fault zones is preferable.**
##### `broker.rack` — rack awareness and its sharp edge [#brokerrack--rack-awareness-and-its-sharp-edge]
```properties
broker.rack=rack-a # or the cloud fault domain
```
Kafka **assigns new partitions in a rack-aware manner, ensuring replicas of a single partition do not share a rack.**
**Three critical limitations:**
> 1. **This only applies to partitions that are newly created.**
> 2. **The cluster does not monitor for partitions that are no longer rack aware** (for example, as a result of a partition reassignment).
> 3. **Nor does it automatically correct this situation.**
**Therefore:** use tools that keep the cluster balanced to maintain rack awareness — **Cruise Control** (Appendix B).
##### Physical best practice [#physical-best-practice]
> Best practice: **each Kafka broker in a different rack** — or at the very least, **not sharing single points of failure for infrastructure services such as power and network.**
Concretely:
* **Dual power connections to two different circuits**
* **Dual network switches**, with a **bonded interface on the servers** to fail over seamlessly
> **Even with dual connections, there is a benefit to having brokers in completely separate racks** — from time to time it may be necessary to take a rack or cabinet **offline** for physical maintenance (moving servers, rewiring power).
#### 8.3 Colocating applications on ZooKeeper [#83-colocating-applications-on-zookeeper]
**What Kafka actually writes to ZooKeeper:** metadata about brokers, topics, and partitions. **Writes only happen on** changes to consumer-group membership, or changes to the Kafka cluster itself.
> This traffic is **generally minimal, and it does not justify a dedicated ZooKeeper ensemble for a single Kafka cluster.** Many deployments **use a single ensemble for multiple Kafka clusters** (with a **chroot path per cluster**).
> ### Kafka, ZooKeeper, and the direction of travel [#kafka-zookeeper-and-the-direction-of-travel]
>
> **Dependency on ZooKeeper is shrinking.**
>
> * **Kafka 2.8.0 introduces an early-access, completely ZooKeeper-less Kafka — but it is NOT production ready.** (This is KRaft.)
> * **Before 0.9.0.0:** consumers (as well as brokers) used ZooKeeper directly to store consumer-group composition, which topics it consumed, and to **periodically commit offsets per partition** (to enable failover between group members).
> * **With 0.9.0.0:** the consumer interface changed, allowing this to be **managed directly with the Kafka brokers.**
> * **Each 2.x release removes ZooKeeper from more required paths.** Administration tools **now connect directly to the cluster**; connecting to ZooKeeper for topic creation, dynamic config changes, etc. is **deprecated**.
> * Command-line tools moved from `--zookeeper` to `--bootstrap-server`. **`--zookeeper` still works but is deprecated and will be removed.**
##### The consumer-offsets-in-ZooKeeper problem [#the-consumer-offsets-in-zookeeper-problem]
Even though deprecated, consumers have **a configurable choice** to commit offsets to **ZooKeeper or Kafka**, plus a configurable commit interval.
> These commits can be **a significant amount of ZooKeeper traffic**, especially with many consumers. **It may be necessary to use a longer commit interval if the ensemble can't handle the traffic. However, it is recommended that consumers using the latest Kafka libraries use Kafka for committing offsets, removing the ZooKeeper dependency.**
Note the tension worth internalizing: **commit interval is simultaneously a load knob and a correctness knob.** Longer interval = less traffic *and* more duplicate reprocessing on failure.
##### Do NOT share the ensemble with non-Kafka applications [#do-not-share-the-ensemble-with-non-kafka-applications]
> Outside of using one ensemble for multiple Kafka clusters, **it is not recommended to share the ensemble with other applications if it can be avoided.**
**The failure chain — memorize this one:**
> Other applications that can stress the ensemble, **either through heavy usage or improper operations, should be segregated to their own ensemble.**
The nastiest part is the last step: ZooKeeper hiccups leave **latent damage** that surfaces later during an unrelated operation. This is why "ZK looked fine when I checked" is not exculpatory evidence during an incident review.
***
# 2.5 Selecting hardware (/docs/kafka/installing-kafka/selecting-hardware)
> *"Selecting an appropriate hardware configuration for a Kafka broker can be more art than science."* Kafka has **no strict hardware requirement** and runs fine on most systems. Once performance matters, the bottlenecks are:
#### 5.1 Disk throughput → producer latency [#51-disk-throughput--producer-latency]
**The causal chain:**
**SSD vs HDD:**
| | SSD | HDD |
| ---------------------- | ---------------------------------------- | ------------------------------------------------------------------- |
| Seek/access time | **Drastically lower** → best performance | Higher |
| Cost per unit capacity | Expensive | **More economical, more capacity** |
| Improve by | — | **more of them per broker**: multiple data directories, or **RAID** |
Other factors: **specific drive technology** (SAS vs SATA) and **the quality of the drive controller**.
**The observed rule of thumb:**
> **HDD** → clusters with **very high storage needs but data not accessed often**.
> **SSD** → when there is a **very large number of client connections**.
#### 5.2 Disk capacity → retention [#52-disk-capacity--retention]
Example from the book: 1 TB/day traffic × 7 days retention = **minimum 7 TB usable** for log segments, **plus at least 10% overhead**, plus a growth buffer.
Total cluster traffic is balanced by **multiple partitions per topic**, letting **additional brokers augment capacity** when single-broker density is insufficient. The requirement is also shaped by the **replication strategy** (Ch. 7).
#### 5.3 Memory → consumer performance (this is the subtle one) [#53-memory--consumer-performance-this-is-the-subtle-one]
**The normal mode of operation:** a consumer reads **from the end of the partition**, caught up, lagging very little or not at all. In that state:
> The messages the consumer is reading are **optimally stored in the system's page cache**, resulting in **faster reads than if the broker had to reread from disk.** Therefore **more memory available for page cache improves consumer performance.**
**Heap requirements are tiny by comparison:**
> Kafka itself does not need much heap. **Even a broker handling 150,000 messages/second and 200 megabits/second can run with a 5 GB heap.** The rest of system memory is used by page cache and **benefits Kafka by caching log segments in use.**
**Hence the colocation rule:**
> This is **the main reason it is not recommended to colocate Kafka with any other significant application** — it will have to **share the page cache**, which **decreases consumer performance for Kafka.**
*Restated:* Kafka's performance model is "the OS is my cache." Any neighbor process that touches a lot of memory is stealing your read path.
#### 5.4 Networking → the throughput ceiling [#54-networking--the-throughput-ceiling]
> Available network throughput specifies the **maximum amount of traffic Kafka can handle.** Combined with disk storage, it's a governing factor for cluster sizing.
**The inherent asymmetry:**
**The dangerous consequence:**
> Should the network interface become saturated, **it is not uncommon for cluster replication to fall behind, which can leave the cluster in a vulnerable state.**
Read that as: *network saturation converts into a durability risk*, because under-replicated partitions mean you're one failure from data loss.
**Recommendation: at least 10 Gb NICs.** *"Older machines with 1 Gb NICs are easily saturated and aren't recommended."*
#### 5.5 CPU → only at extreme scale [#55-cpu--only-at-extreme-scale]
> Processing power is **not as important as disk and memory** until you scale very large.
**Where the CPU actually goes** — and it's a specific, non-obvious path:
Not the primary hardware factor **unless clusters become very large — hundreds of nodes and millions of partitions in a single cluster.** At that point better CPU can **reduce cluster size**.
*(This is also the mechanism behind zero-copy being defeated: if the broker must recompress, it can't just `sendfile()` bytes through. Matching client and broker compression codecs is how you keep this cheap.)*
#### 5.6 Kafka in the cloud [#56-kafka-in-the-cloud]
General method: **prioritize Kafka's performance characteristics, then pick the instance shape** — each instance type is a different mix of CPU, memory, IOPS, and disk.
##### Microsoft Azure [#microsoft-azure]
* **Disks are managed separately from the VM** → storage needs are decoupled from VM type.
* **Decision order:** start with **data retention required**, then **producer performance needed**.
* **Very low latency needed** → I/O optimized instances with **premium SSD**.
* Otherwise → **Azure Managed Disks** or **Azure Blob Storage** may suffice.
| Cluster size | Instance |
| ---------------------------------- | -------------------- |
| Smaller clusters, most use cases | **Standard D16s v3** |
| High performance / larger clusters | **D64s v4** |
* **Build the cluster in an Azure availability set** and **balance partitions across Azure compute fault domains** to ensure availability.
* **Strongly prefer Azure Managed Disks over ephemeral disks** — *"If a VM is moved, you run the risk of losing all the data on your Kafka broker."*
| Storage | Cost | SLA |
| ----------------------- | ---------------------- | ------------------------------------------------------ |
| HDD Managed Disks | relatively inexpensive | **no clearly defined availability SLA from Microsoft** |
| Premium SSD / Ultra SSD | much more expensive | **much quicker, 99.99% SLA** |
| Microsoft Blob Storage | — | option **if not latency sensitive** |
##### Amazon Web Services [#amazon-web-services]
* **Very low latency needed** → I/O optimized instances with **local SSD**.
* Otherwise → ephemeral storage (e.g. **Amazon EBS**) may suffice.
| Instance | Trade |
| ----------- | ------------------------------------------------------------------------------------ |
| **m4** | **greater retention periods**, but **lower disk throughput** (elastic block storage) |
| **r3** | **much better throughput** (local SSD), but **drives limit retained data** |
| **i2 / d2** | best of both worlds, but **significantly more expensive** |
***
# 2.11 Self-test (/docs/kafka/installing-kafka/self-test)
What are the *only two* config requirements for brokers to form one cluster?
Why 5 ZooKeeper nodes rather than 3, when 3 already tolerates a failure?
Why is `vm.swappiness=0` wrong, and when did it become wrong?
Your topic has 7-day retention but you find 17-day-old messages. Explain precisely why, and give two fixes.
`log.dirs` has three mounts; one is 95% full and the other two are near-empty. Why did Kafka do this?
Why does raising `message.max.bytes` risk both a stuck consumer *and* stalled replication?
`acks=all` is set and the producer got a success response, yet the message is gone. Explain the exact sequence.
What does RF++ mean and what two simultaneous events is it designed to survive?
Why does the broker need CPU at all, given it treats messages as opaque bytes?
Why is colocating another service on a broker a *consumer* performance problem specifically?
Give the outbound-network formula for sizing a broker's NIC. What term do people forget?
Rack awareness is configured and RF is 3, yet losing one rack took a partition offline. How?
Why does a ZooKeeper hiccup cause errors *hours later* during an unrelated controlled shutdown?
What are the current per-broker and per-cluster replica ceilings — and why is "replica" not "partition"?
You must set retention to exactly 24 hours for a compliance requirement, on a topic receiving 20 MB/day. What do you configure and why?
**Previous:** [Chapter 1 — Meet Kafka](01-meet-kafka.md)
**Next:** [Chapter 3 — Kafka Producers](03-kafka-producers.md)
# 2.0 The shape of a Kafka deployment (/docs/kafka/installing-kafka/shape-kafka-deployment)
**Two hard requirements to form a cluster — that's all:**
1. **Every broker has the same `zookeeper.connect`** (same ensemble + same path).
2. **Every broker has a unique `broker.id`.** If two brokers try to join with the same ID, **the second logs an error and fails to start.**
Everything else is tuning.
***
# 4.13 What actually breaks in production — Ch. 4 consolidated (/docs/kafka/kafka-consumers-reading-data/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| 1 | **Consumer group falls behind and never catches up** | Single consumer; high-latency per-record work (DB write) | Add consumers **up to the partition count**; beyond that, add partitions (mind Ch. 3 §9.4) |
| 2 | **Added consumers, throughput unchanged** | More consumers **than partitions** → the extras are **idle** | Partition count is the ceiling |
| 3 | **A second application "steals" messages from the first** | Both apps used **the same `group.id`** | One `group.id` **per application** |
| 4 | **Rebalance storm — group never stabilizes** | Processing per batch exceeds **`max.poll.interval.ms`** → evicted → rebalance → new owner also evicted | Lower **`max.poll.records`**; raise `max.poll.interval.ms`; move slow work off the poll thread |
| 5 | **Entire group pauses for seconds on every deploy** | **Eager** rebalance = "stop the world" | **`CooperativeStickyAssignor`** (mind the \<2.3 upgrade path) |
| 6 | **Consumers assigned very uneven partition counts** | **`RangeAssignor`** is the **default** and assigns **per topic independently** | `RoundRobinAssignor` / `StickyAssignor` / `CooperativeStickyAssignor` |
| 7 | **`session.timeout.ms` dead air after every restart** | `close()` not called → coordinator must wait for session timeout | Shutdown hook → `wakeup()` → `close()` |
| 8 | **Local cache/state rebuilt on every restart, taking minutes** | Dynamic membership → new member ID → new partitions | **`group.instance.id`** (static membership) |
| 9 | **After enabling static membership, partitions sit unconsumed after a crash** | Static members **don't leave proactively**; detection waits for `session.timeout.ms` | Tune `session.timeout.ms`: high enough to survive restarts, low enough to reassign on real failure |
| 10 | **`group.instance.id already exists` error** | Two consumers with the same `group.instance.id` | Unique per instance |
| 11 | **Duplicate processing after every crash** | **Autocommit** — last commit up to `auto.commit.interval.ms` (default **5 s**) old | Manual commits after processing; **duplicates can be reduced but never eliminated with autocommit** |
| 12 | **Silent data loss when processing throws mid-batch** | Autocommit + `catch`+`continue` → next `poll()` commits the whole batch including unprocessed records | *"critical to always process all events returned by `poll()` before calling `poll()` again"*; use manual commits |
| 13 | **Records committed but never processed (loss)** | `commitSync()` called **before** finishing the batch | Commit **after** processing |
| 14 | **Committed offset went backwards; duplicates spiked** | `commitAsync()` retried in a callback, landing an **older** offset after a newer one | Don't retry naively; use the **monotonic sequence-number pattern** |
| 15 | **Last batch reprocessed after every clean shutdown** | Only `commitAsync()` used; the final commit failed with no "next commit" to retry it | **`commitAsync()` in the loop + `commitSync()` on exit** |
| 16 | **Batch reprocessed after every rebalance** | No offsets committed before losing partitions | **`commitSync(currentOffsets)` in `onPartitionsRevoked()`** |
| 17 | **Off-by-one: one record duplicated (or skipped) per commit** | Committed `record.offset()` instead of `record.offset()+1` | The committed offset is **the next offset to read** |
| 18 | **A consumer down for a long weekend silently skips everything** | Committed offset aged out of retention + **`auto.offset.reset=latest`** (the default) | Set `earliest` (accept reprocessing) or **`none`** (fail loudly); retention ≥ max outage (Ch. 2) |
| 19 | **A monthly batch consumer group "forgets" everything** | Group empty longer than **`offsets.retention.minutes`** (default **7 days**) → offsets deleted → behaves as brand new | Raise the broker setting, or keep a member alive, or manage offsets externally |
| 20 | **Broker/network overloaded by metadata, not data** | **Regex `subscribe()`** on a cluster with \~30,000+ partitions — filtering is **client-side**, full topic list fetched at intervals | Explicit topic lists; *"bandwidth used by topic metadata can be larger than the bandwidth used to send data"* |
| 21 | **Regex subscription fails with authorization errors** | Regex requires **a full describe grant on the entire cluster** | Explicit lists, or grant cluster-wide describe (usually unacceptable) |
| 22 | **Migrating `poll(0)` to `poll(Duration.ofMillis(0))` broke metadata fetching** | `poll(long)` blocked for metadata regardless of timeout; `poll(Duration)` does not | Move the logic into **`onPartitionsAssigned()`** |
| 23 | **`ConcurrentModificationException` / bizarre consumer behavior** | Sharing one consumer across threads, or two same-group consumers in one thread | **One consumer per thread**; `ExecutorService` or a consumer→queue→workers pattern |
| 24 | **Consumer OOM after a rebalance shrinks the group** | `max.partition.fetch.bytes` × (now-larger) partition count | Use **`fetch.max.bytes`** — an absolute cap |
| 25 | **High consumer CPU on an idle topic** | Tight poll loop with a short timeout and `fetch.min.bytes=1` | Raise `fetch.min.bytes` and/or the poll timeout |
| 26 | **Latency spikes to \~500 ms on a low-traffic topic** | `fetch.min.bytes` raised, `fetch.max.wait.ms` left at the 500 ms default | Lower `fetch.max.wait.ms` to your SLA |
| 27 | **Lowering `request.timeout.ms` made an incident worse** | Disconnect/reconnect churn against an already-overloaded broker | *"recommended not to lower it"* |
| 28 | **Huge cross-AZ data transfer bill** | Consumers fetching from the **leader** across zones | `client.rack` + broker `RackAwareReplicaSelector` |
| 29 | **Rebalance times out during `onPartitionsAssigned()`** | Expensive state loading in the callback | Preparation must return within `max.poll.timeout.ms` |
| 30 | **Corrupted state after a cooperative rebalance edge case** | `onPartitionsLost()` not implemented → `onPartitionsRevoked()` ran instead, while **the new owner already saved its own state** | Implement `onPartitionsLost()` carefully, avoiding conflicts with the new owner |
| 31 | **Garbage/exceptions when deserializing** | Producer serializer ≠ consumer deserializer for that topic | Avro + Schema Registry so compatibility errors surface as messages, not byte diffs |
| 32 | **Standalone consumer silently ignores new partitions** | `assign()` gets **no notification** of partition additions | Poll `partitionsFor()` periodically, or bounce the app on partition changes |
| 33 | **Standalone consumer dies; nothing reads its partitions** | `assign()` has **no failover** | Use a consumer group unless you truly need fixed assignment |
***
# 4.7 Commits and offsets — where consumers actually break (/docs/kafka/kafka-consumers-reading-data/commits-offsets-where-consumers)
#### 7.1 The model [#71-the-model]
> *"One of Kafka's unique characteristics is that **it does not track acknowledgments from consumers the way many JMS queues do.** Instead, it allows consumers to use Kafka to **track their position (offset) in each partition.**"*
> *"Unlike traditional message queues, **Kafka does not commit records individually.** Instead, consumers **commit the last message they've successfully processed** from a partition and **implicitly assume that every message before the last was also successfully processed.**"*
**Mechanically:** the consumer *"sends a message to Kafka, which updates a special **`__consumer_offsets`** topic with the committed offset for each partition."*
**Why any of this matters:** *"As long as all your consumers are up, running, and churning away, this will have no impact. **However, if a consumer crashes or a new consumer joins the consumer group, this will trigger a rebalance.** After a rebalance, each consumer may be assigned a **new set of partitions**. In order to know where to pick up the work, **the consumer will read the latest committed offset of each partition and continue from there.**"*
#### 7.2 The two failure modes — this is the whole ballgame [#72-the-two-failure-modes--this-is-the-whole-ballgame]
> *"Clearly, managing offsets has a big impact on the client application."*
#### 7.3 ⚠️ "Which offset is committed?" — the off-by-one everyone hits [#73-️-which-offset-is-committed--the-off-by-one-everyone-hits]
> *"When committing offsets either automatically or without specifying the intended offsets, the default behavior is to commit **the offset AFTER the last offset that was returned by `poll()`.**"*
>
> The book's convention: *"it is tedious to repeatedly read 'Commit the offset that is one larger than the last offset the client received from `poll()`,' and 99% of the time it does not matter. So, we are going to write **'Commit the last offset'** when we refer to the default behavior — **and if you need to manually manipulate offsets, please keep this note in mind.**"*
**The rule when committing manually:**
```txt
committed offset = offset of the NEXT message your application will read
= last_processed_offset + 1
```
Get this wrong by one and you reprocess one record per commit forever (harmless-ish) or skip one record per commit (data loss).
#### 7.4 Automatic commit [#74-automatic-commit]
```java
enable.auto.commit = true
auto.commit.interval.ms = 5000 // default: 5 seconds
```
**Mechanics:** *"Just like everything else in the consumer, **the automatic commits are driven by the poll loop.** Whenever you poll, the consumer checks if it is time to commit, and if it is, **it will commit the offsets it returned in the last poll.**"*
**The duplicate window:**
> *"It is possible to configure the commit interval to commit more frequently and reduce the window in which records will be duplicated, but **it is impossible to completely eliminate them.**"*
**⚠️ The subtle killer:**
> *"With autocommit enabled, when it is time to commit offsets, **the next poll will commit the last offset returned by the previous poll. It doesn't know which events were actually processed**, so it is **critical to always process all the events returned by `poll()` before calling `poll()` again.** (Just like `poll()`, **`close()` also commits offsets automatically.**) This is usually not an issue, but **pay attention when you handle exceptions or exit the poll loop prematurely.**"*
**Verdict:** *"Automatic commits are convenient, but **they don't give developers enough control to avoid duplicate messages.**"*
#### 7.5 `commitSync()` — commit current offset [#75-commitsync--commit-current-offset]
Set `enable.auto.commit=false`. *"The simplest and most reliable of the commit APIs is `commitSync()`. This API will commit the latest offset returned by `poll()` and **return once the offset is committed, throwing an exception if the commit fails.**"*
```java
Duration timeout = Duration.ofMillis(100);
while (true) {
ConsumerRecords records = consumer.poll(timeout);
for (ConsumerRecord record : records) {
System.out.printf("topic = %s, partition = %d, offset = %d, " +
"customer = %s, country = %s\n",
record.topic(), record.partition(),
record.offset(), record.key(), record.value());
}
try {
consumer.commitSync(); // AFTER processing the whole batch
} catch (CommitFailedException e) {
log.error("commit failed", e);
}
}
```
**Positioning is everything:**
> *"`commitSync()` will commit the **latest offset returned by `poll()`**, so **if you call `commitSync()` before you are done processing all the records in the collection, you risk missing the messages that were committed but not processed**, in case the application crashes. If the application crashes **while it is still processing records** in the collection, **all the messages from the beginning of the most recent batch until the time of the rebalance will be processed twice** — this may or may not be preferable to missing messages."*
**Retry behavior:** *"`commitSync` retries committing as long as there is no error that can't be recovered. If this happens, **there is not much we can do except log an error.**"*
**Note also:** *"You should determine when you are 'done' with a record **according to your use case.**"* — the book is explicit that "processed" is your definition, not Kafka's.
#### 7.6 `commitAsync()` [#76-commitasync]
**The problem with sync:** *"the application is **blocked until the broker responds** to the commit request. This will **limit the throughput** of the application. Throughput can be improved by committing less frequently, but then we are **increasing the number of potential duplicates** that a rebalance may create."*
```java
while (true) {
ConsumerRecords records = consumer.poll(timeout);
for (ConsumerRecord record : records) { /* process */ }
consumer.commitAsync(); // send and carry on
}
```
##### ⚠️ Why `commitAsync()` deliberately does NOT retry [#️-why-commitasync-deliberately-does-not-retry]
This is one of the most instructive explanations in the book:
> *"The reason it does not retry is that **by the time `commitAsync()` receives a response from the server, there may have been a later commit that was already successful.**"*
**With a callback:**
```java
consumer.commitAsync(new OffsetCommitCallback() {
public void onComplete(Map offsets,
Exception e) {
if (e != null)
log.error("Commit failed for offsets {}", offsets, e);
}
});
```
> *"It is common to use the callback to **log commit errors or to count them in a metric**, but **if you want to use the callback for retries, you need to be aware of the problem with commit order.**"*
> ### Retrying async commits safely — the sequence-number pattern [#retrying-async-commits-safely--the-sequence-number-pattern]
>
> *"A simple pattern to get the commit order right for asynchronous retries is to use a **monotonically increasing sequence number**. Increase the sequence number every time you commit, and add the sequence number at the time of the commit to the `commitAsync` callback. **When you're getting ready to send a retry, check if the commit sequence number the callback got is equal to the instance variable; if it is, there was no newer commit and it is safe to retry. If the instance sequence number is higher, don't retry** because a newer commit was already sent."*
#### 7.7 Combining sync and async — the standard production shape [#77-combining-sync-and-async--the-standard-production-shape]
**The insight:** *"Normally, occasional failures to commit without retrying are not a huge problem, because **if the problem is temporary, the following commit will be successful.** But **if we know that this is the LAST commit before we close the consumer, or before a rebalance, we want to make extra sure that the commit succeeds.**"*
```java
Duration timeout = Duration.ofMillis(100);
try {
while (!closing) {
ConsumerRecords records = consumer.poll(timeout);
for (ConsumerRecord record : records) {
/* process */
}
consumer.commitAsync(); // ① fast; a failure is retried by the
// NEXT commit
}
consumer.commitSync(); // ② no "next commit" exists — retry
// until success or unrecoverable
} catch (Exception e) {
log.error("Unexpected error", e);
} finally {
consumer.close();
}
```
#### 7.8 Committing a specified offset [#78-committing-a-specified-offset]
**Why you'd need it:** *"Committing the latest offset only allows you to commit **as often as you finish processing batches.** But what if you want to commit more frequently? **What if `poll()` returns a huge batch and you want to commit offsets in the middle of the batch** to avoid having to process all those rows again if a rebalance occurs? **You can't just call `commitSync()` or `commitAsync()` — this will commit the last offset returned, which you didn't get to process yet.**"*
```java
private Map currentOffsets = new HashMap<>();
int count = 0;
Duration timeout = Duration.ofMillis(100);
while (true) {
ConsumerRecords records = consumer.poll(timeout);
for (ConsumerRecord record : records) {
System.out.printf("topic = %s, partition = %s, offset = %d, " +
"customer = %s, country = %s\n",
record.topic(), record.partition(), record.offset(),
record.key(), record.value());
currentOffsets.put(
new TopicPartition(record.topic(), record.partition()),
new OffsetAndMetadata(record.offset()+1, "no metadata"));
// ▲▲
// THE COMMITTED OFFSET SHOULD ALWAYS BE THE OFFSET OF
// THE NEXT MESSAGE YOUR APPLICATION WILL READ
if (count % 1000 == 0)
consumer.commitAsync(currentOffsets, null); // no callback → null
count++;
}
}
```
**Cost:** *"Since your consumer may be consuming more than a single partition, **you will need to track offsets on all of them, which adds complexity to your code.**"*
> You can commit *"based on time or perhaps content of the records"* — the `count % 1000` is just one policy. And `commitSync()` is *"also completely valid here."*
***
# 4.6 Configuring consumers (/docs/kafka/kafka-consumers-reading-data/configuring-consumers)
#### 6.1 Fetch sizing — the latency/efficiency dial [#61-fetch-sizing--the-latencyefficiency-dial]
##### `fetch.min.bytes` (default **1 byte**) [#fetchminbytes-default-1-byte]
Minimum data the consumer wants to receive from the broker per fetch. *"If a broker receives a request for records from a consumer but the new records amount to fewer bytes than `fetch.min.bytes`, **the broker will wait until more messages are available** before sending records back."*
**Why raise it:**
* *"reduces the load on both the consumer and the broker, as they have to handle fewer back-and-forth messages in cases where the topics don't have much new activity (or for **lower-activity hours of the day**)"*
* *"if the consumer is **using too much CPU when there isn't much data available**"*
* *"to **reduce load on the brokers when you have a large number of consumers**"*
**Cost:** *"increasing this value can increase latency for low-throughput cases."*
##### `fetch.max.wait.ms` (default **500 ms**) [#fetchmaxwaitms-default-500-ms]
The companion cap — **how long the broker will wait** to satisfy `fetch.min.bytes`.
> *"This results in up to 500 ms of extra latency in case there is not enough data flowing... If you want to limit the potential latency (usually due to **SLAs controlling the maximum latency of the application**), set `fetch.max.wait.ms` to a lower value."*
##### `fetch.max.bytes` (default **50 MB**) — prefer this one [#fetchmaxbytes-default-50-mb--prefer-this-one]
Maximum bytes Kafka returns per poll of a broker. *"Used to limit the size of memory that the consumer will use to store data returned from the server, **irrespective of how many partitions or messages were returned.**"*
**The progress guarantee:** *"records are sent to the client in batches, and **if the first record-batch that the broker has to send exceeds this size, the batch will be sent and the limit will be ignored.** This guarantees that the consumer can continue making progress."*
*(Nice design detail: the limit is advisory when honoring it would deadlock the consumer. Contrast with Ch. 2's stuck-consumer bug, where a mismatched `fetch.message.max.bytes` genuinely does stall.)*
**Matching broker config exists too:** *"requests for large amounts of data can result in **large reads from disk and long sends over the network**, which can cause contention and increase load on the broker."*
##### `max.partition.fetch.bytes` (default **1 MB**) — usually the wrong knob [#maxpartitionfetchbytes-default-1-mb--usually-the-wrong-knob]
Max bytes the server returns **per partition**. When `poll()` returns `ConsumerRecords`, the object uses at most this **per partition assigned to the consumer**.
> *"**controlling memory usage using this configuration can be quite complex, as you have no control over how many partitions will be included in the broker response.** Therefore, **we highly recommend using `fetch.max.bytes` instead**, unless you have special reasons to try and process similar amounts of data from each partition."*
##### `max.poll.records` [#maxpollrecords]
*"the maximum **number** of records that a single call to `poll()` will return. Use this to control the amount of data (**but not the size of data**) your application will need to process in one iteration of the poll loop."*
**Its real job** is bounding the time between `poll()` calls, so `max.poll.interval.ms` isn't tripped — see below.
#### 6.2 Liveness and timeouts [#62-liveness-and-timeouts]
##### `session.timeout.ms` (default **10 s**) + `heartbeat.interval.ms` [#sessiontimeoutms-default-10-s--heartbeatintervalms]
**The trade-off:**
| | Effect |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------- |
| **Lower `session.timeout.ms`** | *"allows consumer groups to **detect and recover from failure sooner** but may also cause **unwanted rebalances**"* |
| **Higher `session.timeout.ms`** | *"reduces the chance of accidental rebalance but also means **it will take longer to detect a real failure**"* |
##### `max.poll.interval.ms` (default **5 minutes**) [#maxpollintervalms-default-5-minutes]
The backstop for a **hung main thread with a healthy heartbeat thread**.
> *"The easiest way to know whether the consumer is still processing records is to **check whether it is asking for more records.** However, the intervals between requests for more records are **difficult to predict** and depend on the amount of available data, the type of processing done by the consumer, and sometimes on the latency of additional services."*
**How to size it:** *"It has to be an interval **large enough that it will very rarely be reached by a healthy consumer** but **low enough to avoid significant impact from a hanging consumer.**"*
**What happens on breach:** *"the background thread will send a **'leave group' request** to let the broker know the consumer is dead and the group must rebalance, and then **stop sending heartbeats.**"*
**The relationship with `max.poll.records`:** *"In applications that need to do time-consuming processing on each record returned, `max.poll.records` is used to **limit the amount of data returned and therefore limit the duration before the application is available to `poll()` again.** Even with `max.poll.records` defined, the interval between calls to `poll()` is difficult to predict, and `max.poll.interval.ms` is used as a fail-safe or backstop."*
##### `default.api.timeout.ms` (default **1 minute**) [#defaultapitimeoutms-default-1-minute]
Applies to **(almost) all API calls** when you don't specify an explicit timeout. *"Since it is higher than the request timeout default, **it will include a retry when needed.** The notable exception is the **`poll()` method, that always requires an explicit timeout.**"*
##### `request.timeout.ms` (default **30 s**) [#requesttimeoutms-default-30-s]
Max time the consumer waits for a broker response. On breach: *"the client will assume the broker will not respond at all, **close the connection, and attempt to reconnect.**"*
> **"It is recommended NOT to lower it.** It is important to leave the broker with enough time to process the request before giving up — **there is little to gain by resending requests to an already overloaded broker, and the act of disconnecting and reconnecting adds even more overhead.**"\*
That's an anti-thundering-herd argument: shortening this timeout under load makes an overloaded broker *worse*.
#### 6.3 `auto.offset.reset` — the data-loss/replay switch [#63-autooffsetreset--the-data-lossreplay-switch]
**When it applies:** the consumer starts reading a partition for which it has **no committed offset**, **or** the committed offset is **invalid** — *"usually because the consumer was down for so long that the record with that offset was already aged out of the broker."*
| Value | Behavior | Risk |
| ---------------------- | ------------------------------------------------------------------------ | ---------------------------------------------------------------------------- |
| **`latest`** (default) | Start from the **newest records** (written after the consumer started) | **SILENT DATA LOSS** — everything between the lost offset and now is skipped |
| **`earliest`** | Read **all data in the partition from the very beginning** | **MASSIVE REPROCESSING** — duplicates, possibly hours/days of it |
| **`none`** | **Throw an exception** when attempting to consume from an invalid offset | Fails loudly — *often the right choice for correctness-critical apps* |
This config, combined with Ch. 2's retention discussion, is exactly how "we lost a day of data and nobody noticed" happens: a consumer down longer than retention + `auto.offset.reset=latest` = silent skip.
#### 6.4 `enable.auto.commit` (default **true**) [#64-enableautocommit-default-true]
> *"Set it to **false** if you prefer to control when offsets are committed, which is **necessary to minimize duplicates and avoid missing data.**"* Interval controlled by `auto.commit.interval.ms`. Full discussion in §7.
#### 6.5 `partition.assignment.strategy` — the four assignors [#65-partitionassignmentstrategy--the-four-assignors]
Setup for all examples: **consumers C1, C2** subscribed to **topics T1, T2**, each topic with **3 partitions**.
##### Range (default — `RangeAssignor`) [#range-default--rangeassignor]
*"Assigns to each consumer a **consecutive subset of partitions from each topic** it subscribes to."*
**Note this is the *default* and it is systematically unbalanced.** The imbalance compounds per topic.
##### RoundRobin (`RoundRobinAssignor`) [#roundrobin-roundrobinassignor]
*"Takes **all the partitions from all subscribed topics** and assigns them to consumers **sequentially, one by one.**"*
##### Sticky (`StickyAssignor`) — **two goals** [#sticky-stickyassignor--two-goals]
1. *"an assignment that is as **balanced as possible**"*
2. *"in case of a rebalance, **leave as many assignments as possible in place**, minimizing the overhead associated with moving partition assignments from one consumer to another"*
> *"In the common case where all consumers are subscribed to the same topic, the **initial** assignment from Sticky will be **as balanced as RoundRobin. Subsequent assignments will be just as balanced but will reduce the number of partition movements.** In cases where consumers in the same group subscribe to **different** topics, the assignment achieved by Sticky is **more balanced than RoundRobin.**"*
##### Cooperative Sticky (`CooperativeStickyAssignor`) [#cooperative-sticky-cooperativestickyassignor]
*"Identical to Sticky but **supports cooperative rebalances** in which consumers can continue consuming from the partitions that are not reassigned."*
> ⚠️ *"if you are **upgrading from a version older than 2.3**, you'll need to follow **a specific upgrade path** in order to enable the cooperative sticky assignment strategy, so **pay extra attention to the upgrade guide.**"*
**Summary table:**
| Assignor | Balance | Rebalance cost | Notes |
| ------------------------------- | ---------------------------------------------------------------- | ----------------------------------- | ---------------------------------------------- |
| `RangeAssignor` | **Poor** when consumers don't divide partitions evenly | Eager ("stop the world") | **The default** |
| `RoundRobinAssignor` | Good (±1) when all subscribe to same topics | Eager | |
| `StickyAssignor` | Good; **better than RoundRobin** for heterogeneous subscriptions | Eager, but **minimizes movement** | |
| **`CooperativeStickyAssignor`** | Same as Sticky | **Incremental — no stop-the-world** | **Requires a careful upgrade path from \<2.3** |
You can also **implement your own** and point `partition.assignment.strategy` at your class.
#### 6.6 `client.rack` — fetch from the closest replica [#66-clientrack--fetch-from-the-closest-replica]
> *"By default, consumers will fetch messages from the **leader replica** of each partition. However, when the cluster spans **multiple datacenters or multiple cloud availability zones**, there are advantages both in **performance and in cost** to fetching from a replica located **in the same zone as the consumer.**"*
**Two-part setup — you need BOTH:**
You can also *"implement your own `replica.selector.class` with custom logic for choosing the best replica to consume from, based on client metadata and partition metadata."*
*(Recall Ch. 1: consumers MAY fetch from a follower; producers MUST use the leader. This config is what cashes in that asymmetry.)*
#### 6.7 The rest [#67-the-rest]
| Config | Notes |
| ------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **`client.id`** | Any string; used by brokers to identify requests (e.g. fetch requests). **Logging, metrics, and quotas** |
| **`group.instance.id`** | Any unique string → **static group membership** (§3) |
| **`receive.buffer.bytes` / `send.buffer.bytes`** | TCP socket buffers; **`-1` = OS defaults**. *"It can be a good idea to increase these when producers or consumers communicate with brokers in a different datacenter"* (higher latency, lower bandwidth) |
#### 6.8 ⚠️ `offsets.retention.minutes` — a BROKER config that silently resets your consumers [#68-️-offsetsretentionminutes--a-broker-config-that-silently-resets-your-consumers]
> *"This is a broker configuration, but it is important to be aware of it **due to its impact on consumer behavior.**"*
> *"Note that this behavior changed a few times, so **if you use versions older than 2.1.0, check the documentation for your version** for the expected behavior."*
**Real scenario:** a batch job's consumer group runs monthly. Between runs the group is empty for \~30 days > 7 days → offsets vanish → next run either reprocesses everything or skips everything. This is a classic and very confusing production surprise.
***
# 4.9 Consuming from specific offsets (/docs/kafka/kafka-consumers-reading-data/consuming-specific-offsets)
```java
consumer.seekToBeginning(Collection tp); // replay all
consumer.seekToEnd(Collection tp); // skip to now
consumer.seek(TopicPartition tp, long offset); // exact position
```
**Two motivating use cases from the book:**
* *"a **time-sensitive application** could **skip ahead** a few records when falling behind"*
* *"a consumer that writes data to a file could be **reset back to a specific point in time in order to recover data if the file was lost**"*
#### Seek by timestamp — the recovery pattern [#seek-by-timestamp--the-recovery-pattern]
```java
Long oneHourEarlier = Instant.now().atZone(ZoneId.systemDefault())
.minusHours(1).toEpochSecond();
Map partitionTimestampMap = consumer.assignment()
.stream()
.collect(Collectors.toMap(tp -> tp, tp -> oneHourEarlier)); // ①
Map offsetMap
= consumer.offsetsForTimes(partitionTimestampMap); // ②
for (Map.Entry entry: offsetMap.entrySet()) {
consumer.seek(entry.getKey(), entry.getValue().offset()); // ③
}
```
① Map every partition **assigned to this consumer** (`consumer.assignment()`) to the target timestamp.
② `offsetsForTimes()` — *"sends a request to the broker where **a timestamp index** is used to return the relevant offsets."*
③ `seek()` each partition to the returned offset.
*(Cross-reference Ch. 2: the resolution of timestamp→offset lookups depends on **log segment size**, because Kafka locates the segment that was being written at that time and returns the offset at its beginning. Big segments on a low-volume topic = imprecise seeks.)*
***
# 4.4 Creating a consumer (/docs/kafka/kafka-consumers-reading-data/creating-consumer)
#### Three mandatory + one near-mandatory property [#three-mandatory--one-near-mandatory-property]
| Property | Notes |
| ------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`bootstrap.servers`** | Same as the producer (Ch. 3) |
| **`key.deserializer`** | Class taking a **byte array → Java object** (inverse of a serializer) |
| **`value.deserializer`** | Same, for the value |
| **`group.id`** | *"not strictly mandatory but very commonly used"* — the consumer group this instance belongs to. *"While it is possible to create consumers that do not belong to any consumer group, this is uncommon"* (see §10) |
```java
Properties props = new Properties();
props.put("bootstrap.servers", "broker1:9092,broker2:9092");
props.put("group.id", "CountryCounter");
props.put("key.deserializer",
"org.apache.kafka.common.serialization.StringDeserializer");
props.put("value.deserializer",
"org.apache.kafka.common.serialization.StringDeserializer");
KafkaConsumer consumer =
new KafkaConsumer(props);
```
#### Subscribing [#subscribing]
```java
// explicit list
consumer.subscribe(Collections.singletonList("customerCountries"));
// regular expression
consumer.subscribe(Pattern.compile("test.*"));
```
**Regex subscription behavior:** *"if someone creates a new topic with a name that matches, **a rebalance will happen almost immediately** and the consumers will start consuming from the new topic."* Most commonly used in *"applications that replicate data between Kafka and another system, or streams processing applications."*
> ### ⚠️ WARNING — regex subscription is a client-side scan with real cost [#️-warning--regex-subscription-is-a-client-side-scan-with-real-cost]
>
> If your cluster has a large number of partitions — **perhaps 30,000 or more** — be aware that **the filtering of topics for the subscription is done on the client side.**
>
> **What actually happens:**
>
> ```
> consumer requests THE LIST OF ALL TOPICS AND THEIR PARTITIONS
> from the broker at REGULAR INTERVALS
> → client uses this list to detect new matching topics
> → subscribes to them
> ```
>
> **Consequences when the topic list is large and there are many consumers:**
>
> * *"the size of the list of topics and partitions is significant"*
> * *"the regular expression subscription has **significant overhead on the broker, client, and network**"*
> * **"There are cases where the bandwidth used by the topic metadata is larger than the bandwidth used to send data."**
>
> **And a security consequence people miss:** *"in order to subscribe with a regular expression, the client needs permissions to **describe all topics in the cluster** — that is, **a full describe grant on the entire cluster.**"*
That last point often blocks regex subscriptions outright in locked-down multi-tenant clusters — you cannot grant least-privilege topic access *and* use regex subscription.
***
# 4.14 Deploy / monitor / scale / recover (/docs/kafka/kafka-consumers-reading-data/deploy-monitor-scale-recover)
#### Config recipes [#config-recipes]
#### Monitoring [#monitoring]
| Signal | Why it matters (per this chapter) |
| --------------------------------------------------------------- | -------------------------------------------------------------------------- |
| **Consumer lag per partition** | The master metric. Lag vs **retention** = imminent silent data loss (#18) |
| **Rebalance rate / rebalance duration** | Detects #4/#5; eager rebalances show as group-wide throughput dips |
| **Commit rate + commit failure rate** | `commitAsync` failures are invisible unless you count them in the callback |
| **`records-lag-max`, `records-consumed-rate`, `fetch-latency`** | Fetch tuning feedback |
| **Time between `poll()` calls** (your own metric) | The direct predictor of `max.poll.interval.ms` eviction |
| **Assigned partition count per instance** | Detects assignor imbalance (#6) and idle consumers (#2) |
| **Number of live group members** | Static membership hides deaths until `session.timeout.ms` |
| **`__consumer_offsets` health** | Where all commits actually land |
#### Scaling [#scaling]
#### "Backup" / recovery, consumer edition [#backup--recovery-consumer-edition]
The consumer's recovery story *is* offset manipulation:
**The important framing:** in Kafka, a consumer's "backup" is *the log plus a correct offset*. That's why offset-commit discipline is not a code-style concern — it is your recovery point objective.
***
# 4.11 Deserializers (/docs/kafka/kafka-consumers-reading-data/deserializers)
> *"It should be obvious that **the serializer used to produce events to Kafka must match the deserializer used when consuming events.** Serializing with `IntSerializer` and then deserializing with `StringDeserializer` **will not end well.** This means that, as a developer, **you need to keep track of which serializers were used to write into each topic** and make sure each topic only contains data that your deserializers can interpret."*
**Why Avro + Schema Registry fixes this class of problem:**
> *"the `AvroSerializer` can make sure that **all the data written to a specific topic is compatible with the schema of the topic**, which means it can be deserialized with the matching deserializer and schema. **Any errors in compatibility — on the producer or the consumer side — will be caught easily with an appropriate error message, which means you will not need to try to debug byte arrays for serialization errors.**"*
#### Custom deserializer (shown to argue against it) [#custom-deserializer-shown-to-argue-against-it]
```java
public class CustomerDeserializer implements Deserializer {
@Override
public void configure(Map configs, boolean isKey) { /* nothing */ }
@Override
public Customer deserialize(String topic, byte[] data) {
int id; int nameSize; String name;
try {
if (data == null) return null;
if (data.length < 8)
throw new SerializationException("Size of data received " +
"by deserializer is shorter than expected");
ByteBuffer buffer = ByteBuffer.wrap(data);
id = buffer.getInt();
nameSize = buffer.getInt();
byte[] nameBytes = new byte[nameSize];
buffer.get(nameBytes);
name = new String(nameBytes, "UTF-8");
return new Customer(id, name);
} catch (Exception e) {
throw new SerializationException(
"Error when deserializing byte[] to Customer " + e);
}
}
@Override
public void close() { /* nothing */ }
}
```
> *"**The consumer also needs the implementation of the `Customer` class, and both the class and the serializer need to match on the producing and consuming applications. In a large organization with many consumers and producers sharing access to the data, this can become challenging.**"*
>
> *"implementing a custom serializer and deserializer **is not recommended. It tightly couples producers and consumers and is fragile and error prone.** A better solution would be to use a standard message format, such as **JSON, Thrift, Protobuf, or Avro.**"*
#### Avro deserialization [#avro-deserialization]
```java
Properties props = new Properties();
props.put("bootstrap.servers", "broker1:9092,broker2:9092");
props.put("group.id", "CountryCounter");
props.put("key.deserializer",
"org.apache.kafka.common.serialization.StringDeserializer");
props.put("value.deserializer",
"io.confluent.kafka.serializers.KafkaAvroDeserializer"); // ①
props.put("specific.avro.reader","true"); // ③
props.put("schema.registry.url", schemaUrl); // ②
String topic = "customerContacts";
KafkaConsumer consumer = new KafkaConsumer<>(props);
consumer.subscribe(Collections.singletonList(topic));
while (true) {
ConsumerRecords records = consumer.poll(timeout);
for (ConsumerRecord record: records) {
System.out.println("Current customer name is: " +
record.value().getName()); // ④
}
consumer.commitSync();
}
```
① `KafkaAvroDeserializer` deserializes the Avro messages.
② `schema.registry.url` — *"points to where we store the schemas. **This way, the consumer can use the schema that was registered by the producer to deserialize the message.**"*
③ `specific.avro.reader=true` + generated class `Customer` as the value type.
④ `record.value()` **is a `Customer` instance.**
***
# 4.10 Exiting cleanly — `wakeup()` and `close()` (/docs/kafka/kafka-consumers-reading-data/exiting-cleanly-wakeup-close)
**Mechanics:**
> *"Calling `wakeup` will cause `poll()` to exit with **`WakeupException`**, or **if `consumer.wakeup()` was called while the thread was not waiting on poll, the exception will be thrown on the next iteration when `poll()` is called.** The `WakeupException` **doesn't need to be handled**, but **before exiting the thread, you must call `consumer.close()`.**"*
**Why `close()` is not optional:**
Skip `close()` and you pay `session.timeout.ms` of dead air on every deploy — across every partition that consumer owned.
```java
Runtime.getRuntime().addShutdownHook(new Thread() {
public void run() {
System.out.println("Starting exit...");
consumer.wakeup(); // ① only safe cross-thread call
try {
mainThread.join();
} catch (InterruptedException e) {
e.printStackTrace();
}
}
});
...
Duration timeout = Duration.ofMillis(10000); // ② a long poll timeout
try {
// looping until ctrl-c, the shutdown hook will cleanup on exit
while (true) {
ConsumerRecords records = movingAvg.consumer.poll(timeout);
System.out.println(System.currentTimeMillis() + "-- waiting for data...");
for (ConsumerRecord record : records) {
System.out.printf("offset = %d, key = %s, value = %s\n",
record.offset(), record.key(), record.value());
}
for (TopicPartition tp: consumer.assignment())
System.out.println("Committing offset at position:" +
consumer.position(tp));
movingAvg.consumer.commitSync();
}
} catch (WakeupException e) {
// ignore for shutdown // ③
} finally {
consumer.close(); // ④
System.out.println("Closed consumer and we are done");
}
```
① *"`ShutdownHook` runs in a separate thread, so **the only safe action you can take is to call `wakeup`** to break out of the poll loop."*
② **On long poll timeouts:** *"If the poll loop is short enough and you don't mind waiting a bit before exiting, **you don't need to call `wakeup` — just checking an atomic boolean in each iteration would be enough.** Long poll timeouts are useful when consuming **low-throughput topics**; this way, **the client uses less CPU for constantly looping while the broker has no new data to return.**"*
③ *"You'll want to catch the exception to make sure your application doesn't exit unexpectedly, but **there is no need to do anything with it.**"*
④ *"Before exiting the consumer, make sure you close it cleanly."*
***
# 4. Kafka Consumers: Reading Data from Kafka (/docs/kafka/kafka-consumers-reading-data)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 4*
> **Learning goal:** *"Reading data from Kafka is a bit different than reading data from other messaging systems."* The two things that make it different — **consumer groups + rebalancing** and **offset commits** — are also the two things that cause virtually every consumer bug in production. Master those and the API is trivial.
***
# 4.5 The poll loop (/docs/kafka/kafka-consumers-reading-data/poll-loop)
```java
Duration timeout = Duration.ofMillis(100);
while (true) { // ①
ConsumerRecords records = consumer.poll(timeout); // ②
for (ConsumerRecord record : records) { // ③
System.out.printf("topic = %s, partition = %d, offset = %d, " +
"customer = %s, country = %s\n",
record.topic(), record.partition(), record.offset(),
record.key(), record.value());
int updatedCount = 1;
if (custCountryMap.containsKey(record.value())) {
updatedCount = custCountryMap.get(record.value()) + 1;
}
custCountryMap.put(record.value(), updatedCount);
JSONObject json = new JSONObject(custCountryMap);
System.out.println(json.toString());
}
}
```
**① The infinite loop.** *"Consumers are usually long-running applications that continuously poll Kafka for more data."*
**② The book's own words: "This is the most important line in the chapter."**
> *"The same way that **sharks must keep moving or they die, consumers must keep polling Kafka or they will be considered dead** and the partitions they are consuming will be handed to another consumer in the group."*
The `timeout` parameter controls **how long `poll()` will block if data is not available in the consumer buffer.** If set to 0, or if records are already available, `poll()` returns immediately; otherwise it waits up to that many ms.
**③ Each record contains:** topic, partition, **offset within the partition**, key, value.
#### What `poll()` secretly does — the reason exceptions surface here [#what-poll-secretly-does--the-reason-exceptions-surface-here]
> *"This means that **almost everything that can go wrong with a consumer, or in the callbacks used in its listeners, is likely to show up as an exception thrown by `poll()`.**"*
**And the constraint that follows:**
> *"if `poll()` is not invoked for longer than `max.poll.interval.ms`, the consumer will be considered dead and evicted from the consumer group, so **avoid doing anything that can block for unpredictable intervals inside the poll loop.**"*
#### Thread safety — the rule [#thread-safety--the-rule]
*(Note the contrast with Ch. 3: a **producer** IS thread-safe and shareable. A **consumer** is not.)*
**Two patterns for concurrency:**
1. *"wrap the consumer logic in its own object and then use Java's **`ExecutorService`** to start multiple threads, each with its own consumer."*
2. *"have **one consumer populate a queue of events** and have **multiple worker threads** perform work from this queue."* (Pattern documented by Igor Buzatović.)
> ### ⚠️ WARNING — `poll(long)` vs `poll(Duration)` semantics changed [#️-warning--polllong-vs-pollduration-semantics-changed]
>
> The old `poll(long)` signature is **deprecated**; the new API is `poll(Duration)`. **The blocking semantics subtly changed:**
>
> | | Behavior |
> | ---------------------- | ------------------------------------------------------------------------------------------------------ |
> | `poll(long)` (old) | **Blocks as long as it takes to get the needed metadata from Kafka — even if longer than the timeout** |
> | `poll(Duration)` (new) | **Adheres to the timeout restrictions and does NOT wait for metadata** |
>
> **The migration trap:** *"If you have existing consumer code that uses `poll(0)` as a method to **force Kafka to get the metadata without consuming any records** (a rather common hack), **you can't just change it to `poll(Duration.ofMillis(0))` and expect the same behavior.** You'll need to figure out a new way to achieve your goals."*
>
> **The recommended replacement:** *"Often the solution is placing the logic in the **`rebalanceListener.onPartitionAssignment()`** method, which is guaranteed to get called **after you have metadata for the assigned partitions but before records start arriving.**"*
***
# 4.1 What problem do consumer groups solve? (/docs/kafka/kafka-consumers-reading-data/problem-consumer-groups-solve)
**The starting point:** an application reads from a topic, runs validations, writes results to another data store. One consumer object, one subscription. Works fine.
**Then:** *"what if the rate at which producers write messages to the topic exceeds the rate at which your application can validate them? If you are limited to a single consumer reading and processing the data, your application may fall further and further behind, unable to keep up with the rate of incoming messages."*
**Why this is the normal case, not an edge case:**
> *"It is common for Kafka consumers to do **high-latency operations** such as write to a database or a time-consuming computation on the data. In these cases, **a single consumer can't possibly keep up** with the rate data flows into a topic."*
**The solution:** *"Just like multiple producers can write to the same topic, we need to allow multiple consumers to read from the same topic, splitting the data among them."*
#### The scaling ladder — memorize these four states [#the-scaling-ladder--memorize-these-four-states]
> *"**The main way we scale data consumption from a Kafka topic is by adding more consumers to a consumer group.** ... This is a good reason to create topics with a large number of partitions — it allows adding more consumers when the load increases. **Keep in mind that there is no point in adding more consumers than you have partitions in a topic** — some of the consumers will just be idle."*
**This closes the loop with Ch. 2:** partition count is your permanent ceiling on consumer parallelism, and partitions can only be added (which breaks keyed routing, per Ch. 3). Hence: *size partitions for future consumer parallelism, up front.*
#### Multiple groups = multiple independent readers [#multiple-groups--multiple-independent-readers]
> *"One of the main design goals in Kafka was to make the data produced to Kafka topics available for **many use cases throughout the organization.** ... **To make sure an application gets all the messages in a topic, ensure the application has its own consumer group.** Unlike many traditional messaging systems, **Kafka scales to a large number of consumers and consumer groups without reducing performance.**"*
#### The rule, stated once [#the-rule-stated-once]
***
# 4.8 Rebalance listeners (/docs/kafka/kafka-consumers-reading-data/rebalance-listeners)
**Why:** *"a consumer will want to do some cleanup work **before exiting** and also **before partition rebalancing.** If you know your consumer is about to lose ownership of a partition, you will want to **commit offsets of the last event you've processed.** Perhaps you also need to **close file handles, database connections**, and such."*
Pass a `ConsumerRebalanceListener` to **`subscribe()`**.
#### The three methods [#the-three-methods]
| Method | When called | What to do there |
| -------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`onPartitionsAssigned(partitions)`** | **After** partitions are reassigned to the consumer, **before** it starts consuming | *"prepare or load any state that you want to use with the partition, **seek to the correct offsets** if needed."* ⚠️ *"Any preparation done here should be **guaranteed to return within `max.poll.timeout.ms`** so the consumer can successfully join the group."* |
| **`onPartitionsRevoked(partitions)`** | When the consumer must give up partitions it owned — **rebalance or close**. **Eager:** invoked *before* rebalancing starts and *after* the consumer stopped consuming. **Cooperative:** invoked *at the END* of the rebalance, with **just the subset** being given up | **"This is where you want to commit offsets, so whoever gets this partition next will know where to start."** |
| **`onPartitionsLost(partitions)`** | **Cooperative only**, and **only in exceptional cases** where partitions were assigned to other consumers **without first being revoked** | *"clean up any state or resources used with these partitions."* ⚠️ *"has to be done carefully — **the new owner may have already saved its own state**, and you'll need to avoid conflicts."* **If you don't implement it, `onPartitionsRevoked()` is called instead** |
> ### TIP — cooperative rebalancing callback semantics [#tip--cooperative-rebalancing-callback-semantics]
>
> * **`onPartitionsAssigned()`** — invoked on **every** rebalance, as a way of **notifying the consumer that a rebalance happened**. If there are no new partitions assigned, **it is called with an empty collection.**
> * **`onPartitionsRevoked()`** — invoked in normal rebalancing conditions, **but only if the consumer gave up ownership of partitions. It will NOT be called with an empty collection.**
> * **`onPartitionsLost()`** — invoked in exceptional rebalancing conditions; **the partitions in the collection will already have new owners by the time the method is invoked.**
>
> **The ordering guarantee:** *"If you implemented all three methods, you are guaranteed that during a normal rebalance, **`onPartitionsAssigned()` will be called by the new owner of the partitions only AFTER the previous owner completed `onPartitionsRevoked()` and gave up its ownership.**"*
That guarantee is what makes handoff of external state (locks, file handles, DB transactions) actually safe.
#### The canonical example [#the-canonical-example]
```java
private Map currentOffsets = new HashMap<>();
Duration timeout = Duration.ofMillis(100);
private class HandleRebalance implements ConsumerRebalanceListener {
public void onPartitionsAssigned(Collection partitions) {
// nothing needed — we just start consuming
}
public void onPartitionsRevoked(Collection partitions) {
System.out.println("Lost partitions in rebalance. " +
"Committing current offsets:" + currentOffsets);
consumer.commitSync(currentOffsets); // ← SYNC, on purpose
}
}
try {
consumer.subscribe(topics, new HandleRebalance()); // ← THE KEY LINE
while (true) {
ConsumerRecords records = consumer.poll(timeout);
for (ConsumerRecord record : records) {
System.out.printf("topic = %s, partition = %s, offset = %d, " +
"customer = %s, country = %s\n",
record.topic(), record.partition(), record.offset(),
record.key(), record.value());
currentOffsets.put(
new TopicPartition(record.topic(), record.partition()),
new OffsetAndMetadata(record.offset()+1, null));
}
consumer.commitAsync(currentOffsets, null);
}
} catch (WakeupException e) {
// ignore, we're closing
} catch (Exception e) {
log.error("Unexpected error", e);
} finally {
try {
consumer.commitSync(currentOffsets);
} finally {
consumer.close();
System.out.println("Closed consumer and we are done");
}
}
```
Two authorial notes worth keeping:
* *"We are committing offsets for **all** partitions, not just the partitions we are about to lose — because the offsets are for events that were **already processed**, there is no harm in that."*
* *"we are using **`commitSync()`** to make sure the offsets are committed **before the rebalance proceeds.**"*
* **"The most important part: pass the `ConsumerRebalanceListener` to the `subscribe()` method so it will get invoked."**
***
# 4.2 Rebalancing — the mechanism, and why it hurts (/docs/kafka/kafka-consumers-reading-data/rebalancing-mechanism-hurts)
**When does a rebalance happen?**
1. A **new consumer joins** the group → it starts consuming partitions previously owned by another consumer
2. A consumer **shuts down or crashes** → it leaves; its partitions go to a remaining consumer
3. **The topics the group consumes are modified** — e.g. an administrator **adds new partitions**
> *"Moving partition ownership from one consumer to another is called a **rebalance**. Rebalances are important because they provide the consumer group with **high availability and scalability** ... but **in the normal course of events they can be fairly undesirable.**"*
#### 2.1 Eager rebalance — "stop the world" [#21-eager-rebalance--stop-the-world]
> *"During an eager rebalance, **all consumers stop consuming, give up their ownership of all partitions, rejoin the consumer group, and get a brand-new partition assignment.** This is essentially **a short window of unavailability of the entire consumer group.** The length of the window depends on the size of the consumer group as well as on several configuration parameters."*
#### 2.2 Cooperative (incremental) rebalance [#22-cooperative-incremental-rebalance]
> *"Cooperative rebalances (also called **incremental rebalances**) typically involve **reassigning only a small subset of the partitions** from one consumer to another, and **allowing consumers to continue processing records from all the partitions that are not reassigned.** ... **This is especially important in large consumer groups where rebalances can take a significant amount of time.**"*
**Sequence:** (1) group leader informs all consumers they will lose ownership of a subset; consumers stop consuming those and give up ownership. (2) Group leader assigns the now-orphaned partitions to new owners.
#### 2.3 Heartbeats, the group coordinator, and the two death-detection paths [#23-heartbeats-the-group-coordinator-and-the-two-death-detection-paths]
> *"As long as the consumer is sending heartbeats at regular intervals, it is assumed to be alive. If the consumer **stops sending heartbeats for long enough, its session will timeout** and the group coordinator will consider it dead and trigger a rebalance."*
**The two failure paths — and why there must be two:**
| Path | Detects | Config | Why it exists |
| --------------------------------------- | ------------------------------------------------------------------------------------------------ | ---------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Session timeout** (missed heartbeats) | The **process is dead / unreachable** | `session.timeout.ms` + `heartbeat.interval.ms` | The obvious case |
| **Poll interval timeout** | The **main thread is stuck** while the background heartbeat thread is still happily heartbeating | `max.poll.interval.ms` | *"There is a possibility that the main thread consuming from Kafka is **deadlocked, but the background thread is still sending heartbeats.** This means that records from partitions owned by this consumer **are not being processed.**"* |
**Crash vs clean shutdown:**
**This is why `close()` matters operationally**, not just for tidiness. See §7.
#### 2.4 How partition assignment actually works [#24-how-partition-assignment-actually-works]
**Design note worth appreciating:** assignment logic runs **on a client**, not the broker. That's why you can plug in your own `PartitionAssignor` without touching the cluster — and why the *leader's* Kafka client version determines which strategies are available.
***
# 4.15 Self-test (/docs/kafka/kafka-consumers-reading-data/self-test)
Why does adding a 5th consumer to a 4-partition topic do nothing? What's the general rule?
When do you create a new consumer group vs add a consumer to an existing one?
List the three events that trigger a rebalance.
Contrast eager and cooperative rebalance phase-by-phase. Which partitions keep flowing in each?
There are two independent mechanisms for detecting a dead consumer. Name both, and explain the specific failure each one catches that the other cannot.
Why does a clean `close()` reduce downtime compared to a crash?
Who computes partition assignments — a broker or a client? What does each consumer get to see?
What does `group.instance.id` change, and what new risk does it introduce?
Why is tuning `session.timeout.ms` harder under static membership?
Why is the `poll()` line called "the most important line in the chapter"?
Name three things `poll()` does besides fetching records. Why does that make it the source of most consumer exceptions?
Producers are thread-safe; consumers are not. What are the two supported concurrency patterns?
What changed between `poll(long)` and `poll(Duration)`, and what's the standard replacement for the `poll(0)` hack?
Explain `fetch.min.bytes` + `fetch.max.wait.ms` as a pair. What does each protect?
Why does the book prefer `fetch.max.bytes` over `max.partition.fetch.bytes`?
What's the relationship between `max.poll.records` and `max.poll.interval.ms`?
Why is lowering `request.timeout.ms` counterproductive during an incident?
Give the three `auto.offset.reset` values and the specific risk of each.
Draw the two offset failure modes (committed \< processed, committed > processed) and name the consequence of each.
Precisely which offset gets committed by default, and what's the manual-commit rule?
Autocommit + an exception mid-batch + `continue` = what bug? Trace it.
Why does `commitAsync()` deliberately not retry? Give the 2000/3000 trace.
Describe the sequence-number pattern for safe async commit retries.
Why is `commitAsync()` in the loop plus `commitSync()` on shutdown the standard shape?
You need to commit mid-batch. Why can't you just call `commitSync()`?
Which rebalance-listener method commits offsets, and why must it be *sync*?
What's the ordering guarantee across `onPartitionsRevoked()` and `onPartitionsAssigned()`, and why does it matter?
When is `onPartitionsLost()` called, what's the hazard, and what happens if you don't implement it?
What's the *only* consumer method safe to call from another thread? What exception does it produce, and where?
Two ways regex subscription can bite you on a large cluster.
`offsets.retention.minutes` is a broker config. Describe the consumer-visible failure it causes.
`subscribe()` vs `assign()`: list what you gain and what you give up. What silently stops working with `assign()`?
Your output database was wiped. How do you rebuild it from Kafka, and what bounds whether you can?
**Previous:** [Chapter 3 — Kafka Producers](03-kafka-producers.md)
**Next:** [Chapter 5 — Managing Kafka Programmatically](05-managing-kafka-programmatically.md)
# 4.12 Standalone consumer — `assign()` instead of `subscribe()` (/docs/kafka/kafka-consumers-reading-data/standalone-consumer-assign-instead)
**When:** *"Sometimes you know you have a **single consumer that always needs to read data from all the partitions in a topic, or from a specific partition** in a topic. In this case, **there is no reason for groups or rebalances** — just assign the consumer-specific topic and/or partitions, consume messages, and commit offsets on occasion."*
> ⚠️ *"although **you still need to configure `group.id` to commit offsets** — without calling `subscribe`, the consumer won't join any group."*
```java
Duration timeout = Duration.ofMillis(100);
List partitionInfos = null;
partitionInfos = consumer.partitionsFor("topic"); // ① ask the cluster
if (partitionInfos != null) {
for (PartitionInfo partition : partitionInfos)
partitions.add(new TopicPartition(partition.topic(),
partition.partition()));
consumer.assign(partitions); // ② assign, don't subscribe
while (true) {
ConsumerRecords records = consumer.poll(timeout);
for (ConsumerRecord record: records) {
/* process */
}
consumer.commitSync();
}
}
```
> ⚠️ **"Keep in mind that if someone adds new partitions to the topic, the consumer will NOT be notified.** You will need to handle this by **checking `consumer.partitionsFor()` periodically**, or simply by **bouncing the application whenever partitions are added.**"\*
**Trade-off summary:**
***
# 4.3 Static group membership — avoiding rebalance on restart (/docs/kafka/kafka-consumers-reading-data/static-group-membership-avoiding)
**Default behavior:** *"the identity of a consumer as a member of its consumer group is **transient**. When consumers leave a consumer group, the partitions that were assigned to the consumer are revoked, and when it rejoins, it is assigned a **new member ID** and a **new set of partitions** through the rebalance protocol."*
**With `group.instance.id` set to a unique string → the consumer becomes a *static* member:**
> If two consumers join the same group with the same `group.instance.id`, **the second gets an error saying a consumer with this ID already exists.**
#### When to use it — and the sharp trade-off [#when-to-use-it--and-the-sharp-trade-off]
**Use it when:** *"your application maintains **local state or cache** that is populated by the partitions assigned to each consumer. When **re-creating this cache is time-consuming**, you don't want this process to happen every time a consumer restarts."*
**⚠️ The cost — read this twice:**
> *"On the flip side, it is important to remember that **the partitions owned by each consumer will not get reassigned when a consumer is restarted.** For a certain duration, **no consumer will consume messages from these partitions**, and when the consumer finally starts back up, **it will lag behind the latest messages** in these partitions. **You should be confident that the consumer that owns these partitions will be able to catch up with the lag after the restart.**"*
**Tuning `session.timeout.ms` for static membership is a genuine dilemma:**
***
# 6.8 What actually breaks in production — Ch. 6 consolidated (/docs/kafka/kafka-internals/actually-breaks-production-ch)
| # | Symptom | Internals-level cause | Fix |
| -- | ------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------ |
| 1 | **Brokers drop out of the cluster with no restart** | A **long GC pause** or **network partition** looks identical to "stopped" — the ZK ephemeral node vanishes | G1GC tuning (Ch. 2); raise `zookeeper.session.timeout.ms`; dedicated ZK ensemble |
| 2 | **New broker won't start: node already exists** | Duplicate `broker.id` — the ephemeral node is already registered | Unique IDs |
| 3 | **Replaced a dead broker; it doesn't pick up its old partitions** | You gave it a **new** broker ID; replica lists still reference the old one | **Reuse the dead broker's ID** — assignments are inherited instantly |
| 4 | **Two brokers behaved as controller; contradictory commands** | Old controller resumed after a GC pause as a **zombie** | Nothing to fix — **controller epoch fencing** already discards its messages. Understand it so you don't misdiagnose |
| 5 | **Controller failover takes many seconds; cluster-wide stall** | Controller must **read the full replica state map from ZooKeeper** before it can act; *"in clusters with large numbers of partitions... several seconds"* | Limit partition counts (Ch. 2's 14k/broker, 1M/cluster); ultimately, **KRaft** |
| 6 | **Metadata inconsistent between brokers, controller, and ZooKeeper** | ZK writes sync, broker pushes async, ZK reads async — *"edge cases... challenging to detect"* | KRaft; until then, avoid direct ZooKeeper access (Ch. 5) |
| 7 | **Client produced to a broker that was no longer the leader** | The broker was too out-of-date to know it lost leadership | KRaft's **fenced state**; today, rely on `NotLeaderForPartition` + metadata refresh |
| 8 | **Consumer latency higher than expected after enabling follower fetch** | The **high-water mark propagates to followers with a delay** — followers are always slightly behind | Expected trade-off: cheaper cross-AZ reads for slightly staler data |
| 9 | **A partition becomes unavailable on leader failure despite RF 3** | Followers were **out of sync** (>`replica.lag.time.max.ms`) so were **ineligible** for election | Fix replication throughput (NIC, Ch. 2); monitor under-replicated partitions; RF++ |
| 10 | **One broker holds most leadership; it's hot** | The **first replica in the list is the preferred leader**; a manual reassignment put the same broker first everywhere | *"make sure you spread those around different brokers"*; `auto.leader.rebalance.enable`; Cruise Control |
| 11 | **Producers get `NotLeaderForPartition` in bursts** | Leader election happened; the client's **cached metadata is stale** | Normal and retriable — the client refreshes and retries; check `metadata.max.age.ms` if persistent |
| 12 | **Acknowledged data lost after a correlated power failure** | Kafka **writes to the filesystem cache and does NOT fsync** — durability comes from *replication*, not disk | RF ≥ 3 across **racks/AZs** so failures aren't correlated; `min.insync.replicas=2` |
| 13 | **`acks=all` produce latency spikes** | Request sits in **purgatory** until followers replicate | Fix replication speed; watch purgatory size + under-replicated partitions |
| 14 | **Consumers see an empty response even though the leader has data** | The data is **above the high-water mark** — not yet on all ISR, therefore "unsafe" | Expected. If persistent, replication is slow — see #9 |
| 15 | **New messages take unusually long to reach consumers** | Visibility waits for ISR replication; bounded by `replica.lag.time.max.ms` | Same root cause as #9 |
| 16 | **Broker CPU high and zero-copy seemingly not working** | **Message format down-conversion** for old consumers | Check **`FetchMessageConversionsPerSec`** / **`MessageConversionsTimeMs`** (KIP-188); **upgrade clients** |
| 17 | **Clients break after a Kafka upgrade** | You upgraded **clients before brokers**; old brokers can't parse newer request versions | **Always upgrade brokers first** — *"new brokers know how to handle old requests, but not vice versa"* |
| 18 | **Fetch overhead high on consumers with many partitions** | **Fetch session** not created or **evicted** (limited cache space; followers and large-partition consumers are prioritized) → fell back to **full fetch requests** | Expected degradation; reduce partitions per consumer, or accept it. Monitor |
| 19 | **A partition can't grow past a certain size** | *"Partitions cannot be split between multiple brokers, and not even between multiple disks"* — bounded by **one mount point** | More partitions; bigger mounts; eventually **tiered storage** |
| 20 | **One disk fills while others sit empty** | Directory placement counts **partitions, not bytes** — and *"if you add a new disk, ALL new partitions will be created on that disk"* | Monitor per-mount usage; equal-size disks; manual reassignment |
| 21 | **Brokers with more disk space get no more data** | **Broker** allocation ignores available space and existing load entirely | Don't mix heterogeneous hardware casually; use a balancer |
| 22 | **A backfill/historical read destroys everyone's latency** | Old reads **evict the hot page cache** and compete for disk I/O (21 ms → 60 ms p99 in KIP-405's measurement) | **Tiered storage** (network path, leaves page cache intact); until then, isolate historical consumers |
| 23 | **"7-day retention" keeps far more** | **The active segment is never deleted**, and only closed segments are eligible | Size segments so retention ÷ roll ≈ several segments; `log.roll.ms` for low-volume topics |
| 24 | **"Too many open files"** | Broker keeps **an open handle to every segment of every partition, including inactive ones** | Raise ulimits / `vm.max_map_count`; don't over-shrink segments |
| 25 | **Rebalancing/expanding the cluster is glacially slow** | Move time is driven by **partition size** — *"large partitions make the cluster less elastic"* | Smaller partitions; tiered storage; throttle + plan |
| 26 | **Compaction stops working; error in the logs** | **Not even one full segment fits** in the per-thread offset map (total memory ÷ thread count) | Allocate more offset-map memory **or use FEWER cleaner threads** |
| 27 | **Compacted topic keeps growing** | Compaction only touches **inactive** segments, and only fires at \~**50% dirty** | Tune the dirty ratio; check cleaner threads are alive; consider `delete.and.compact` |
| 28 | **Compaction fails outright** | The topic contains **null keys** | Compaction requires a key on every record |
| 29 | **GDPR deletion didn't reach the downstream database** | The consumer was **offline while the tombstone came and went**; on restart the key just doesn't exist, so no delete event is ever seen | **Tombstone retention > worst-case consumer downtime**; alert on long consumer outages |
| 30 | **A delete request wasn't honored within the legal window** | No **`max.compaction.lag.ms`** — the tombstone sat in a segment that wasn't eligible | Set `max.compaction.lag.ms` below your legal deadline (e.g. GDPR 30 days) |
| 31 | **Records "deleted" via `deleteRecords` but consumers weren't notified** | `deleteRecords` **moves the low-water mark** — it emits **no event** | Use **tombstones** when downstream systems must learn about deletions |
| 32 | **Suspected index corruption / weird offset lookups** | Index files have **no checksums** | **Delete the index segments** — they regenerate automatically (cost: recovery time) |
| 33 | **Poor compression ratio and high per-message overhead** | Batches of one; per-batch header amortized over a single record; deltas useless | `linger.ms` > 0; write to **fewer partitions per producer** (sticky partitioner) |
***
# 6.1 Cluster membership — ZooKeeper ephemeral nodes (/docs/kafka/kafka-internals/cluster-membership-zookeeper-ephemeral)
**Mechanics:**
* Every broker has a unique ID — *"either set in the broker configuration file or automatically generated."*
* **Every time a broker process starts, it registers itself with its ID in ZooKeeper by creating an ephemeral node.**
* Start another broker with the same ID → **error**, because *"we already have a ZooKeeper node for the same broker ID."*
**How a broker "disappears":**
> *"When a broker loses connectivity to ZooKeeper (**usually as a result of the broker stopping, but this can also happen as a result of network partition or a long garbage-collection pause**), the ephemeral node... **will be automatically removed** from ZooKeeper. Kafka components watching the list of brokers will be notified that the broker is gone."*
Note the three causes lumped together: **stopped**, **network partition**, **long GC pause**. From ZooKeeper's perspective they are indistinguishable — which is exactly why Ch. 2's GC tuning is a *cluster stability* concern, not just a latency concern.
#### 💡 Broker IDs outlive broker nodes — and this is a feature [#-broker-ids-outlive-broker-nodes--and-this-is-a-feature]
> *"Even though the node representing the broker is gone when the broker is stopped, **the broker ID still exists in other data structures.** For example, **the list of replicas of each topic contains the broker IDs for the replica.** This way, **if you completely lose a broker and start a brand-new broker with the ID of the old one, it will immediately join the cluster in place of the missing broker with the same partitions and topics assigned to it.**"*
This is the recovery pattern for a dead broker: **reuse the ID.** It also explains Ch. 2's advice to derive `broker.id` from the hostname — ID↔host mapping is load-bearing during exactly this operation.
***
# 6.7 Compaction (/docs/kafka/kafka-internals/compaction)
#### 7.1 Why it exists — two use cases [#71-why-it-exists--two-use-cases]
#### 7.2 The three retention policies [#72-the-three-retention-policies]
| Policy | Behavior |
| ------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`delete`** | *"deletes events older than retention time"* |
| **`compact`** | *"only stores the **most recent value for each key** in the topic"* |
| **`delete.and.compact`** | *"combines compaction with a retention period. **Messages older than the retention period will be removed EVEN IF THEY ARE THE MOST RECENT VALUE FOR A KEY.**"* Purpose: *"prevents compacted topics from growing overly large and is also used when the business requires removing records after a certain time period."* |
> ⚠️ *"setting the policy to `compact` **only makes sense on topics for which applications produce events that contain BOTH a key and a value. If the topic contains NULL KEYS, compaction will FAIL.**"*
#### 7.3 How compaction works [#73-how-compaction-works]
**Clean and dirty:**
**Thread setup:** if compaction is enabled at startup (*"using the awkwardly named `log.cleaner.enabled` configuration"*), *"each broker will start a **compaction manager thread** and a number of **compaction threads**."*
**Partition selection:** *"Each thread chooses the partition with the **highest ratio of dirty messages to total partition size**."*
**The offset map — and its beautiful efficiency:**
**Memory configuration and the failure mode:**
> *"the administrator configures **how much memory compaction threads can use** for this offset map. **Even though each thread has its own map, the configuration is for TOTAL memory across all threads.** If you configured 1 GB and you have 5 cleaner threads, **each thread will get 200 MB.**"*
>
> ⚠️ *"Kafka **doesn't require the entire dirty section to fit** into the size allocated for this map, **but AT LEAST ONE FULL SEGMENT HAS TO FIT. If it doesn't, Kafka will log an error, and the administrator will need to either allocate more memory for the offset maps OR USE FEWER CLEANER THREADS.** If only a few segments fit, Kafka will start by compacting the oldest segments that fit into the map. The rest will remain dirty and wait for the next compaction."*
**The counterintuitive fix:** if compaction fails for lack of memory, **reducing the thread count may help** — because total memory is divided among threads. More threads = less memory each.
**The copy phase:**
The swap is what makes compaction crash-safe: the original segment is intact until the replacement is complete.
#### 7.4 Deleted events — tombstones [#74-deleted-events--tombstones]
**The question:** *"If we always keep the latest message for each key, what do we do when we really want to delete ALL messages for a specific key, such as if a user left our service and **we are legally obligated to remove all traces of that user**?"*
#### ⚠️ Why the tombstone retention window matters so much [#️-why-the-tombstone-retention-window-matters-so-much]
> *"It is important to give consumers enough time to see the tombstone message, because **if our consumer was down for a few hours and MISSED the tombstone message, it will simply NOT SEE THE KEY when consuming and therefore NOT KNOW that it was deleted from Kafka or that it needs to be deleted from the database.**"*
**This is a compliance bug disguised as a config default.** Tombstone retention must exceed your worst-case consumer downtime.
#### 7.5 `deleteRecords` — a completely different mechanism [#75-deleterecords--a-completely-different-mechanism]
> *"Kafka's admin client also includes a **`deleteRecords`** method. This method deletes all records before a specified offset, and **it uses a COMPLETELY DIFFERENT MECHANISM.** When called, **Kafka will move the LOW-WATER MARK — its record of the first offset of a partition — to the specified offset.** This will **prevent consumers from consuming the records below the new low-water mark and effectively makes these records inaccessible until they get deleted by a cleaner thread.** This method can be used **on topics with a retention policy AND on compacted topics.**"*
**Practical guidance:** "delete this user everywhere" → **tombstone** (downstream systems learn about it). "Delete everything older than 30 days" → **`deleteRecords`** (Ch. 5 §7.3).
#### 7.6 When are topics compacted? [#76-when-are-topics-compacted]
**The default trigger and its rationale:**
> *"By default, Kafka will start compacting when **50% of the topic contains dirty records.** The goal is **not to compact too often (since compaction can impact the read/write performance on a topic)** but also **not to leave too many dirty records around (since they consume disk space). Wasting 50% of the disk space used by a topic on dirty records and then compacting them in one go seems like a reasonable trade-off**, and it can be tuned."*
**Two timing controls:**
| Config | Guarantees |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------- |
| **`min.compaction.lag.ms`** | *"the **MINIMUM** length of time that must pass after a message is written before it **could be** compacted"* |
| **`max.compaction.lag.ms`** | *"the **MAXIMUM** delay between the time a message is written and the time the message **becomes eligible** for compaction"* |
> **The compliance use case, named explicitly:** *"`max.compaction.lag.ms` is often used in situations where there is a business reason to guarantee compaction within a certain period; for example, **GDPR requires that certain information will be deleted within 30 days after a request to delete has been made.**"*
***
# 6.9 Deploy / monitor / scale / backup — through the internals lens (/docs/kafka/kafka-internals/deploy-monitor-scale-backup)
#### The metrics this chapter explains [#the-metrics-this-chapter-explains]
#### Tuning knobs and the mechanism each controls [#tuning-knobs-and-the-mechanism-each-controls]
| Knob | The mechanism it actually touches |
| -------------------------------------------------------- | --------------------------------------------------------------------------- |
| `num.network.threads` | **Processor threads** in §5.1 — moving requests to/from queues |
| `num.io.threads` | **Request handler threads** — actual request processing |
| `replica.lag.time.max.ms` | ISR membership **and** the consumer visibility delay ceiling |
| `min.insync.replicas` | The **third produce-request validation** (§5.3) |
| `acks=all` | Whether the response waits in **purgatory** |
| `linger.ms` | **Batch fill** → per-record overhead + compression ratio (§6.5) |
| `log.segment.bytes` / `log.roll.ms` | Segment count → retention precision, file handles, timestamp-seek precision |
| `log.cleaner.enabled` + offset-map memory + thread count | Whether compaction can run at all |
| `min/max.compaction.lag.ms` | Compaction timing guarantees (compliance) |
| `broker.rack` | The **rack-alternating broker list** in partition allocation (§6.3) |
| `client.rack` + `replica.selector.class` | Follower fetching, at the cost of HW-propagation delay |
| `metadata.max.age.ms` | Client metadata cache freshness → `NotLeaderForPartition` frequency |
#### "Backup" — the internals answer [#backup--the-internals-answer]
Ch. 6 finally makes precise what Kafka's durability actually *is*:
***
# 6. Kafka Internals (/docs/kafka/kafka-internals)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 6*
> **The book's own framing of why this chapter exists:** *"It is not strictly necessary to understand Kafka's internals in order to run Kafka in production... However, knowing how Kafka works does provide context when troubleshooting... **Understanding these topics in-depth will be especially useful when tuning Kafka — understanding the mechanisms that the tuning knobs control goes a long way toward using them with precise intent rather than fiddling with them randomly.**"*
**The four topics:** the **controller**, **replication**, **request processing**, and **storage** (file format + indexes + compaction).
***
# 6.3 KRaft — the Raft-based controller (/docs/kafka/kafka-internals/kraft-raft-based-controller)
**Timeline:** work started **2019**; **preview in Kafka 2.8**; *"The Apache Kafka 3.0 release, planned for mid 2021, will include the first production version of KRaft, and Kafka clusters will be able to run with either the traditional ZooKeeper-based controller or KRaft."*
#### 3.1 The four motivating problems — each is a real production pain [#31-the-four-motivating-problems--each-is-a-real-production-pain]
> *"Kafka's existing controller already underwent several rewrites, but... **it became clear that the existing model will not scale to the number of partitions we want Kafka to support.**"*
#### 3.2 What has to be replaced [#32-what-has-to-be-replaced]
#### 3.3 The core idea — eat your own dog food [#33-the-core-idea--eat-your-own-dog-food]
> *"The core idea behind the new controller design is that **Kafka itself has a log-based architecture, where users represent state as a stream of events.** The benefits are well understood: **multiple consumers can quickly catch up to the latest state by replaying events. The log establishes a clear ordering between events and ensures that the consumers always move along a single timeline.** The new controller architecture brings the same benefits to the management of Kafka's metadata."*
**Key changes, one by one:**
| Old | New |
| -------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| ZooKeeper elects the controller | *"Using the Raft algorithm, the controller nodes will **elect a leader from among themselves, without relying on any external system.**"* |
| Controller failover = lengthy full metadata reload | *"**Because the controllers will now all track the latest state, controller failover will NOT require a lengthy reloading period** in which we transfer all the state to the new controller."* |
| Controller **pushes** updates to brokers | *"brokers will **FETCH** updates from the active controller via a new **`MetadataFetch` API**"* — offset-tracked, incremental, just like a normal fetch |
| Brokers rebuild metadata at startup | *"Brokers will **persist the metadata to disk**"* → fast startup at millions of partitions |
| Broker registration is ephemeral | *"Brokers will register with the controller quorum and will **remain registered until unregistered by an admin**, so **once a broker shuts down, it is offline but still registered.**"* |
| A stale broker can serve requests | **NEW FENCED STATE** — see below |
#### 3.4 The new fenced state — closing a real correctness hole [#34-the-new-fenced-state--closing-a-real-correctness-hole]
> *"Brokers that are **online but are not up-to-date with the latest metadata will be FENCED and will not be able to serve client requests.** The new fenced state will **prevent cases where a client produces events to a broker that is no longer a leader but is too out-of-date to be aware that it isn't a leader.**"*
That is the *broker-level* analogue of controller zombie fencing: a lagging broker can currently accept writes it has no right to accept, because it doesn't yet know it lost leadership.
#### 3.5 Migration strategy [#35-migration-strategy]
> *"As part of the migration to the controller quorum, **all operations that previously involved either clients or brokers communicating directly to ZooKeeper will be routed via the controller.** This will allow **seamless migration by replacing the controller without having to change anything on any broker.**"*
**Reference KIPs:**
| KIP | Content |
| ----------- | -------------------------------------------------------------------------------- |
| **KIP-500** | Overall design of the new architecture |
| **KIP-595** | How the Raft protocol was adapted for Kafka |
| **KIP-631** | Controller quorum design, controller configuration, new CLI for cluster metadata |
***
# 6.6 Physical storage (/docs/kafka/kafka-internals/physical-storage)
#### 6.1 The fundamental storage constraint [#61-the-fundamental-storage-constraint]
> *"The basic storage unit of Kafka is **a partition replica**. **Partitions cannot be split between multiple brokers, and not even between multiple disks on the same broker.** So **the size of a partition is limited by the space available on a single mount point.**"*
This single sentence explains why partition count is a *capacity* decision (Ch. 2's "≤6 GB per day of retention" heuristic) and why tiered storage (§6.2) is such a big deal.
> Also: `log.dirs` is *"not to be confused with the location in which Kafka stores its error log, which is configured in the `log4j.properties` file."* *"The usual configuration includes a directory for each mount point that Kafka will use."*
#### 6.2 Tiered storage (planned for 3.0) [#62-tiered-storage-planned-for-30]
**Work started late 2018.** Three motivating concerns:
**The two-tier design:**
**What it buys:**
> *"allows **scaling storage independent of memory and CPUs**... enables Kafka to be a **long-term storage solution**... reduces the amount of data stored locally, and hence **the amount of data that needs to be copied during recovery and rebalancing**... **increasing the retention period no longer requires scaling the Kafka cluster storage and adding new nodes**... **eliminating the need for separate data pipelines to copy the data from Kafka to external stores**, as done currently in many deployments."*
#### 💡 The counterintuitive performance result (from KIP-405) [#-the-counterintuitive-performance-result-from-kip-405]
> *"in addition to infinite storage, lower costs, and elasticity, tiered storage also delivers **ISOLATION BETWEEN HISTORICAL READS AND REAL-TIME READS.**"*
**This is the real prize.** The classic Kafka production incident is *"someone started a backfill and now everyone's latency is terrible"* — because the historical reader evicts the hot page cache and competes for disk I/O. Tiered storage moves that traffic to a different resource entirely.
**Design detail:** KIP-405 introduces a new component, the **`RemoteLogManager`**, and specifies its interactions with replicas catching up to the leader, and leader elections.
#### 6.3 Partition allocation [#63-partition-allocation]
**Example: 6 brokers, 10 partitions, RF 3 → 30 replicas to place.**
**Three goals:**
**The algorithm, without rack awareness:**
**With rack awareness — a rack-alternating broker list:**
**Then: which directory?**
> *"We do this **independently for each partition**, and the rule is very simple: **we count the number of partitions on each directory and add the new partition to the directory with the FEWEST PARTITIONS.** This means that **if you add a new disk, ALL the new partitions will be created on that disk** — because, until things balance out, the new disk will always have the fewest partitions."*
> ### ⚠️ MIND THE DISK SPACE [#️-mind-the-disk-space]
>
> *"the allocation of **partitions to brokers does NOT take available space or existing load into account**, and the allocation of **partitions to disks takes the NUMBER of partitions into account but NOT THE SIZE of the partitions.** This means that if:*
>
> * *some brokers have **more disk space than others** (perhaps because the cluster is a mix of older and newer servers),*
> * *some **partitions are abnormally large**, or*
> * *you have **disks of different sizes on the same broker**,*
>
> *you need to be careful with the partition allocation."*
This is the internals-level explanation of Ch. 2's "one disk fills while others are empty" failure. **Kafka's placement is count-based and space-blind, at both levels.**
#### 6.4 File management — segments [#64-file-management--segments]
> *"Because **finding the messages that need purging in a large file and then deleting a portion of the file is both time-consuming and error prone**, we instead split each partition into **segments**."*
**Default segment: 1 GB of data OR a week of data, whichever is smaller.**
#### ⚠️ The active-segment retention trap (restated with numbers) [#️-the-active-segment-retention-trap-restated-with-numbers]
> *"**The active segment is never deleted**, so if you set log retention to only store **a day** of data, but each segment contains **five days** of data, **you will really keep data for FIVE DAYS** because we can't delete the data before the segment is closed."*
>
> *"If you choose to store data for **a week** and **roll a new segment every day**, you will see that every day we roll a new segment while deleting the oldest — so **most of the time the partition will have seven segments.**"*
That second sentence is the *design target*: **retention period ÷ segment roll interval ≈ number of segments.** Aim for a handful of segments per retention window, not one giant one.
**And the file-handle consequence (from Ch. 2):**
> *"a Kafka broker will keep an **open file handle to EVERY segment in EVERY partition — even inactive segments.** This leads to an unusually high number of open file handles, and **the OS must be tuned accordingly.**"*
#### 6.5 File format — and why it's identical to the wire format [#65-file-format--and-why-its-identical-to-the-wire-format]
> *"Each segment is stored in a single data file. Inside the file, we store Kafka messages and their offsets. **The format of the data on disk is IDENTICAL to the format of the messages that we send from the producer to the broker and later from the broker to the consumers.**"*
**What that buys — two things:**
**And what it costs:**
> *"if we decide to change the message format, **BOTH the wire protocol AND the on-disk format need to change**, and Kafka brokers need to know how to handle cases in which **files contain messages of two formats due to upgrades.**"*
#### Batching is mandatory since v0.11 / message format v2 [#batching-is-mandatory-since-v011--message-format-v2]
> *"Starting with version 0.11 (and the **v2 message format**), **Kafka producers ALWAYS send messages in batches.** If you send a single message, **the batching adds a bit of overhead. But with two messages or more per batch, the batching SAVES space**, which reduces network and disk usage."*
**Three consequences the book draws out:**
#### Batch header contents (v2) [#batch-header-contents-v2]
| Field | Notes |
| ----------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Magic number** | current message format version |
| **Offset of the first message + the difference to the last message's offset** | *"**preserved even if the batch is later compacted and some messages are removed.**"* First offset is set to **0** by the producer; **the partition leader that first persists the batch replaces it with the real offset** |
| **Timestamp of the first message + the highest timestamp in the batch** | *"can be set by the broker if the timestamp type is set to **append time** rather than **create time**"* |
| **Size of the batch, in bytes** | |
| **The epoch of the leader that received the batch** | *"used when truncating messages after leader election"* — **KIP-101, KIP-279** |
| **Checksum** | validating the batch is not corrupted |
| **16 bits of attributes** | **compression type**, **timestamp type**, whether the batch is **part of a transaction** or is a **control batch** |
| **Producer ID, producer epoch, first sequence in the batch** | *"all used for **exactly-once guarantees**"* (Ch. 8) |
| **The set of messages** | |
Two things worth noticing: **offsets are assigned by the leader, not the producer** (which is why a producer can't know its offset until the ack returns), and the **leader epoch is stored per batch** specifically so post-election truncation can be done correctly.
#### Record contents — and why per-record overhead is tiny [#record-contents--and-why-per-record-overhead-is-tiny]
| Field | Notes |
| ------------------------------------------------------------------------------------------ | -------------------------------------------------------------------------- |
| Size of the record, in bytes | |
| Attributes | *"currently there are **no record-level attributes**, so this isn't used"* |
| **The DIFFERENCE between this record's offset and the batch's first offset** | delta encoding |
| **The DIFFERENCE, in ms, between this record's timestamp and the batch's first timestamp** | delta encoding |
| User payload: **key, value, headers** | |
> *"**there is very little overhead to each record, and most of the system information is at the BATCH level.** Storing the first offset and timestamp of the batch in the header and **only storing the DIFFERENCE in each record dramatically reduces the overhead of each record, making larger batches more efficient.**"*
#### Control batches [#control-batches]
> *"Kafka also has **control batches** — indicating transactional commits, for instance. Those are **handled by the consumer and NOT passed to the user application**, and currently they include a **version** and a **type indicator: 0 for an aborted transaction, 1 for a commit.**"*
*(These are why Ch. 1 warns that offsets are "not necessarily monotonically greater" — control batches consume offsets that consumers never surface.)*
#### Inspecting segments yourself [#inspecting-segments-yourself]
```bash
bin/kafka-run-class.sh kafka.tools.DumpLogSegments
```
> *"allows you to look at a partition segment in the filesystem and examine its contents."* With **`--deep-iteration`** *"it will show you information about messages compressed inside the wrapper messages."*
#### ⚠️ Message format down-conversion — a CPU landmine [#️-message-format-down-conversion--a-cpu-landmine]
> *"Since Kafka supports upgrading brokers before all the clients are upgraded, it had to support any combination of versions... **But there is a challenging situation when a NEW PRODUCER sends v2 messages to NEW BROKERS: the message is stored in v2 format, but an OLD CONSUMER that doesn't support v2 tries to read it. In this scenario, THE BROKER WILL NEED TO CONVERT the message from v2 to v1**, so the consumer can parse it. **This conversion uses FAR MORE CPU AND MEMORY than normal consumption, so it is best avoided.**"*
This is a genuinely nasty one: your brokers get slower and hotter, and the cause is *a client you forgot about*. Two metrics tell you instantly.
#### 6.6 Indexes [#66-indexes]
> *"Kafka allows consumers to start fetching messages from **any available offset**. This means that if a consumer asks for 1 MB messages starting at offset 100, **the broker must be able to quickly locate the message for offset 100 (which can be in ANY of the segments for the partition).**"*
*(The timestamp index is what powers `offsetsForTimes()` from Ch. 4 and `OffsetSpec.forTimestamp()` from Ch. 5.)*
**Indexes are disposable — a genuinely useful operational fact:**
> *"Indexes are also broken into segments, so we can delete old index entries when the messages are purged. **Kafka does NOT attempt to maintain checksums of the index. If the index becomes corrupted, IT WILL GET REGENERATED from the matching log segment** simply by rereading the messages and recording the offsets and locations. **It is also completely safe (albeit, it can cause a lengthy recovery) for an administrator to DELETE INDEX SEGMENTS if needed — they will be regenerated automatically.**"*
***
# 6.4 Replication (/docs/kafka/kafka-internals/replication)
> *"Replication is at the heart of Kafka's architecture. Indeed, Kafka is often described as **'a distributed, partitioned, replicated commit log service.'** Replication is critical because it is **the way Kafka guarantees availability and durability when individual nodes inevitably fail.**"*
Scale note: *"each broker typically stores **hundreds or even thousands of replicas** belonging to different topics and partitions."*
#### 4.1 Leader and follower replicas [#41-leader-and-follower-replicas]
| | Leader replica | Follower replica |
| ------------------- | ------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
| Count per partition | **Exactly one** | All the rest |
| Produce requests | **All produce requests go through the leader — to guarantee consistency** | — |
| Consume requests | Yes | *"Unless configured otherwise, followers don't serve client requests"* |
| Main job | Serve clients; track follower progress | *"replicate messages from the leader and stay up-to-date with the most recent messages the leader has"* |
| On leader crash | — | *"one of the follower replicas will be promoted to become the new leader"* |
#### 4.2 Read from follower (KIP-392) — and its hidden latency cost [#42-read-from-follower-kip-392--and-its-hidden-latency-cost]
**Goal:** *"decrease network traffic costs by allowing clients to consume from the **nearest in-sync replica** rather than from the lead replica."*
**How correctness is preserved — and why it costs latency:**
> *"The replication protocol was extended to guarantee that **only committed messages will be available when consuming from a follower replica**... To provide this guarantee, **all replicas need to know which messages were committed by the leader.** To achieve this, **the leader includes the current HIGH-WATER MARK (latest committed offset) in the data that it sends to the follower.**"*
>
> ⚠️ *"**The propagation of the high-water mark introduces a small delay, which means that data is available for consuming from the leader EARLIER than it is available on the follower.** It is important to remember this additional delay, since **it is tempting to attempt to decrease consumer latency by consuming from the leader replica.**"*
#### 4.3 The ISR mechanism — how "in sync" is actually determined [#43-the-isr-mechanism--how-in-sync-is-actually-determined]
**The elegant part:** followers use **exactly the same Fetch requests consumers use.**
> *"To stay in sync with the leader, the replicas send the leader **Fetch requests — the exact same type of requests that consumers send** in order to consume messages... **Those Fetch requests contain the offset of the message that the replica wants to receive next, and will always be in order.** This means that **the leader can know that a replica got all messages up to the last message that the replica fetched, and none of the messages that came after. By looking at the last offset requested by each replica, the leader can tell how far behind each replica is.**"*
**The out-of-sync rule — two conditions:**
> *"replicas that are consistently asking for the latest messages are called **in-sync replicas**. **Only in-sync replicas are eligible to be elected as partition leaders** in case the existing leader fails."*
**Why followers fall behind (the book's examples):** *"network congestion slows down replication"*, or *"a broker crashes and all replicas on that broker start falling behind until we start the broker and they can start replicating again."*
#### 4.4 Preferred leader [#44-preferred-leader]
> *"each partition has a **preferred leader — the replica that was the leader when the topic was originally created.** It is preferred because **when partitions are first created, the leaders are balanced among brokers.** As a result, we expect that **when the preferred leader is indeed the leader for all partitions in the cluster, load will be evenly balanced between brokers.**"*
Default `auto.leader.rebalance.enable=true` *"will check if the preferred leader replica is **not the current leader but is in sync**, and will trigger leader election to make the preferred leader the current leader."*
> ### FINDING THE PREFERRED LEADER [#finding-the-preferred-leader]
>
> \*"The best way to identify the current preferred leader is by looking at the list of replicas for a partition (see `kafka-topics.sh`). **THE FIRST REPLICA IN THE LIST IS ALWAYS THE PREFERRED LEADER.** This is true **no matter who is the current leader** and **even if the replicas were reassigned** to different brokers using the replica reassignment tool.
>
> **In fact, if you manually reassign replicas, it is important to remember that the replica you specify first will be the preferred replica, so make sure you spread those around different brokers to avoid overloading some brokers with leaders while other brokers are not handling their fair share of the work.**"\*
***
# 6.5 Request processing (/docs/kafka/kafka-internals/request-processing)
> *"**Most of what a Kafka broker does is process requests** sent to the partition leaders from clients, partition replicas, and the controller."*
**Protocol basics:**
* **Binary protocol over TCP.**
* *"**Clients always initiate connections and send requests**, and the broker processes the requests and responds."*
* **The ordering guarantee that makes Kafka a queue:** *"All requests sent to the broker from a specific client will be **processed in the order in which they were received** — this guarantee is what allows Kafka to behave as a message queue and provide ordering guarantees on the messages it stores."*
**Standard request header:**
*(Correlation ID is the underrated one: it's how you tie a client-side timeout to a specific broker-side log line.)*
#### 5.1 The threading model — memorize this diagram [#51-the-threading-model--memorize-this-diagram]
This is the map for every broker metric and thread-pool config you'll ever tune.
**Why purgatory exists:** *"At times, responses to clients have to be delayed — **consumers only receive responses when data is available**, and **admin clients receive a response to a `DeleteTopic` request after topic deletion is underway.** The delayed responses are held in a **purgatory** until they can be completed."*
*(Purgatory is also where `acks=all` produce requests wait — §5.3.)*
#### 5.2 Request routing — metadata requests and the "Not a Leader" loop [#52-request-routing--metadata-requests-and-the-not-a-leader-loop]
**The three most common client request types:**
| Type | Sent by | Contains |
| ----------- | ----------------------------------- | ------------------------------------------ |
| **Produce** | producers | messages the clients write |
| **Fetch** | **consumers AND follower replicas** | read requests |
| **Admin** | admin clients | metadata operations (create/delete topics) |
**The routing constraint:**
> *"**Both produce requests and fetch requests have to be sent to the LEADER replica of a partition.** If a broker receives a produce request for a specific partition and the leader for this partition is on a different broker, the client will get an error response of **'Not a Leader for Partition.'** The same error will occur if a fetch request... arrives at a broker that does not have the leader. **Kafka's clients are responsible for sending produce and fetch requests to the broker that contains the leader** for the relevant partition."*
**This is the machinery behind Ch. 3's "retriable error."** `NotLeaderForPartition` is retriable precisely because the client's recovery action is *refresh metadata and retry* — and the underlying condition (a leader election) resolves on its own.
#### 5.3 Produce request handling [#53-produce-request-handling]
**Validations the leader runs first:**
#### ⚠️ The durability model, stated bluntly [#️-the-durability-model-stated-bluntly]
> *"Then the broker will write the new messages to local disk. **On Linux, the messages are written to the FILESYSTEM CACHE, and there is NO GUARANTEE about when they will be written to disk. Kafka DOES NOT WAIT for the data to get persisted to disk — IT RELIES ON REPLICATION FOR MESSAGE DURABILITY.**"*
**Then the `acks` branch:**
*(So `acks=all` latency is literally "time spent in purgatory." That's the metric to watch.)*
#### 5.4 Fetch request handling [#54-fetch-request-handling]
**The request shape:**
> *"something like **'Please send me messages starting at offset 53 in partition 0 of topic Test and messages starting at offset 64 in partition 3 of topic Test.'**"*
**Why the upper limit is mandatory, not optional:**
> *"Clients also specify a **limit to how much data the broker can return for each partition.** The limit is important because **clients need to allocate memory that will hold the response.** **Without this limit, brokers could send back replies large enough to cause clients to run out of memory.**"*
**Validation:** *"does this offset even exist for this particular partition? If the client is asking for a message that is **so old it got deleted** from the partition, or **an offset that does not exist yet**, the broker will respond with an error."*
#### 💡 Zero-copy — why the on-disk format equals the wire format [#-zero-copy--why-the-on-disk-format-equals-the-wire-format]
> *"Kafka famously uses a **zero-copy** method to send the messages to the clients — this means that **Kafka sends messages from the file (or more likely, the Linux filesystem cache) DIRECTLY TO THE NETWORK CHANNEL without any intermediate buffers.** This is different than most databases where data is stored in a local cache before being sent to clients. **This technique removes the overhead of copying bytes and managing buffers in memory, and results in much improved performance.**"*
**The precondition, from §7.2:** the on-disk format is **identical** to the wire format. If the broker had to transform bytes (decompress/recompress, or **down-convert message versions** — §7.3), zero-copy is defeated. That single fact links compression choice, client version skew, and broker CPU.
#### 5.5 Lower boundary (`fetch.min.bytes`) — the mechanism [#55-lower-boundary-fetchminbytes--the-mechanism]
> *"clients can also set a **lower boundary** on the amount of data returned. Setting the lower boundary to 10K is the client's way of telling the broker, **'Only return results once you have at least 10K bytes to send me.'** This is a great way to **reduce CPU and network utilization when clients are reading from topics that are not seeing much traffic.**"*
Plus the timeout: *"If you didn't satisfy the minimum amount of data to send within x milliseconds, **just send what you got.**"* (= `fetch.max.wait.ms`, Ch. 4.)
#### 5.6 ⚠️ The high-water mark — consumers can't read everything on the leader [#56-️-the-high-water-mark--consumers-cant-read-everything-on-the-leader]
> *"It is interesting to note that **not all the data that exists on the leader of the partition is available for clients to read.** **Most clients can only read messages that were written to ALL IN-SYNC REPLICAS** (follower replicas, even though they are consumers, are **exempt** from this — otherwise replication would not work)... **until a message was written to all in-sync replicas, it will not be sent to consumers — attempts to fetch those messages will result in an EMPTY RESPONSE rather than an error.**"*
**The consistency argument — this is the reasoning that matters:**
> *"messages not replicated to enough replicas yet are considered **'unsafe'** — if the leader crashes and another replica takes its place, **these messages will no longer exist in Kafka.** If we allowed clients to read messages that only exist on the leader, we could see inconsistent behavior. **For example, if a consumer reads a message and the leader crashed and no other broker contained this message, the message is gone. No other consumer will be able to read this message, which can cause inconsistency with the consumer who did read it.**"*
> ⚠️ **The latency consequence:** *"if replication between brokers is slow for some reason, **it will take longer for new messages to arrive to consumers** (since we wait for the messages to replicate first). **This delay is limited to `replica.lag.time.max.ms`** — the amount of time a replica can be delayed while still being considered in sync."*
**This is the mechanism behind Ch. 3's biggest insight:** *end-to-end latency is identical for all `acks` settings*, because visibility is gated on ISR replication regardless of what the producer waited for. Now you know exactly why: **the high-water mark.**
**And it gives you a latency bound worth remembering:** worst-case added consume latency from slow replication ≈ `replica.lag.time.max.ms`. Beyond that, the slow replica drops out of the ISR and stops holding the HW back.
#### 5.7 Fetch session cache — incremental fetch requests [#57-fetch-session-cache--incremental-fetch-requests]
**The problem:**
> *"In some cases, a consumer consumes events from a large number of partitions. **Sending the list of all the partitions it is interested in to the broker with every request and having the broker send all its metadata back can be very inefficient** — the set of partitions rarely changes, their metadata rarely changes, and in many cases there isn't that much data to return."*
**The solution:**
Note the graceful degradation: eviction is not an outage, it's a silent efficiency loss. Which means it's the kind of thing you only find via metrics.
#### 5.8 The protocol surface and version negotiation [#58-the-protocol-surface-and-version-negotiation]
> *"The Kafka protocol currently handles **61 different request types**, and more will be added. **Consumers alone use 15 request types** to form groups, coordinate consumption, and allow developers to manage the consumer groups."*
**The same protocol is used broker-to-broker:** *"Those requests are internal and should not be used by clients. For example, when the controller announces that a partition has a new leader, it sends a `LeaderAndIsr` request to the new leader (so it will know to start accepting client requests) and to the followers (so they will know to follow the new leader)."*
**Protocol evolution — the ZooKeeper-removal trail:**
#### ⚠️ 5.9 Why you must upgrade brokers BEFORE clients [#️-59-why-you-must-upgrade-brokers-before-clients]
The book walks a concrete example (Metadata request v0 → v1, adding controller info between 0.9.0 and 0.10.0):
**Mitigation added in 0.10.0:** `ApiVersionRequest` — *"allows clients to ask the broker which versions of each request are supported and to use the correct version accordingly. Clients that use this new capability correctly will be able to talk to older brokers."*
**Future work:** **KIP-584** — APIs for clients to discover which *features* brokers support, and for brokers to gate features per version. *"at this time it seems likely to be part of version 3.0.0."*
***
# 6.10 Self-test (/docs/kafka/kafka-internals/self-test)
Name the three causes of a broker's ZK ephemeral node disappearing. Why does it matter that they're indistinguishable?
A broker dies permanently. What's the fastest way to restore its partition assignments, and why does it work?
Walk the controller election protocol. What ZooKeeper primitive guarantees a single controller?
Trace the zombie-controller scenario end to end. What mechanism fences it, and what ZooKeeper operation makes the epoch safe?
Why does controller *restart* get slower as a cluster grows? Which KRaft change fixes it?
List the four problems that motivated KRaft, and which one is a *human* problem rather than a technical one.
In KRaft, do brokers get pushed metadata or pull it? Name two benefits of that direction.
What is the "fenced" broker state, and which specific bug does it eliminate?
Followers use the same Fetch requests consumers use. What does the leader infer from a fetch at offset N, and why is that elegant?
State both conditions under which a replica is declared out of sync. What config controls them?
Which replica is the preferred leader, and how do you identify it from `kafka-topics.sh` output? Why does this matter when reassigning replicas manually?
Reading from followers saves money. What does it cost, and what mechanism causes that cost?
Draw the broker's threading model: acceptor, processor/network, request queue, I/O/handler, response queue, purgatory. What two situations put a response in purgatory?
Why must produce and fetch requests go to the leader? How does a client find out where that is, and what happens when it's wrong?
Name the three validations a leader runs on a produce request.
Complete the sentence and explain the implication: "Kafka does not wait for the data to get persisted to disk — it relies on \_\_\_ for message durability."
Explain zero-copy, and state the precondition in the storage format that makes it possible.
What is the high-water mark? Give the consistency argument for why consumers can't read past it, in terms of what a consumer would observe after a leader crash.
What is the worst-case extra consume latency caused by slow replication, and what bounds it?
How does the high-water mark explain Ch. 3's claim that end-to-end latency is identical for all `acks` values?
What is the fetch session cache for? What happens when it can't serve you, and how would you find out?
Why must you upgrade brokers before clients? Give the version-negotiation walkthrough.
What is message-format down-conversion, why is it expensive, and which two metrics reveal it?
What is the hard upper bound on a single partition's size, and why?
State the three goals of partition allocation, and how the rack-alternating list achieves the third.
Kafka picks a directory for a new partition by counting what? What does that mean the moment you add a fresh disk?
Why are partitions split into segments at all? Which segment can never be deleted, and what retention bug does that cause?
Why is the batch header large but the per-record header tiny? What does that imply about `linger.ms` and about partition fan-out per producer?
What are control batches, who consumes them, and what does their existence explain about offsets?
Are index files worth protecting? What's the recovery procedure if one is corrupt, and what does it cost?
Give the offset-map arithmetic for compaction (bytes per entry, and the 1 GB segment example). Why might *reducing* cleaner threads fix a compaction failure?
What is a tombstone, and what's the precise failure if a consumer is offline for its entire retention window?
Contrast tombstones with `deleteRecords`: mechanism, granularity, downstream visibility, and which topics each works on.
Which two configs together make a compacted topic GDPR-capable, and what does each guarantee?
Why does tiered storage *improve* p99 latency in the historical-read scenario when it slightly worsens it in the steady-state one?
**Previous:** [Chapter 5 — Managing Kafka Programmatically](05-managing-kafka-programmatically.md)
**Next:** [Chapter 7 — Reliable Data Delivery](07-reliable-data-delivery.md)
# 6.2 The controller (/docs/kafka/kafka-internals/the-controller)
#### 2.1 Election via a single ephemeral node [#21-election-via-a-single-ephemeral-node]
**On controller loss:**
> *"When the controller broker is stopped or loses connectivity to ZooKeeper, the ephemeral node will disappear. **This includes any scenario in which the ZooKeeper client used by the controller stops sending heartbeats to ZooKeeper for longer than `zookeeper.session.timeout.ms`.** When the ephemeral node disappears, other brokers... will be notified through the ZooKeeper watch... and will attempt to create the controller node themselves. **The first node to create the new controller becomes the next controller**, while the others receive 'node already exists' and re-create the watch on the new node."*
#### 2.2 ⚠️ The controller epoch — zombie fencing [#22-️-the-controller-epoch--zombie-fencing]
This is one of the most important distributed-systems ideas in Kafka.
> *"Each time a controller is elected, it receives a **new, higher controller epoch number** through a **ZooKeeper conditional increment operation.** The brokers know the current controller epoch, and **if they receive a message from a controller with an older number, they know to ignore it.**"*
**Why it's necessary — the zombie scenario:**
> *"The controller epoch in the message, which allows brokers to ignore messages from old controllers, is a form of **zombie fencing**."* And: *"The controller uses the epoch number to prevent a **'split brain'** scenario where two nodes believe each is the current controller."*
**The generalizable lesson:** you cannot prevent a paused process from waking up and acting stale. You can only make its stale actions **rejectable**. Monotonic epochs + conditional increment is how.
#### 2.3 Controller startup cost [#23-controller-startup-cost]
> *"When the controller first comes up, **it has to read the latest replica state map from ZooKeeper before it can start managing** the cluster metadata and performing leader elections. The loading process uses **async APIs, and pipelines the read requests to ZooKeeper to hide latencies.** But even so, **in clusters with large numbers of partitions, the loading process can take SEVERAL SECONDS.**"*
#### 2.4 What the controller does on broker failure [#24-what-the-controller-does-on-broker-failure]
**Broker *startup* is the mirror image, with one difference:**
> *"the main difference is that **all replicas in the broker start as FOLLOWERS and need to catch up to the leader before they are eligible to be elected as leaders themselves.**"*
**Summary from the book:** *"Kafka uses ZooKeeper's ephemeral node feature to elect a controller and to notify the controller when nodes join and leave the cluster. The controller is responsible for electing leaders among the partitions and replicas whenever it notices nodes join and leave. The controller uses the epoch number to prevent a 'split brain' scenario."*
***
# 3.13 What actually breaks in production — Ch. 3 consolidated (/docs/kafka/kafka-producers-writing-messages/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ----------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ | --------------------------------------------------------------------------------------------------------------- |
| 1 | **Messages silently vanish; no error, no log** | Fire-and-forget `send()` — errors at/after the broker are invisible | Always use `send(record, callback)` |
| 2 | **Producer throughput is \~1/RTT messages per second** | Synchronous `send().get()` copied from a tutorial | Async + callback |
| 3 | **Producer throughput collapses; the code "looks async"** | **Blocking operation inside the callback** — callbacks run on the producer's main thread | Hand off to another thread inside the callback |
| 4 | **Producer got a success ack; message is gone** | `acks=1` (or the pre-3.0 default) + leader crashed before replicating | `acks=all` **and** `min.insync.replicas≥2` (Ch. 2/7). Remember: **`acks=all` costs no end-to-end latency** |
| 5 | **Total data loss under broker restart, no signal** | `acks=0` | Never in a durability-sensitive path |
| 6 | **$100 withdrawn before it was deposited** | `retries>0` + `max.in.flight>1` → **retry reordering** | **`enable.idempotence=true`** (ordering up to 5 in-flight, no dupes) |
| 7 | **Duplicate records despite everything succeeding** | Broker wrote + replicated, then **crashed before acking**; producer retried to the new leader | `enable.idempotence=true` |
| 8 | **`ConfigException` on producer startup after enabling idempotence** | `max.in.flight>5`, or `retries=0`, or `acks≠all` | Satisfy all three preconditions |
| 9 | **Producer gives up during a broker failover** | `delivery.timeout.ms` shorter than actual leader-election/recovery time | **Measure recovery time**, set `delivery.timeout.ms` above it (book's example: 30s election → 120s timeout) |
| 10 | **Confusing timeout exceptions; can't tell where time went** | Synchronous send blocks across *both* intervals indistinguishably | Async + callback; reason about `max.block.ms` vs `delivery.timeout.ms` separately |
| 11 | **Producer startup fails with an inconsistent-timeout exception** | `delivery.timeout.ms` ≤ `linger.ms + request.timeout.ms` | Keep the inequality |
| 12 | **Broker rejects messages: "message too large"** | `max.request.size` > broker `message.max.bytes` | Make them match (see Ch. 2 — also `replica.fetch.max.bytes` and consumer fetch size) |
| 13 | **Poor compression ratio despite compression enabled** | `linger.ms=0` → batches of 1 → nothing to compress against | Raise `linger.ms`; compression *"is much better"* with real batches |
| 14 | **High memory use, no latency benefit** | `batch.size` set huge, expecting it to trigger waiting | `batch.size` is a **ceiling**, not a wait trigger. **Only `linger.ms` makes the producer wait** |
| 15 | **`send()` blocks, then throws `TimeoutException`** | `buffer.memory` exhausted — sending faster than the broker accepts (capacity **or quota**) | Capacity-plan; monitor; catch it at `send()` (**not** in the `Future`) |
| 16 | **Records expire before being sent** | Sat in a batch longer than `delivery.timeout.ms` while the buffer was backed up | same as above |
| 17 | **Producer errors on a keyed write while the cluster is "mostly fine"** | Keyed records hash over **ALL** partitions, not just available ones — an offline partition is fatal for its keys | Expected trade-off; ensure replication/availability (Ch. 7). Null keys avoid it but lose key routing |
| 18 | **One partition is 10× the others; a broker fills its disk** | **Hot key** (the "Banana" problem) | Custom partitioner isolating hot keys, or `UniformStickyPartitioner` if key-routing isn't needed |
| 19 | **After adding partitions, a key's history splits; ordering and state break** | `hash(key) % N` changed when N changed | **Create with sufficient partitions and never add them** when key partitioning matters |
| 20 | **Cross-team byte-format war; nobody can add a field** | Hand-written custom serializer; *"you need to compare arrays of raw bytes"* to debug; all teams must change code simultaneously | Avro/Protobuf/Thrift + Schema Registry |
| 21 | **Consumers crash after a producer schema change** | Incompatible schema evolution, or the deserializer lacks the *writer's* schema | Follow Avro compatibility rules; Schema Registry supplies the writer schema by ID |
| 22 | **Record size more than doubled** | Embedding the full schema in every record | Schema Registry — only the **schema ID** travels with the record |
| 23 | **`KafkaAvroSerializer` fails on your class** | *"The Avro serializer can only serialize Avro objects, not POJO"* | Generate classes (`avro-tools.jar` / Avro Maven plugin), or use `GenericRecord` + explicit schema |
| 24 | **Producer can't connect even though "the cluster is up"** | `bootstrap.servers` lists a single broker, and that broker is down | *"It is recommended to include at least two"* |
| 25 | **Security alert names an IP, nobody knows which service** | `client.id` unset or meaningless | Set a descriptive `client.id` — it drives logs, metrics, **and quotas** |
| 26 | **Client mysteriously slow; broker healthy** | **Quota throttling** — broker delaying responses and muting the channel | Check `produce-throttle-time-avg/max`, `fetch-throttle-time-*`; review quotas |
| 27 | **Quota change requires a full cluster restart** | Quotas set statically in `server.properties` | Use **dynamic** config (`kafka-configs` / AdminClient) |
| 28 | **Interceptor leaks threads / file handles** | `close()` not implemented | Clean up in `close()` |
| 29 | **Hot-key list change requires recompiling the partitioner** | Values hardcoded in `partition()` | Pass them through `configure()` — the book flags its own example for this |
***
# 3.4 Configuring producers — every important knob (/docs/kafka/kafka-producers-writing-messages/configuring-producers-every-important)
#### 4.1 `client.id` [#41-clientid]
A **logical identifier for the client and the application it's used in**. Any string. Used by brokers to identify messages from the client — **in logging, metrics, and quotas**.
The book's argument for taking this seriously is the most practical sentence in the chapter:
> Choosing a good client name is *"the difference between 'We are seeing a high rate of authentication failures from IP 104.27.155.134' and **'Looks like the Order Validation service is failing to authenticate — can you ask Laura to take a look?'**"*
#### 4.2 `acks` — the durability knob [#42-acks--the-durability-knob]
Controls **how many partition replicas must receive the record before the producer considers the write successful.**
**Default: leader-only received** (`acks=1`) — *"release 3.0 of Apache Kafka is expected to change this default"* (it changed to `acks=all`).
#### 4.3 💡 The single best insight in the chapter: producer latency ≠ end-to-end latency [#43--the-single-best-insight-in-the-chapter-producer-latency--end-to-end-latency]
> \*"You will see that with lower and less reliable `acks` configuration, the producer will be able to send records faster. This means that **you trade off reliability for producer latency.**
>
> **However, end-to-end latency is measured from the time a record was produced until it is available for consumers to read — and is IDENTICAL for all three options.**
>
> The reason is that, **in order to maintain consistency, Kafka will not allow consumers to read records until they are written to all in-sync replicas.** Therefore, if you care about end-to-end latency, rather than just the producer latency, \**there is no trade-off to make: you will get the same end-to-end latency if you choose the most reliable option."*
**Practical implication:** if your SLO is "event visible to consumers within X ms" — which is almost always the real SLO — then **`acks=all` is free**. Choosing `acks=1` for latency is usually optimizing a metric nobody cares about while accepting real data-loss risk.
***
# 3.2 Constructing a producer (/docs/kafka/kafka-producers-writing-messages/constructing-producer)
#### Three mandatory properties [#three-mandatory-properties]
| Property | Purpose | Notes |
| ----------------------- | ---------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **`bootstrap.servers`** | `host:port` pairs for the **initial connection** | **Doesn't need all brokers** — the producer gets more info after connecting. **Include at least two**, so if one broker is down the producer can still reach the cluster |
| **`key.serializer`** | Class name implementing `org.apache.kafka.common.serialization.Serializer`, used to turn the **key object** into bytes | **Required even if you only send values** — use type `Void` with `VoidSerializer` |
| **`value.serializer`** | Same, for the **value object** | |
**Why serializers exist at all:** *"Kafka brokers expect byte arrays as keys and values. However, the producer interface allows, using parameterized types, any Java object to be sent as a key and value. This makes for very readable code, but it also means that the producer has to know how to convert these objects to byte arrays."*
Built in: `ByteArraySerializer` (*"which doesn't do much"*), `StringSerializer`, `IntegerSerializer`, **and much more**.
```java
Properties kafkaProps = new Properties();
kafkaProps.put("bootstrap.servers", "broker1:9092,broker2:9092");
kafkaProps.put("key.serializer",
"org.apache.kafka.common.serialization.StringSerializer");
kafkaProps.put("value.serializer",
"org.apache.kafka.common.serialization.StringSerializer");
producer = new KafkaProducer(kafkaProps);
```
> *"With such a simple interface, it is clear that most of the control over producer behavior is done by setting the correct configuration properties."*
***
# 3.14 Deploy / monitor / scale — the producer's operational surface (/docs/kafka/kafka-producers-writing-messages/deploy-monitor-scale-producer)
#### Config recipes by requirement [#config-recipes-by-requirement]
#### Monitoring the producer [#monitoring-the-producer]
| Signal | Why |
| -------------------------------------------------- | --------------------------------------------------------- |
| **`produce-throttle-time-avg` / `-max`** | You are being quota-throttled (or request-time throttled) |
| **Buffer available bytes / buffer exhausted rate** | Precursor to `send()` blocking and `TimeoutException` |
| **Record error rate / retry rate** | Retriable errors churning; possible leader instability |
| **Batch size avg, records per request** | Whether `linger.ms`/`batch.size` are actually batching |
| **Compression rate** | Whether compression is achieving anything (see #13) |
| **Request latency avg/max** | vs `request.timeout.ms` headroom |
| **Callback latency** (your own metric) | Guards against the blocking-callback trap (#3) |
| **`client.id` on every metric** | The whole point of naming clients well |
#### Scaling the producer [#scaling-the-producer]
#### "Backup" from the producer's perspective [#backup-from-the-producers-perspective]
There is no producer-side backup, but there *is* a producer-side **durability boundary**, and knowing where it sits is what matters:
**Practical consequence:** anything in the accumulator is unreplicated, unacked, and gone if the JVM dies. If you cannot tolerate that, you need either a durable local outbox before the producer, or `send().get()` at the boundary (accepting the throughput cost) — and the "errors file for later analysis" pattern the book mentions for the callback failure path.
***
# 3.10 Headers (/docs/kafka/kafka-producers-writing-messages/headers)
> *"Record headers give you the ability to add some metadata about the Kafka record, **without adding any extra information to the key/value pair of the record itself.**"*
**Uses:**
* **Lineage** — indicating the source of the data in the record
* **Routing or tracing messages based on header information without having to parse the message itself** — *"perhaps the message is encrypted and the router doesn't have permissions to access the data."*
That last clause is the compelling one: headers let an intermediary route/trace **data it is not authorized to read**.
**Implementation:** *"an **ordered collection of key/value pairs**. The **keys are always a String**, and the **values can be any serialized object** — just like the message value."*
```java
ProducerRecord record =
new ProducerRecord<>("CustomerCountry", "Precision Products", "France");
record.headers().add("privacy-level",
"YOLO".getBytes(StandardCharsets.UTF_8));
```
***
# 3.1 How does it work internally? Producer architecture (/docs/kafka/kafka-producers-writing-messages/how-does-work-internally)
#### The sequence in words [#the-sequence-in-words]
1. Create a **`ProducerRecord`** — must include **topic** and **value**; optionally **key, partition, timestamp, headers**.
2. Producer **serializes key and value objects to byte arrays** so they can go over the network.
3. If **no partition was explicitly specified**, the data goes to a **partitioner**, which chooses a partition, **usually based on the key**.
4. Now the producer knows the topic *and* partition → it **adds the record to a batch of records bound for that same topic and partition**.
5. **A separate thread** sends those batches to the appropriate brokers.
6. The broker responds:
* **Success** → `RecordMetadata` with **topic, partition, and the offset of the record within the partition**.
* **Failure** → an error. On a retriable error, the producer **may retry a few more times before giving up and returning an error**.
**Two facts that surprise people:**
* **A single producer object is thread-safe and can be used by multiple threads.** (All chapter examples are single-threaded, but this is explicit.)
* **`send()` is *always* asynchronous internally.** "Synchronous send" is just you blocking on the returned `Future`.
***
# 3. Kafka Producers: Writing Messages to Kafka (/docs/kafka/kafka-producers-writing-messages)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 3*
> **Learning goal:** the producer API is \~10 lines. Everything that matters is (a) what happens on the internal threads, and (b) how \~15 configs interact to define *durability*, *ordering*, *latency*, and *throughput*. Most production Kafka bugs are producer misconfigurations.
***
# 3.11 Interceptors (/docs/kafka/kafka-producers-writing-messages/interceptors)
**The problem they solve:**
> *"There are times when you want to **modify the behavior of your Kafka client application without modifying its code**, perhaps because you want to add identical behavior to all applications in the organization. Or perhaps **you don't have access to the original code.**"*
`ProducerInterceptor` has two key methods:
| Method | When | Can it modify? |
| -------------------------------------------------------------------------- | ---------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`ProducerRecord onSend(ProducerRecord record)`** | **Before the record is sent to Kafka — indeed before it is even serialized** | **Yes.** *"you can capture information about the sent record and even modify it. Just be sure to return a valid `ProducerRecord` — the record this method returns will be serialized and sent to Kafka."* |
| **`void onAcknowledgement(RecordMetadata metadata, Exception exception)`** | **If and when Kafka responds with an acknowledgment** | **No.** *"does not allow modifying the response from Kafka, but you can capture information about the response."* |
**Common use cases:**
* capturing **monitoring and tracing** information
* **enhancing the message with standard headers, especially for lineage tracking**
* **redacting sensitive information**
#### Example: a counting interceptor [#example-a-counting-interceptor]
```java
public class CountingProducerInterceptor implements ProducerInterceptor {
ScheduledExecutorService executorService =
Executors.newSingleThreadScheduledExecutor();
static AtomicLong numSent = new AtomicLong(0);
static AtomicLong numAcked = new AtomicLong(0);
public void configure(Map map) {
Long windowSize = Long.valueOf(
(String) map.get("counting.interceptor.window.size.ms"));
executorService.scheduleAtFixedRate(CountingProducerInterceptor::run,
windowSize, windowSize, TimeUnit.MILLISECONDS);
}
public ProducerRecord onSend(ProducerRecord producerRecord) {
numSent.incrementAndGet();
return producerRecord; // unmodified
}
public void onAcknowledgement(RecordMetadata recordMetadata, Exception e) {
numAcked.incrementAndGet();
}
public void close() {
executorService.shutdownNow(); // clean up! avoid leaks
}
public static void run() {
System.out.println(numSent.getAndSet(0));
System.out.println(numAcked.getAndSet(0));
}
}
```
Notes from the book:
* `ProducerInterceptor` is a **`Configurable`** interface — override `configure` to set up **before any other method is called**. It **receives the entire producer configuration**, so you can read any parameter, including **your own custom ones** (here `counting.interceptor.window.size.ms`).
* `close()` is *"the place to close everything and avoid leaks"* — threads, **file handles, connections to remote data stores**.
#### Applying an interceptor with zero code changes [#applying-an-interceptor-with-zero-code-changes]
Using it with `kafka-console-producer` (which ships with Kafka):
```bash
# 1. add your jar to the classpath
export CLASSPATH=$CLASSPATH:~./target/CountProducerInterceptor-1.0-SNAPSHOT.jar
# 2. create a config file containing:
# interceptor.classes=com.shapira.examples.interceptors.CountProducerInterceptor
# counting.interceptor.window.size.ms=10000
# 3. run normally, including the config
bin/kafka-console-producer.sh --broker-list localhost:9092 \
--topic interceptor-test --producer.config producer.config
```
*(Note: `--broker-list` here is the older flag form; Ch. 2 covers the `--bootstrap-server` migration.)*
***
# 3.5 Message delivery time — the timeout configs and how they interact (/docs/kafka/kafka-producers-writing-messages/message-delivery-time-timeout)
> *"The producer has multiple configuration parameters that interact to control one of the behaviors that are of most interest to developers: **how long will it take until a call to `send()` will succeed or fail.**"*
>
> The configs *"were modified several times over the years"* — this describes **the latest implementation, introduced in Apache Kafka 2.1.**
#### The two intervals [#the-two-intervals]
> **Note:** *"If you use `send()` synchronously, the sending thread will block for **both** time intervals continuously, and **you won't be able to tell how much time was spent in each.**"*
#### Detailed timeline [#detailed-timeline]
#### The configs, one by one [#the-configs-one-by-one]
##### `max.block.ms` [#maxblockms]
How long the producer may block when calling **`send()`** and when explicitly requesting metadata via **`partitionsFor()`**. Those methods block when:
* **the producer's send buffer is full**, or
* **metadata is not available.**
When reached, **a timeout exception is thrown**.
##### `delivery.timeout.ms` — the one you should actually tune [#deliverytimeoutms--the-one-you-should-actually-tune]
Limits the time **from the point a record is ready for sending** (`send()` returned successfully and the record is in a batch) **until either the broker responds or the client gives up — including time spent on retries.**
* **Must be greater than `linger.ms` + `request.timeout.ms`.** *"If you try to create a producer with an inconsistent timeout configuration, you will get an exception."*
* *"Messages can be successfully sent much faster than `delivery.timeout.ms`, and typically will."*
**Which exception you get depends on where the clock ran out:**
| Timeout exceeded... | Callback receives |
| ------------------------------------------------------- | ---------------------------------------------------------------------------------------- |
| **while retrying** | *"the exception that corresponds to the error that the broker returned before retrying"* |
| **while the record batch was still waiting to be sent** | **a timeout exception** |
> ### 💡 TIP — the right mental model for retries [#-tip--the-right-mental-model-for-retries]
>
> \*"Configure the delivery timeout to the maximum time you'll want to wait for a message to be sent, typically **a few minutes**, and then **leave the default number of retries (virtually infinite).** With this configuration, the producer will keep retrying for as long as it has time to keep trying (or until it succeeds). **This is a much more reasonable way to think about retries.**
>
> Our normal process for tuning retries is: *'In case of a broker crash, it typically takes leader election **30 seconds** to complete, so let's keep retrying for **120 seconds** just to be on the safe side.'* Instead of converting this mental dialog to number of retries and time between retries, \**you just configure `delivery.timeout.ms` to 120\[s]."*
This reframing is the single most useful operational idea in the chapter: **you reason in *wall-clock recovery time*, not in retry counts.** Retry count × backoff is a derived quantity you should never hand-compute.
##### `request.timeout.ms` [#requesttimeoutms]
How long the producer waits for a reply from the server **when sending data**.
> *"Note that this is the time spent waiting on **each producer request** before giving up; it does **not** include retries, time spent before sending, and so on."*
If reached without a reply, the producer **either retries sending or completes the callback with a `TimeoutException`.**
##### `retries` and `retry.backoff.ms` [#retries-and-retrybackoffms]
`retries` = how many times to retry before giving up. Default wait between retries: **100 ms** (`retry.backoff.ms`).
> **"We recommend against using these parameters in the current version of Kafka.**
> Instead, **test how long it takes to recover from a crashed broker** (i.e., how long until all partitions get new leaders), and set `delivery.timeout.ms` such that the total time spent retrying will be **longer than the time it takes the cluster to recover from the crash** — otherwise, **the producer will give up too soon.**"
> *"Because the producer handles retries for you, **there is no point in handling retries within your own application logic.** You will want to focus your efforts on handling **non-retriable errors** or cases where **retry attempts were exhausted.**"*
> **TIP:** *"If you want to completely disable retries, setting `retries=0` is the **only** way to do so."*
##### `linger.ms` [#lingerms]
Time to wait for additional messages before sending the current batch. `KafkaProducer` sends a batch **either when the batch is full, or when `linger.ms` is reached.**
**Default behavior (linger.ms = 0):** *"the producer will send messages as soon as there is a sender thread available to send them, **even if there's just one message in the batch.**"*
Setting it **> 0**:
> *"increases latency a little and **significantly increases throughput** — the overhead per message is much lower, and **compression, if enabled, is much better.**"*
The compression point is the underrated one: compression ratio is a function of *how much similar data is in the same block*. `linger.ms=0` largely defeats compression.
##### `buffer.memory` [#buffermemory]
Memory the producer uses to **buffer messages waiting to be sent**.
> If the application sends faster than messages can be delivered, *"the producer may run out of space, and **additional `send()` calls will block for `max.block.ms`** and wait for space to free up before throwing an exception."*
>
> **"Note that unlike most producer exceptions, this timeout is thrown by `send()` and NOT by the resulting `Future`."**
That distinction matters: a `try/catch` around `send()` catches this one; a callback does not.
##### `compression.type` [#compressiontype]
**Default: messages are sent uncompressed.** Options: **`snappy`, `gzip`, `lz4`, `zstd`**.
| Codec | Character | Recommended when |
| ---------- | ---------------------------------------------------------------------------------------------------------- | ------------------------------------------------ |
| **snappy** | *"invented by Google to provide decent compression ratios with **low CPU overhead** and good performance"* | **both performance and bandwidth are a concern** |
| **gzip** | *"typically use more CPU and time but results in **better compression ratios**"* | **network bandwidth is more restricted** |
| lz4, zstd | (also available) | |
> *"By enabling compression, you reduce network utilization and storage, **which is often a bottleneck when sending messages to Kafka.**"*
*(Cross-reference Ch. 2: the broker decompresses to validate checksums and assign offsets, then recompresses. Compression is not free on the broker either.)*
##### `batch.size` [#batchsize]
**Memory in bytes (not messages!)** used for each batch. When the batch is full, all messages in it are sent.
**The crucial clarification most people get wrong:**
> *"However, this does **not** mean that the producer will wait for the batch to become full. **The producer will send half-full batches and even batches with just a single message in them.** Therefore, **setting the batch size too large will not cause delays in sending messages; it will just use more memory** for the batches. Setting the batch size too small will add some overhead because the producer will need to send messages more frequently."*
So: `batch.size` is a **ceiling and a memory cost**, not a trigger to wait for. **`linger.ms` is the only thing that makes the producer wait.**
##### `max.in.flight.requests.per.connection` [#maxinflightrequestsperconnection]
How many message batches the producer will send **without receiving responses**.
> *"Higher settings can increase memory usage while improving throughput. Apache's wiki experiments show that **in a single-DC environment, the throughput is maximized with only 2 in-flight requests; however, the default value is 5 and shows similar performance.**"*
***
# 3.0 The motivating example (keep this in your head) (/docs/kafka/kafka-producers-writing-messages/motivating-example-keep-head)
A credit-card transaction processing system:
**Its requirements, stated precisely:**
* **Never lose a single message. Never duplicate any message.**
* Latency low, but **up to 500 ms tolerable**.
* **Very high throughput — up to a million messages/second.**
**Contrast: website click tracking.**
* **Some loss and a few duplicates are tolerable.**
* **Latency can be high**, as long as the user experience isn't affected — "we don't mind if it takes a few seconds for the message to arrive at Kafka, as long as the next page loads immediately after the user clicks."
* Throughput depends on anticipated activity.
**This is the chapter's core lesson:** *"The different requirements will influence the way you use the producer API to write messages to Kafka and the configuration you use."* There is no universally correct producer config. There is only a config that matches a stated requirement.
The questions to answer *before* configuring a producer:
1. Is every message critical, or can we tolerate loss?
2. Are we OK with accidentally duplicating messages?
3. Are there strict latency or throughput requirements?
> ### Third-party clients [#third-party-clients]
>
> Kafka has a **binary wire protocol** — applications can read/write simply by sending the correct byte sequences to Kafka's network port. Multiple clients implement it in **C++, Python, Go, and many more**. These are **not part of the Apache Kafka project**; a list of non-Java clients is maintained in the project wiki.
***
# 3.6 ⚠️ ORDERING GUARANTEES — the reordering bug (/docs/kafka/kafka-producers-writing-messages/ordering-guarantees-reordering-bug)
> *"Apache Kafka preserves the order of messages **within a partition**. If messages are sent from the producer in a specific order, the broker will write them to a partition in that order and all consumers will read them in that order. For some use cases, order is very important. **There is a big difference between depositing $100 in an account and later withdrawing it, and the other way around!**"*
#### How retries silently reorder your data [#how-retries-silently-reorder-your-data]
#### The bind, and the resolution [#the-bind-and-the-resolution]
*(Ch. 8 covers the idempotent producer in depth.)*
**Do not use the old folklore fix** (`max.in.flight.requests.per.connection=1`) — it costs you throughput and is strictly worse than idempotence.
***
# 3.9 Partitions (/docs/kafka/kafka-producers-writing-messages/partitions)
#### 9.1 What keys are for [#91-what-keys-are-for]
> *"Kafka messages are key-value pairs, and while it is possible to create a `ProducerRecord` with just a topic and a value, with the key set to null by default, **most applications produce records with keys.**"*
**Keys serve two goals:**
1. **Additional information stored with the message.**
2. **Deciding which partition the message is written to.** *(Keys also play an important role in **compacted topics** — Ch. 6.)*
> **"All messages with the same key will go to the same partition. This means that if a process is reading only a subset of the partitions in a topic, all the records for a single key will be read by the same process."**
```java
// with key
ProducerRecord record =
new ProducerRecord<>("CustomerCountry", "Laboratory Equipment", "USA");
// null key — just leave it out
ProducerRecord record =
new ProducerRecord<>("CustomerCountry", "USA");
```
#### 9.2 Default partitioner behavior [#92-default-partitioner-behavior]
##### Null key → round-robin, and since 2.4, *sticky* round-robin [#null-key--round-robin-and-since-24-sticky-round-robin]
> When the key is null and the default partitioner is used, the record goes to **one of the *available* partitions at random**; a **round-robin algorithm** balances messages among partitions.
>
> **Starting in the Apache Kafka 2.4 producer, the round-robin algorithm for null keys is *sticky*.** *"This means that it will **fill a batch of messages sent to a single partition before switching to the next partition.** This allows sending the same number of messages to Kafka in **fewer requests, leading to lower latency and reduced CPU utilization on the broker.**"*
##### Non-null key → hash [#non-null-key--hash]
> Kafka will **hash the key** — *"using its own hash algorithm, so **hash values will not change when Java is upgraded**"* — and use the result to map to a partition.
**A crucial detail with a real consequence:**
> *"Since it is important that a key is always mapped to the same partition, **we use ALL the partitions in the topic to calculate the mapping — not just the available partitions.** This means that **if a specific partition is unavailable when you write data to it, you might get an error.**"* (Fairly rare — see Ch. 7 on replication and availability.)
That's a genuine, deliberate availability-for-consistency trade. You cannot have both stable key routing and immunity to offline partitions.
#### 9.3 Other built-in partitioners: `RoundRobinPartitioner` and `UniformStickyPartitioner` [#93-other-built-in-partitioners-roundrobinpartitioner-and-uniformstickypartitioner]
These give **random** and **sticky random** partition assignment **even when messages have keys**.
**Why you'd want that:**
> *"These are useful when **keys are important for the consuming application** (for example, there are **ETL applications that use the key from Kafka records as the primary key when loading data from Kafka to a relational database**), but **the workload may be skewed, so a single key may have a disproportionately large workload.** Using the `UniformStickyPartitioner` will result in an even distribution of workload across all partitions."*
This decouples two things people conflate: **the key as data** vs **the key as routing**. You can keep the key for downstream semantics while abandoning key-based routing.
#### 9.4 ⚠️ Adding partitions breaks key→partition mapping [#94-️-adding-partitions-breaks-keypartition-mapping]
> \*"When the default partitioner is used, **the mapping of keys to partitions is consistent only as long as the number of partitions in a topic does not change.** As long as the number of partitions is constant, you can be sure that records regarding **user 045189 will always get written to partition 34**. This allows all kinds of optimization when reading data from partitions.
>
> **However, the moment you add new partitions to the topic, this is no longer guaranteed** — the old records will stay in partition 34 while new records may get written to a different partition."\*
> **"When partitioning keys is important, the easiest solution is to create topics with sufficient partitions and NEVER ADD PARTITIONS."**
That is about as blunt as this book gets. Combine with Ch. 2's guidance: size on **expected future** throughput, and remember partition count can only go **up**, never down — so you must get it right roughly once.
#### 9.5 Custom partitioner — the "Banana" hot-key problem [#95-custom-partitioner--the-banana-hot-key-problem]
**The scenario:** you're a B2B vendor; your biggest customer manufactures handheld devices called **Bananas**. *"You do so much business with customer 'Banana' that **over 10% of your daily transactions** are with this customer."*
**With default hash partitioning:** Banana's records land in the same partition as *other* accounts → **one partition much larger than the rest** → *"can cause servers to run out of space, processing to slow down, etc."*
**What you want:** *give Banana its own partition, then hash the rest of the accounts across all other partitions.*
```java
public class BananaPartitioner implements Partitioner {
public void configure(Map configs) {}
public int partition(String topic, Object key, byte[] keyBytes,
Object value, byte[] valueBytes,
Cluster cluster) {
List partitions = cluster.partitionsForTopic(topic);
int numPartitions = partitions.size();
if ((keyBytes == null) || (!(key instanceOf String)))
throw new InvalidRecordException("We expect all messages " +
"to have customer name as key");
if (((String) key).equals("Banana"))
return numPartitions - 1; // Banana always goes to last partition
// Other records get hashed to the rest of the partitions
return Math.abs(Utils.murmur2(keyBytes)) % (numPartitions - 1);
}
public void close() {}
}
```
**Interface:** `configure`, `partition`, `close`.
**The book's own self-critique, worth internalizing:** *"Here we only implement `partition`, although **we really should have passed the special customer name through `configure` instead of hardcoding it in `partition`.**"* — i.e. hot keys change; recompiling a partitioner to change a customer name is a deploy you shouldn't need.
Note also `Utils.murmur2` — this is the same hash Kafka's own default partitioner uses, which is why it's Java-version-stable.
***
# 3.12 Quotas and throttling (/docs/kafka/kafka-producers-writing-messages/quotas-throttling)
> *"Kafka brokers have the ability to limit the rate at which messages are produced and consumed. This is done via the **quota** mechanism."*
#### Three quota types [#three-quota-types]
| Type | Limits | Unit |
| ----------- | ------------------------------------------------------------------- | -------------------- |
| **produce** | rate at which clients can **send** data | **bytes per second** |
| **consume** | rate at which clients can **receive** data | **bytes per second** |
| **request** | **percentage of time the broker spends processing client requests** | **% of broker time** |
**Scope:** can be applied to **all clients (defaults)**, **specific client-ids**, **specific users**, or **both**.
> *"User-specific quotas are only meaningful in clusters where security is configured and clients authenticate."*
#### Static configuration (broker config file) — and why it's bad [#static-configuration-broker-config-file--and-why-its-bad]
```properties
# default for all clients
quota.producer.default=2M
# per-client overrides (while not recommended)
quota.producer.override="clientA:4M,clientB:10M"
```
> *"Quotas specified in Kafka's configuration file are **static**, and you can only modify them by **changing the configuration and then restarting all the brokers.** **Since new clients can arrive at any time, this is very inconvenient.**"*
#### Dynamic configuration — the usual method [#dynamic-configuration--the-usual-method]
Via **`kafka-config.sh`** or the **AdminClient API**:
```bash
# limit clientC (by client-id) to produce only 1024 B/s
bin/kafka-configs --bootstrap-server localhost:9092 --alter \
--add-config 'producer_byte_rate=1024' \
--entity-name clientC --entity-type clients
# limit user1 (by authenticated principal): produce 1024 B/s, consume 2048 B/s
bin/kafka-configs --bootstrap-server localhost:9092 --alter \
--add-config 'producer_byte_rate=1024,consumer_byte_rate=2048' \
--entity-name user1 --entity-type users
# limit ALL users to consume 2048 B/s, except users with a more specific
# override — this is how you dynamically modify the DEFAULT quota
bin/kafka-configs --bootstrap-server localhost:9092 --alter \
--add-config 'consumer_byte_rate=2048' --entity-type users
```
#### How throttling actually works — two mechanisms [#how-throttling-actually-works--two-mechanisms]
Mechanism 1 is elegant: it exploits the client's own in-flight limit as a natural rate limiter, so well-behaved clients need no throttling logic at all. Mechanism 2 is the enforcement backstop for clients that ignore the hint.
#### Throttling metrics exposed to clients [#throttling-metrics-exposed-to-clients]
```txt
produce-throttle-time-avg produce-throttle-time-max
fetch-throttle-time-avg fetch-throttle-time-max
```
The **average and maximum** time a produce/fetch request was delayed due to throttling.
> *"Note that this time can represent throttling due to **produce and consume throughput quotas, request time quotas, or both.** Other types of client requests **can only be throttled due to request time quotas**, and those will also be exposed via similar metrics."*
#### ⚠️ WARNING — the buffer-exhaustion cascade [#️-warning--the-buffer-exhaustion-cascade]
This is the most important operational warning in the chapter. Trace it:
> *"**It is therefore important to plan and monitor to make sure that the broker capacity over time will match the rate at which producers are sending data.**"*
**And note the echo of Ch. 1's ActiveMQ lesson:** step 4 is *the producer being stalled by the broker*. Async sends do not protect you — they only move the stall point from "every send" to "when the buffer fills."
***
# 3.7 Remaining configs (/docs/kafka/kafka-producers-writing-messages/remaining-configs)
##### `max.request.size` [#maxrequestsize]
Controls the size of a **produce request**. It caps **both**:
1. **the size of the largest message that can be sent**, and
2. **the number of messages the producer can send in one request.**
> *"With a default maximum request size of **1 MB**, the largest message you can send is 1 MB, **or** the producer can batch **1,024 messages of size 1 KB each** into one request."*
**Coordinate with the broker:**
> *"the broker has its own limit on the size of the largest message it will accept (`message.max.bytes`). **It is usually a good idea to have these configurations match**, so the producer will not attempt to send messages of a size that will be rejected by the broker."*
##### `receive.buffer.bytes` and `send.buffer.bytes` [#receivebufferbytes-and-sendbufferbytes]
TCP send/receive buffer sizes used by the sockets. **`-1` → use OS defaults.**
> *"It is a good idea to increase these when producers or consumers communicate with brokers **in a different datacenter**, because those network links typically have **higher latency and lower bandwidth**."*
##### `enable.idempotence` [#enableidempotence]
Kafka has supported exactly-once semantics **since version 0.11**. The idempotent producer is *"a simple and highly beneficial part of it."*
**The duplicate scenario — trace it carefully:**
Note what makes this so nasty: **nothing failed.** The data was written and replicated correctly. Only the *acknowledgment* was lost. Retry-on-timeout is fundamentally unable to distinguish "write failed" from "ack lost."
**The fix:**
> *"When the idempotent producer is enabled, **the producer will attach a sequence number to each record it sends.** If the broker receives records with the same sequence number, **it will reject the second copy** and the producer will receive the harmless **`DuplicateSequenceException`**."*
> **NOTE — required preconditions:**
>
> ```
> max.in.flight.requests.per.connection ≤ 5
> retries > 0
> acks = all
> ```
>
> *"If incompatible values are set, a `ConfigException` will be thrown."*
***
# 3.15 Self-test (/docs/kafka/kafka-producers-writing-messages/self-test)
Name the three mandatory producer properties. Why is `key.serializer` required even when you never set a key?
Walk the record through the producer: what happens on the app thread vs the sender thread, and in what order?
`send()` is described as "always asynchronous." What then is a "synchronous send," and what does it cost?
Fire-and-forget still throws some exceptions. Which ones, and what do they have in common?
Give two retriable errors and one non-retriable one. What determines the category?
Why does `acks=all` cost nothing in end-to-end latency? What *does* it cost?
Draw the delivery-time timeline and place `max.block.ms`, `linger.ms`, `request.timeout.ms`, `delivery.timeout.ms`.
Why does the book tell you to stop tuning `retries` and tune `delivery.timeout.ms` instead? Give the mental script.
Exactly how do `retries > 0` and `max.in.flight > 1` reorder your data? What's the correct fix, and why isn't `max.in.flight=1`?
A broker wrote *and replicated* your record, then crashed before acking. What happens next, and why can't retry logic ever solve this on its own?
What's the difference between `batch.size` and `linger.ms`? Which one makes the producer *wait*?
Why does raising `linger.ms` improve compression ratio specifically?
Why is `buffer.memory` exhaustion thrown by `send()` rather than the `Future`, and why does that matter for your error handling?
Callbacks run on the producer's main thread. Name the guarantee that buys you and the failure mode it creates.
Keyed records hash over *all* partitions, not just available ones. Why, and what does that cost you?
What changed about null-key partitioning in Kafka 2.4, and what three benefits does it produce?
When would you deliberately use `UniformStickyPartitioner` on records that *have* keys?
You need to add partitions to a keyed topic. What breaks? What's the book's advice?
List four reasons the hand-written `CustomerSerializer` is a liability.
In the Avro fax→email example, what happens when a new app reads an old record? An old app reads a new record?
Why can't you just embed the Avro schema in each Kafka record like you would in an Avro file?
What are the two Avro caveats the book insists on?
What can `onSend()` do that `onAcknowledgement()` cannot?
Name a use case where headers are strictly better than putting the metadata in the value.
Describe both mechanisms a broker uses to throttle a client over quota. Why are there two?
Trace the full buffer-exhaustion cascade from "sending too fast" to "TimeoutException," naming every config involved.
**Previous:** [Chapter 2 — Installing Kafka](02-installing-kafka.md)
**Next:** [Chapter 4 — Kafka Consumers](04-kafka-consumers.md)
# 3.8 Serializers (/docs/kafka/kafka-producers-writing-messages/serializers)
#### 8.1 Why you should not write your own — demonstrated by writing one [#81-why-you-should-not-write-your-own--demonstrated-by-writing-one]
```java
public class Customer {
private int customerID;
private String customerName;
public Customer(int ID, String name) { this.customerID = ID; this.customerName = name; }
public int getID() { return customerID; }
public String getName() { return customerName; }
}
```
```java
public class CustomerSerializer implements Serializer {
@Override
public void configure(Map configs, boolean isKey) { /* nothing to configure */ }
@Override
/**
We are serializing Customer as:
4 byte int — customerId
4 byte int — length of customerName in UTF-8 bytes (0 if name is Null)
N bytes — customerName in UTF-8
**/
public byte[] serialize(String topic, Customer data) {
try {
byte[] serializedName;
int stringSize;
if (data == null) return null;
else {
if (data.getName() != null) {
serializedName = data.getName().getBytes("UTF-8");
stringSize = serializedName.length;
} else {
serializedName = new byte[0];
stringSize = 0;
}
}
ByteBuffer buffer = ByteBuffer.allocate(4 + 4 + stringSize);
buffer.putInt(data.getID());
buffer.putInt(stringSize);
buffer.put(serializedName);
return buffer.array();
} catch (Exception e) {
throw new SerializationException("Error when serializing Customer to byte[] " + e);
}
}
@Override
public void close() { /* nothing to close */ }
}
```
**Now count the ways this fails:**
> \*"This example is pretty simple, but you can see **how fragile the code is.**
>
> * If we ever have too many customers and need to **change `customerID` to `Long`**...
> * or if we ever decide to **add a `startDate` field**...
> * ...we will have a **serious issue maintaining compatibility between old and new messages.**
>
> **Debugging compatibility issues between different versions of serializers and deserializers is fairly challenging: you need to compare arrays of raw bytes.**
>
> To make matters even worse, **if multiple teams in the same company end up writing `Customer` data to Kafka, they will all need to use the same serializers and modify the code at the exact same time.**"\*
That last point is the organizational killer — the same lockstep-deploy problem Ch. 1 identified, now embedded in your byte layout.
**Recommendation: use existing serializers — JSON, Apache Avro, Thrift, or Protobuf.**
#### 8.2 Apache Avro [#82-apache-avro]
**What it is:** a **language-neutral** data serialization format, created by **Doug Cutting** to provide a way to share data files with a large audience.
* Data described in a **language-independent schema**, usually in **JSON**.
* Serialization usually to **binary** (JSON output also supported).
* **Avro assumes the schema is present when reading and writing files** — usually by **embedding the schema in the files themselves**.
**The killer feature for messaging:**
> *"When the application writing messages switches to a **new but compatible** schema, the applications **reading** the data **can continue processing messages without requiring any change or update.**"*
##### Worked schema evolution example [#worked-schema-evolution-example]
**Original schema** (used for months, **a few terabytes of data** generated):
```json
{"namespace": "customerManagement.avro",
"type": "record",
"name": "Customer",
"fields": [
{"name": "id", "type": "int"},
{"name": "name", "type": "string"},
{"name": "faxNumber", "type": ["null", "string"], "default": "null"}
]
}
```
`id` and `name` are **mandatory**; `faxNumber` is **optional, defaults to null**.
**New schema** — *"upgrade to the 21st century"*: drop fax, add email:
```json
{"namespace": "customerManagement.avro",
"type": "record",
"name": "Customer",
"fields": [
{"name": "id", "type": "int"},
{"name": "name", "type": "string"},
{"name": "email", "type": ["null", "string"], "default": "null"}
]
}
```
**The real-world constraint:** *"In many organizations, upgrades are done slowly and over many months."* So old records have `faxNumber` and new records have `email`, **simultaneously**, and both old and new readers must cope.
##### Two caveats — do not skip these [#two-caveats--do-not-skip-these]
> 1. **"The schema used for writing the data and the schema expected by the reading application must be compatible."** (The Avro documentation includes the compatibility rules.)
> 2. **"The deserializer will need access to the schema that was used when writing the data, even when it is different from the schema expected by the application."** In Avro *files*, the writing schema is included in the file itself — **but there is a better way for Kafka messages.**
#### 8.3 Avro + Schema Registry [#83-avro--schema-registry]
**The problem with embedding schemas per record:**
> *"Unlike Avro files, where storing the entire schema in the data file is associated with a fairly reasonable overhead, **storing the entire schema in each record will usually more than double the record size.** However, Avro still requires the entire schema to be present when reading the record, so we need to **locate the schema elsewhere.**"*
**The pattern:**
> **The Schema Registry is NOT part of Apache Kafka**, but there are several open source options. The book uses the **Confluent Schema Registry** (on GitHub, or as part of the Confluent Platform).
##### Producing generated Avro objects [#producing-generated-avro-objects]
```java
Properties props = new Properties();
props.put("bootstrap.servers", "localhost:9092");
props.put("key.serializer",
"io.confluent.kafka.serializers.KafkaAvroSerializer");
props.put("value.serializer",
"io.confluent.kafka.serializers.KafkaAvroSerializer");
props.put("schema.registry.url", schemaUrl); // where schemas live
String topic = "customerContacts";
Producer producer = new KafkaProducer<>(props);
// keep producing new events until someone ctrl-c
while (true) {
Customer customer = CustomerGenerator.getNext();
System.out.println("Generated customer " + customer.toString());
ProducerRecord record =
new ProducerRecord<>(topic, customer.getName(), customer);
producer.send(record);
}
```
Notes:
* `KafkaAvroSerializer` **can also handle primitives** — that's why `String` works as the key while `Customer` is the value.
* **`Customer` is NOT a POJO.** It is *"a specialized Avro object, generated from a schema using Avro code generation."* **"The Avro serializer can only serialize Avro objects, not POJO."** Generate classes with **`avro-tools.jar`** or the **Avro Maven plug-in** (both part of Apache Avro).
##### Producing generic Avro objects (no codegen) [#producing-generic-avro-objects-no-codegen]
```java
Properties props = new Properties();
props.put("bootstrap.servers", "localhost:9092");
props.put("key.serializer", "io.confluent.kafka.serializers.KafkaAvroSerializer");
props.put("value.serializer", "io.confluent.kafka.serializers.KafkaAvroSerializer");
props.put("schema.registry.url", url);
String schemaString =
"{\"namespace\": \"customerManagement.avro\","
+ "\"type\": \"record\", "
+ "\"name\": \"Customer\","
+ "\"fields\": ["
+ "{\"name\": \"id\", \"type\": \"int\"},"
+ "{\"name\": \"name\", \"type\": \"string\"},"
+ "{\"name\": \"email\", \"type\": [\"null\",\"string\"], \"default\":\"null\" }"
+ "]}";
Producer producer = new KafkaProducer(props);
Schema.Parser parser = new Schema.Parser();
Schema schema = parser.parse(schemaString);
for (int nCustomers = 0; nCustomers < customers; nCustomers++) {
String name = "exampleCustomer" + nCustomers;
String email = "example " + nCustomers + "@example.com";
GenericRecord customer = new GenericData.Record(schema);
customer.put("id", nCustomers);
customer.put("name", name);
customer.put("email", email);
ProducerRecord data =
new ProducerRecord<>("customerContacts", name, customer);
producer.send(data);
}
```
Generic Avro objects are **used as key-value maps** rather than generated objects with getters/setters. You must **provide the schema** yourself, since no generated object carries it. *"The serializer will know how to get the schema from this record, store it in the Schema Registry, and serialize the object data."*
***
# 3.3 The three ways to send (/docs/kafka/kafka-producers-writing-messages/three-ways-send)
#### 3.1 Fire-and-forget [#31-fire-and-forget]
```java
ProducerRecord record =
new ProducerRecord<>("CustomerCountry", "Precision Products", "France");
try {
producer.send(record);
} catch (Exception e) {
e.printStackTrace();
}
```
> *"Most of the time it will arrive successfully, since Kafka is highly available and the producer will retry sending messages automatically. **However, in case of nonretriable errors or timeout, messages will get lost and the application will not get any information or exceptions about this.**"*
**Subtle but important:** even here you can still get an exception — but only for failures **before** the message reached the broker path:
* **`SerializationException`** — failed to serialize
* **`BufferExhaustedException` / `TimeoutException`** — the buffer is full
* **`InterruptException`** — the sending thread was interrupted
Everything that goes wrong *at or after* the broker is invisible. *"This method of sending messages can be used when dropping a message silently is acceptable. This is not typically the case in production applications."*
#### 3.2 Synchronous send [#32-synchronous-send]
```java
producer.send(record).get(); // ← the .get() is the whole difference
```
**What it buys you:** catches exceptions when Kafka responds to the produce request with an error, **or when send retries were exhausted**.
**What it costs — and the numbers are worth memorizing:**
> *"Depending on how busy the Kafka cluster is, brokers can take anywhere from **2 ms to a few seconds** to respond to produce requests. If you send messages synchronously, the sending thread will spend this time **waiting and doing nothing else, not even sending additional messages.** This leads to very poor performance, and as a result, **synchronous sends are usually not used in production applications (but are very common in code examples).**"*
That parenthetical is a warning about copy-paste: the pattern you see in every tutorial is the one you must not ship.
#### 3.3 Retriable vs non-retriable errors [#33-retriable-vs-non-retriable-errors]
`KafkaProducer` has **two types of errors**:
> `KafkaProducer` can be configured to retry retriable errors automatically, *"so the application code will get retriable exceptions **only when the number of retries was exhausted and the error was not resolved.**"*
#### 3.4 Asynchronous send with a callback — the production pattern [#34-asynchronous-send-with-a-callback--the-production-pattern]
**The arithmetic that makes the case:**
> *"Suppose the network round-trip time between our application and the Kafka cluster is 10 ms. If we wait for a reply after sending each message, sending 100 messages will take around **1 second**. On the other hand, if we just send all our messages and not wait for any replies, then sending 100 messages will **barely take any time at all.**"*
> *"In most cases, we really don't need a reply — Kafka sends back the topic, partition, and offset of the record after it was written, which is usually not required by the sending app. **On the other hand, we do need to know when we failed to send a message completely** so we can throw an exception, log an error, or perhaps write the message to an 'errors' file for later analysis."*
```java
private class DemoProducerCallback implements Callback {
@Override
public void onCompletion(RecordMetadata recordMetadata, Exception e) {
if (e != null) {
e.printStackTrace(); // production: real error handling here
}
}
}
ProducerRecord record =
new ProducerRecord<>("CustomerCountry", "Biomedical Materials", "USA");
producer.send(record, new DemoProducerCallback());
```
Implement `org.apache.kafka.clients.producer.Callback` — **a single method, `onCompletion()`**. If Kafka returned an error, the exception argument is **non-null**.
> ### ⚠️ WARNING — the callback threading rule [#️-warning--the-callback-threading-rule]
>
> **Callbacks execute in the producer's main thread.**
>
> **Benefit:** guarantees that when you send two messages to the same partition one after another, **their callbacks execute in the same order you sent them**.
>
> **Cost:** *"the callback should be reasonably fast to avoid delaying the producer and preventing other messages from being sent. **It is not recommended to perform a blocking operation within the callback.** Instead, you should use another thread to perform any blocking operation concurrently."*
**This is a top-tier production landmine.** A callback that writes failures to a database, calls an HTTP endpoint, or logs synchronously to a slow appender will **throttle your entire producer**. Symptom: producer throughput collapses and nobody can see why, because the code "looks async."
***
# 5.9 What actually breaks in production — Ch. 5 consolidated (/docs/kafka/managing-apache-kafka-programmatically/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ---------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| 1 | **`createTopics().get()` succeeds, then `listTopics()` doesn't show it** | **Eventual consistency** — the Future completes when the *controller* is updated; the read went to a broker that hasn't got the metadata yet | Retry with backoff; never assert immediately after a mutation |
| 2 | **`deleteTopics().get()` succeeds, topic still described** | Same — *"due to the async nature of deletes, it is possible that at this point the topic still exists"* | Same |
| 3 | **Catch block never matches; you can't identify the error** | You checked the `ExecutionException` type instead of **`e.getCause()`** | **Always inspect `getCause()`** — *all* AdminClient results wrap errors this way |
| 4 | **`TopicExistsException` on startup, intermittently** | **Race:** two app instances both describe (not found), both create | Catch it and treat as success — then *describe* to validate config |
| 5 | **Irrecoverable data loss from a wrong topic name** | *"deletion of topics is FINAL — no recycle bin, no checks that the topic is empty"* | Broker-side `delete.topic.enable=false`; confirmation + audit in any tooling |
| 6 | **HTTP/API server thread pool exhausted while Kafka is slow** | Blocking `.get()` inside a request handler | `KafkaFuture.whenComplete()` + per-call `Options().timeoutMs()` |
| 7 | **Service fails to start because Kafka is slow** | Startup topic validation with the 120 s default `request.timeout.ms`, treated as fatal | Lower the timeout **and** start anyway, validating later or skipping |
| 8 | **SASL authentication fails via a DNS alias** | Client authenticates the alias; server principal is the real hostname → **SASL treats the mismatch as a possible MITM** | `client.dns.lookup=resolve_canonical_bootstrap_servers_only` |
| 9 | **Client can't connect even though brokers are healthy (K8s/LB)** | Client tries **only the first resolved IP**; that LB IP is down | `client.dns.lookup=use_all_dns_ips` |
| 10 | **Consumer group offsets won't update; `UnknownMemberIdException`** | **The group is still active** — Kafka blocks offset edits to live groups because consumers would overwrite them | Shut the consuming application down first (**no admin command exists** for this) |
| 11 | **Stateful stream app double-counts after an offset reset** | Offsets reset, **but the state store still holds the old aggregate** | Reset **both**; in dev, delete the state store entirely first |
| 12 | **Deleting offsets produced unpredictable behavior** | Post-delete behavior is decided by the *consumer's* `auto.offset.reset`, which the operator may not know | **Set** offsets explicitly (e.g. to earliest) instead of deleting |
| 13 | **`listConsumerGroups` throws on one bad group and returns nothing** | Used `.all()` — throws on the first error | Use `.valid()` (+ `.errors()` to inspect), for tooling that must degrade |
| 14 | **Group listing empty / describe fails** | Authorization, or **the group's coordinator is unavailable** | Check ACLs and coordinator health |
| 15 | **Lag monitoring broke after a Kafka upgrade** | Custom code **parsed `__consumer_offsets` internal messages** — *"Kafka does not guarantee compatibility of the internal message formats"* | `listConsumerGroupOffsets` + `listOffsets` |
| 16 | **`createPartitions` created the wrong number of partitions** | The argument is the **TOTAL after expansion**, not the count to add | `describe` first, then `increaseTo(current + n)` |
| 17 | **Multi-topic expansion partially applied** | *"some of the topics will be successfully expanded, while others will fail"* | Make tooling idempotent and re-runnable; check each result |
| 18 | **Keyed consumers break after adding partitions** | `hash(key) % N` changes (Ch. 3 §9.4) | *"check that the operation will not break any application that consumes from the topic"* |
| 19 | **Regulator finds 90-day-old data on a 30-day-retention topic** | *"retention policies were not implemented in a way that guarantees legal compliance"* — a single unclosed segment retains everything | `listOffsets(forTimestamp)` + `deleteRecords(beforeOffset)`; also fix segment rolling (Ch. 2) |
| 20 | **Records "deleted" but disk usage unchanged** | *"Full cleanup from disk will happen **asynchronously**"* — deletion first only makes records inaccessible | Expected; don't gate disk-space alarms on it |
| 21 | **Cluster network saturated; replication falls behind after a reassignment** | Replica reassignment copies **large amounts of data** with no throttle | **Throttle replication using quotas** (broker config, editable via AdminClient) |
| 22 | **Leadership didn't move after a reassignment** | The **first element** of the replica list is the preferred leader; you kept the old broker first | Order the list intentionally; then run **preferred** leader election |
| 23 | **`ElectionNotNeededException`** | Cluster healthy; the preferred leader already *is* the leader | Not an error — handle it as a no-op |
| 24 | **Partition permanently unavailable, no eligible leader** | Leader down; all other replicas are **missing data** so are ineligible | Either wait for the old leader, or accept **unclean leader election** and **permanent silent data loss** |
| 25 | **Reassignment/election results look inconsistent right after the call** | Async metadata propagation | Poll `listPartitionReassignments()` / re-describe over time |
| 26 | **Broker config file destroyed during an upgrade, no backup** | No config backup process | **A surviving broker IS your backup:** `describeConfigs` against it (the book's war story) |
| 27 | **A topic silently stopped being compacted and data aged out** | Config drift on a topic your app depends on | Periodically validate topic config from the app — *"more frequently than the default retention period, just to be safe"* |
| 28 | **`UnsupportedOperationException: Not implemented yet` in unit tests** | `MockAdminClient` doesn't mock everything (e.g. `incrementalAlterConfigs` ≤ 2.5) | Inject your own implementation (Mockito `doReturn`) |
| 29 | **`MockAdminClient` not found on the test classpath** | It ships in a **test jar** | Add `test` to the dependency |
| 30 | **Admin tooling broke on a ZooKeeper-less (KRaft) cluster** | Code manipulated **ZooKeeper directly** | *"NEVER use ZooKeeper directly"* — AdminClient's API survives the migration |
| 31 | **All mutations fail while all reads succeed** | **Writes go to the controller**; reads go to any (least-loaded) broker → a sick controller shows exactly this asymmetry | Check controller health (`describeCluster().controller()`) |
| 32 | **Ran a destructive tool against the wrong cluster** | No cluster identity check | Compare `cluster.clusterId()` (a GUID) before destructive operations |
***
# 5.7 Cluster metadata and advanced operations (/docs/kafka/managing-apache-kafka-programmatically/cluster-metadata-advanced-operations)
#### 7.1 Cluster metadata [#71-cluster-metadata]
> *"It is rare that an application has to explicitly discover anything at all about the cluster to which it connected. **You can produce and consume messages without ever learning how many brokers exist and which one is the controller. Kafka clients abstract away this information — clients only need to be concerned with topics and partitions.**"*
```java
DescribeClusterResult cluster = admin.describeCluster();
System.out.println("Connected to cluster " + cluster.clusterId().get());
System.out.println("The brokers in the cluster are:");
cluster.nodes().get().forEach(node -> System.out.println(" * " + node));
System.out.println("The controller is: " + cluster.controller().get());
```
> *"Cluster identifier is a **GUID** and therefore is **not human readable. It is still useful to check whether your client connected to the correct cluster.**"*
*(That's a real safety check: "am I about to run this destructive tool against staging or prod?" — compare the cluster ID.)*
***
#### Advanced operations — the SRE toolkit [#advanced-operations--the-sre-toolkit]
> *"a few methods that are **rarely used, and can be risky to use, but are incredibly useful when needed.** Those are mostly important for **SREs during incidents — but don't wait until you are in an incident to learn how to use them. Read and practice before it is too late.**"*
#### 7.2 Adding partitions to a topic [#72-adding-partitions-to-a-topic]
**Why it's usually unnecessary and always risky:**
> *"Usually the number of partitions is set when a topic is created. And **since each partition can have very high throughput, bumping against the capacity limits of a topic is rare.** In addition, **if messages in the topic have keys, then consumers can assume that all messages with the same key will always go to the same partition and will be processed in the same order by the same consumer.** For these reasons, **adding partitions to a topic is rarely needed and can be risky. You'll need to check that the operation will not break any application that consumes from the topic.**"*
*(This is the same warning as Ch. 3 §9.4 — `hash(key) % N` changes when N changes.)*
```java
Map newPartitions = new HashMap<>();
newPartitions.put(TOPIC_NAME, NewPartitions.increaseTo(NUM_PARTITIONS+2));
admin.createPartitions(newPartitions).all().get();
```
> ⚠️ **"When expanding topics, you need to specify the TOTAL number of partitions the topic will have AFTER the partitions are added, NOT the number of new partitions."**
>
> **TIP:** *"you may need to **describe the topic and find out how many partitions exist prior to expanding it.**"*
> ⚠️ *"if you try to expand multiple topics at once, **it is possible that some of the topics will be successfully expanded, while others will fail.**"*
**The `increaseTo` naming is a mercy** — but the partial-failure property means multi-topic expansion is not atomic. Write your tooling to be re-runnable.
#### 7.3 Deleting records — the compliance tool [#73-deleting-records--the-compliance-tool]
**The compliance gap, stated plainly:**
> *"Current privacy laws mandate specific retention policies for data. **Unfortunately, while Kafka has retention policies for topics, they were not implemented in a way that guarantees legal compliance. A topic with a retention policy of 30 days can store older data if all the data fits into a single segment in each partition.**"*
*(This is exactly the Ch. 2 low-volume-topic bug: retention only applies to **closed** segments, so a slow topic keeps data far longer than configured. Here the book names the legal consequence.)*
```java
Map olderOffsets =
admin.listOffsets(requestOlderOffsets).all().get(); // forTimestamp()
Map recordsToDelete = new HashMap<>();
for (Map.Entry e:
olderOffsets.entrySet())
recordsToDelete.put(e.getKey(),
RecordsToDelete.beforeOffset(e.getValue().offset()));
admin.deleteRecords(recordsToDelete).all().get();
```
**What `deleteRecords` actually does:**
> *"will **mark as deleted** all the records with offsets older than those specified... and **make them inaccessible by Kafka consumers.** The method **returns the highest deleted offsets**, so we can check if the deletion indeed happened as expected. **Full cleanup from disk will happen asynchronously.**"*
Note the honest distinction: records become **inaccessible to consumers immediately**, but **disk cleanup is asynchronous**. For strict "the bytes are gone" requirements, that gap matters.
#### 7.4 Leader election [#74-leader-election]
```java
Set electableTopics = new HashSet<>();
electableTopics.add(new TopicPartition(TOPIC_NAME, 0));
try {
admin.electLeaders(ElectionType.PREFERRED, electableTopics).all().get();
} catch (ExecutionException e) {
if (e.getCause() instanceof ElectionNotNeededException) {
System.out.println("All leaders are preferred already");
}
}
```
> *"If you call the command with **`null` instead of a collection of partitions, it will trigger the election type you chose for ALL partitions.**"*
>
> *"If the cluster is in a healthy state, the command will do nothing. **Preferred leader election and unclean leader election only take effect when a replica other than the preferred leader is the current leader.**"*
##### Type 1: Preferred leader election — safe [#type-1-preferred-leader-election--safe]
> *"Each partition has a replica designated as the **preferred leader**. It is preferred because **if all partitions use their preferred leader replica as the leader, the number of leaders on each broker should be balanced.**"*
>
> *"By default, Kafka will **check every five minutes** if the preferred leader replica is indeed the leader, and if it isn't but it **is eligible** to become the leader, it will elect it. **If `auto.leader.rebalance.enable` is false, or if you want this to happen faster, `electLeader()` can trigger this process.**"*
##### Type 2: Unclean leader election — DATA LOSS BY DESIGN [#type-2-unclean-leader-election--data-loss-by-design]
> *"If the leader replica of a partition becomes unavailable, **and the other replicas are NOT eligible to become leaders (usually because they are MISSING DATA)**, the partition will be **without a leader and therefore unavailable.** One way to resolve this is to trigger **unclean leader election**, which means **electing a replica that is otherwise ineligible to become a leader as the leader anyway. THIS WILL CAUSE DATA LOSS — all the events that were written to the old leader and were not replicated to the new leader will be LOST.**"*
**Async caveat for both types:**
> *"even after it returns successfully, **it takes a while until all brokers become aware of the new state, and calls to `describeTopics()` can return inconsistent results.** If you trigger leader election for multiple partitions, it is possible that **the operation will be successful for some partitions and fail for others.**"*
#### 7.5 Reassigning replicas [#75-reassigning-replicas]
**The four reasons you'd do this:**
> *"Maybe **a broker is overloaded** and you want to move some replicas. Maybe you want to **add more replicas**. Maybe you want to **move all replicas off a broker so you can remove the machine.** Or maybe **a few topics are so noisy that you need to isolate them from the rest of the workload.**"*
> ⚠️ *"reassigning replicas from one broker to another **may involve copying LARGE amounts of data.** Be mindful of the available network bandwidth, and **throttle replication using quotas if needed; quotas are a broker configuration, so you can describe them and update them with AdminClient.**"*
**Scenario:** single broker ID 0 holds one replica of every partition; a new broker (ID 1) was added.
```java
Map> reassignment = new HashMap<>();
reassignment.put(new TopicPartition(TOPIC_NAME, 0),
Optional.of(new NewPartitionReassignment(Arrays.asList(0,1)))); // ①
reassignment.put(new TopicPartition(TOPIC_NAME, 1),
Optional.of(new NewPartitionReassignment(Arrays.asList(1)))); // ②
reassignment.put(new TopicPartition(TOPIC_NAME, 2),
Optional.of(new NewPartitionReassignment(Arrays.asList(1,0)))); // ③
reassignment.put(new TopicPartition(TOPIC_NAME, 3), Optional.empty()); // ④
admin.alterPartitionReassignments(reassignment).all().get();
System.out.println("currently reassigning: " +
admin.listPartitionReassignments().reassignments().get()); // ⑤
demoTopic = admin.describeTopics(TOPIC_LIST);
topicDescription = demoTopic.values().get(TOPIC_NAME).get();
System.out.println("Description of demo topic:" + topicDescription); // ⑥
```
**The key insight about the list:** **the FIRST element of the replica list is the preferred leader.** `[0,1]` keeps broker 0 preferred; `[1,0]` makes broker 1 preferred. That single ordering detail controls leadership migration.
⑥ *"remember that **it can take a while until it shows consistent results**"* — eventual consistency again.
***
# 5.5 Configuration management (/docs/kafka/managing-apache-kafka-programmatically/configuration-management)
**Config resources come in three types:**
> *"Checking and modifying **broker and broker logging** configuration is typically done using tools like `kafka-config.sh` or other Kafka management tools, but **checking and updating TOPIC configuration from the applications that use them is quite common.**"*
**The motivating self-healing pattern:**
> *"many applications rely on **compacted topics for correct operation.** It makes sense that **periodically (more frequently than the default retention period, just to be safe), those applications will check that the topic is indeed compacted and take action to correct the topic configuration if it is not.**"*
That parenthetical is the interesting bit: *check more often than retention*, because if the topic silently reverted to delete-retention, you want to notice **before your data ages out**.
```java
ConfigResource configResource =
new ConfigResource(ConfigResource.Type.TOPIC, TOPIC_NAME); // ①
DescribeConfigsResult configsResult =
admin.describeConfigs(Collections.singleton(configResource));
Config configs = configsResult.all().get().get(configResource); // ②
// print nondefault configs
configs.entries().stream().filter(
entry -> !entry.isDefault()).forEach(System.out::println);
// Check if topic is compacted
ConfigEntry compaction = new ConfigEntry(TopicConfig.CLEANUP_POLICY_CONFIG,
TopicConfig.CLEANUP_POLICY_COMPACT);
if (!configs.entries().contains(compaction)) {
// if topic is not compacted, compact it
Collection configOp = new ArrayList();
configOp.add(new AlterConfigOp(compaction, AlterConfigOp.OpType.SET)); // ③
Map> alterConf = new HashMap<>();
alterConf.put(configResource, configOp);
admin.incrementalAlterConfigs(alterConf).all().get();
} else {
System.out.println("Topic " + TOPIC_NAME + " is compacted topic");
}
```
① *"You can specify **multiple different resources from different types in the same request.**"*
② Result is *"a map from each `ConfigResource` to a collection of configurations."*
#### `isDefault()` — subtler than it looks [#isdefault--subtler-than-it-looks]
> *"Each configuration entry has an **`isDefault()`** method that lets us know which configs were modified. **A topic configuration is considered nondefault if a user configured the topic to have a nondefault value, OR if a BROKER-LEVEL configuration was modified and the topic that was created inherited this nondefault value from the broker.**"*
#### ③ The four `AlterConfigOp.OpType` operations [#-the-four-alterconfigopoptype-operations]
| OpType | Effect |
| -------------- | ------------------------------------------------------------------------------------------------------ |
| **`SET`** | Sets the configuration value |
| **`DELETE`** | **Removes the value and resets to the default** |
| **`APPEND`** | *"apply only to configurations with a **List** type"* — add values **without sending the entire list** |
| **`SUBTRACT`** | Same, for removal |
APPEND/SUBTRACT exist to avoid read-modify-write races on list configs (a lost-update hazard if two operators edit the same list simultaneously).
#### 💡 The war story — why `describeConfigs` is an SRE tool [#-the-war-story--why-describeconfigs-is-an-sre-tool]
> *"Describing the configuration can be **surprisingly handy in an emergency.** We remember a time when **during an upgrade, the configuration file for the brokers was accidentally replaced with a broken copy.** This was discovered **after restarting the first broker and noticing that it failed to start. The team did not have a way to recover the original**, and we prepared for significant trial and error as we attempted to reconstruct the correct configuration and bring the broker back to life. **A site reliability engineer (SRE) saved the day by connecting to one of the remaining brokers and dumping its configuration using the AdminClient.**"*
**The generalizable lesson:** *a running broker is a live, authoritative backup of its own configuration.* If you have one surviving node, you have your config — but only if you know this API exists **before** the incident.
***
# 5.6 Consumer group management (/docs/kafka/managing-apache-kafka-programmatically/consumer-group-management)
**Why programmatic offset control matters:**
> *"unlike most message queues, **Kafka allows you to reprocess data in the exact order in which it was consumed and processed earlier.** In Chapter 4 we explained how to use the Consumer APIs to go back and reread older messages. **But using these APIs means that you programmed the ability to reprocess data in advance into your application. Your application itself must expose the 'reprocess' functionality.**"*
**Two scenarios where you need it and the app doesn't provide it:**
1. **Troubleshooting a malfunctioning application during an incident**
2. **Preparing an application to start running on a new cluster during a disaster recovery failover**
#### 6.1 List consumer groups — and the `valid()` / `errors()` / `all()` distinction [#61-list-consumer-groups--and-the-valid--errors--all-distinction]
```java
admin.listConsumerGroups().valid().get().forEach(System.out::println);
```
> **Likely causes of such errors:** *"**authorization**, where you don't have permission to view the group, or cases when **the coordinator for some of the consumer groups is not available.**"*
`valid()` is the right choice for tooling that must degrade gracefully; `all()` is the right choice when partial results would be misleading.
#### 6.2 Describe a group [#62-describe-a-group]
```java
ConsumerGroupDescription groupDescription = admin
.describeConsumerGroups(CONSUMER_GRP_LIST)
.describedGroups().get(CONSUMER_GROUP).get();
System.out.println("Description of group " + CONSUMER_GROUP
+ ":" + groupDescription);
```
**What you get:**
> *"This description is very useful when troubleshooting consumer groups. **One of the most important pieces of information about a consumer group is MISSING from this description** — inevitably, we'll want to know **what was the last offset committed by the group for each partition and how much it is lagging behind** the latest messages in the log."*
#### 6.3 Compute consumer lag — the correct way [#63-compute-consumer-lag--the-correct-way]
> **The deprecated approach:** *"In the past, the only way to get this information was to **parse the commit messages that the consumer groups wrote to an internal Kafka topic.** While this method accomplished its intent, **Kafka does not guarantee compatibility of the internal message formats, and therefore the old method is not recommended.**"*
```java
Map offsets =
admin.listConsumerGroupOffsets(CONSUMER_GROUP) // ①
.partitionsToOffsetAndMetadata().get();
Map requestLatestOffsets = new HashMap<>();
for (TopicPartition tp: offsets.keySet()) {
requestLatestOffsets.put(tp, OffsetSpec.latest()); // ②
}
Map latestOffsets =
admin.listOffsets(requestLatestOffsets).all().get();
for (Map.Entry e: offsets.entrySet()) {
String topic = e.getKey().topic();
int partition = e.getKey().partition();
long committedOffset = e.getValue().offset();
long latestOffset = latestOffsets.get(e.getKey()).offset();
System.out.println("Consumer group " + CONSUMER_GROUP
+ " has committed offset " + committedOffset
+ " to topic " + topic + " partition " + partition
+ ". The latest offset in the partition is "
+ latestOffset + " so consumer group is "
+ (latestOffset - committedOffset) + " records behind"); // ③
}
```
① ⚠️ *"unlike `describeConsumerGroups`, **`listConsumerGroupOffsets` only accepts a SINGLE consumer group and not a collection.**"*
② **`OffsetSpec` has three very convenient implementations:**
| Spec | Returns |
| -------------------- | ----------------------------------------------------------------------------------- |
| **`earliest()`** | the earliest offset in the partition |
| **`latest()`** | the latest offset in the partition |
| **`forTimestamp()`** | *"the offset of the record written **on or immediately after** the time specified"* |
`forTimestamp()` is the building block for both time-based offset resets and GDPR-style record deletion (§7.2).
#### 6.4 Modifying consumer groups [#64-modifying-consumer-groups]
**Available operations:** *"deleting groups, removing members, deleting committed offsets, and modifying offsets. These are **commonly used by SREs to build ad hoc tooling to recover from an emergency.**"*
##### Why *modify* rather than *delete* offsets [#why-modify-rather-than-delete-offsets]
> *"From all those, **modifying offsets is the most useful.** **Deleting offsets might seem like a simple way to get a consumer to 'start from scratch,' but this really depends on the configuration of the consumer** — if the consumer starts and no offsets are found, will it start from the beginning? Or jump to the latest message? **Unless we have the value of `auto.offset.reset`, we can't know.** Explicitly modifying the committed offsets to the earliest available offsets **will force the consumer to start processing from the beginning of the topic**, and essentially cause the consumer to 'reset.'"*
##### ⚠️ Constraint 1: you must stop the group first [#️-constraint-1-you-must-stop-the-group-first]
> *"consumer groups **don't receive updates when offsets change in the offset topic. They only read offsets when a consumer is assigned a new partition or on startup.** To prevent you from making changes to offsets that the consumers will not know about (and will therefore override), **Kafka will PREVENT you from modifying offsets while the consumer group is active.**"*
And note: *"this has to be done by **shutting down the consuming application directly; there is no admin command for shutting down a consumer group.**"*
##### ⚠️ Constraint 2: resetting offsets corrupts stateful applications [#️-constraint-2-resetting-offsets-corrupts-stateful-applications]
> *"if the consumer application maintains state (**and most stream processing applications maintain state**), resetting the offsets and causing the consumer group to start from the beginning **can have a strange impact on the stored state.**"*
**The shoe-counting example — internalize this one:**
> *"You need to take care to update the stored state accordingly. **In a development environment, we usually delete the state store completely before resetting the offsets to the start of the input topic.**"*
##### The reset code [#the-reset-code]
```java
Map earliestOffsets =
admin.listOffsets(requestEarliestOffsets).all().get(); // ①
Map resetOffsets = new HashMap<>();
for (Map.Entry e:
earliestOffsets.entrySet()) {
resetOffsets.put(e.getKey(), new OffsetAndMetadata(e.getValue().offset())); // ②
}
try {
admin.alterConsumerGroupOffsets(CONSUMER_GROUP, resetOffsets).all().get(); // ③
} catch (ExecutionException e) {
System.out.println("Failed to update the offsets committed by group "
+ CONSUMER_GROUP + " with error " + e.getMessage());
if (e.getCause() instanceof UnknownMemberIdException) // ④
System.out.println("Check if consumer group is still active.");
}
```
② *"we convert the map with `ListOffsetsResultInfo` values returned by `listOffsets` into a map with `OffsetAndMetadata` values required by `alterConsumerGroupOffsets`."* (A small type-shuffling annoyance worth knowing about.)
④ **The diagnostic to remember:**
> *"One of the most common reasons that `alterConsumerGroupOffsets` fails is that **we didn't stop the consumer group first.** ... If the group is still active, **our attempt to modify the offsets will appear to the consumer coordinator as if a client that is not a member of the group is committing an offset for that group.** In this case, we'll get **`UnknownMemberIdException`**."*
`UnknownMemberIdException` from `alterConsumerGroupOffsets` = **"the group is still running."** That mapping is not obvious from the exception name; memorize it.
***
# 5.10 Deploy / monitor / scale / backup — AdminClient's operational role (/docs/kafka/managing-apache-kafka-programmatically/deploy-monitor-scale-backup)
#### AdminClient as the monitoring primitive [#adminclient-as-the-monitoring-primitive]
This chapter is, quietly, where **lag monitoring** comes from:
**The important point:** this is the *supported* way. Parsing `__consumer_offsets` is unsupported and breaks on upgrade (#15).
#### AdminClient as the recovery toolkit [#adminclient-as-the-recovery-toolkit]
#### "Backup" — what this chapter contributes [#backup--what-this-chapter-contributes]
Kafka still has no backup command, but Ch. 5 adds two genuinely useful things:
1. **Configuration is recoverable from any running broker.** `describeConfigs` turns every live broker into a config backup. Consider dumping it to version control on a schedule — cheap insurance, per the war story.
2. **Deletion is *not* recoverable.** `deleteTopics` is final; `deleteRecords` is one-way. The only protections are `delete.topic.enable=false` and your own tooling discipline.
And one anti-pattern to retire: **do not back up or restore Kafka state by touching ZooKeeper.** That path is being removed.
#### Scaling AdminClient usage itself [#scaling-adminclient-usage-itself]
***
# 5.2 Design principles — understand these and every method becomes obvious (/docs/kafka/managing-apache-kafka-programmatically/design-principles-understand-these)
#### 2.1 Asynchronous and eventually consistent [#21-asynchronous-and-eventually-consistent]
> *"Perhaps **the most important thing to understand** about Kafka's AdminClient is that it is **asynchronous.** Each method **returns immediately** after delivering a request to the **cluster controller**, and each method returns one or more `Future` objects."*
**The wrapping hierarchy:**
#### ⚠️ 2.2 Eventual consistency — the read-your-own-write trap [#️-22-eventual-consistency--the-read-your-own-write-trap]
This is the single most important operational fact in the chapter:
> *"Because **Kafka's propagation of metadata from the controller to the brokers is asynchronous**, the Futures that AdminClient APIs return are considered complete **when the controller state has been fully updated.** At that point, **not every broker might be aware of the new state**, so a `listTopics` request **may end up handled by a broker that is not up-to-date and will not contain a topic that was very recently created.** This property is also called **eventual consistency**: eventually every broker will know about every topic, but **we can't guarantee exactly when.**"*
**Practical rule:** never write `createTopics(...).get(); listTopics(...)` and assert. Never write `deleteTopics(...).get()` and assert absence. Retry-with-backoff, or accept the ambiguity.
#### 2.3 Which broker handles what — and why it matters when debugging [#23-which-broker-handles-what--and-why-it-matters-when-debugging]
> *"This shouldn't impact you as an API user, but it can be good to know in case you are **seeing unexpected behavior**, you notice that **some operations succeed while others fail**, or if you are trying to figure out **why an operation is taking too long.**"*
That's the debugging hook: *writes all funnel through one node (the controller), reads spray across the cluster.* A sick controller breaks all mutations while all reads look fine.
#### 2.4 Options objects [#24-options-objects]
> *"Every method in AdminClient takes as an argument an **`Options` object specific to that method**"* — `listTopics` → `ListTopicsOptions`, `describeCluster` → `DescribeClusterOptions`.
**The universal setting: `timeoutMs`**
> *"this controls how long the client will wait for a response from the cluster before throwing a `TimeoutException`. **This limits the time in which your application may be blocked by an AdminClient operation.**"*
Other examples: whether `listTopics` should also return **internal topics**; whether `describeCluster` should also return **which operations the client is authorized to perform** on the cluster.
#### 2.5 Flat hierarchy — a deliberate, "controversial" choice [#25-flat-hierarchy--a-deliberate-controversial-choice]
> *"All admin operations supported by the Apache Kafka protocol are implemented in `KafkaAdminClient` **directly. There is no object hierarchy or namespaces.** This is a bit controversial as the interface can be quite large and perhaps a bit overwhelming, **but the main benefit is that if you want to know how to programmatically perform any admin operation on Kafka, you have exactly one JavaDoc to search, and your IDE autocomplete will be quite handy. You don't have to wonder whether you are just missing the right place to look. If it isn't in AdminClient, it was not implemented yet.**"*
#### 2.6 🚫 Never touch ZooKeeper directly [#26--never-touch-zookeeper-directly]
> *"At the time we are writing this chapter (Apache Kafka 2.5 is about to be released), most admin operations can be performed **either through AdminClient or directly by modifying the cluster metadata in ZooKeeper. We highly encourage you to NEVER use ZooKeeper directly, and if you absolutely have to, report this as a bug to Apache Kafka.**"*
**The reason is a forward-compatibility argument, not a stylistic one:**
AdminClient is the abstraction that makes the ZooKeeper→KRaft migration invisible to your code. That is its most underrated value.
***
# 5.4 Essential topic management (/docs/kafka/managing-apache-kafka-programmatically/essential-topic-management)
#### 4.1 List topics [#41-list-topics]
```java
ListTopicsResult topics = admin.listTopics();
topics.names().get().forEach(System.out::println);
```
> `admin.listTopics()` returns `ListTopicsResult` — *"a **thin wrapper over a collection of Futures**."* `topics.names()` returns a **Future set of names**. *"When we call `get()` on this Future, the executing thread will **wait until the server responds** with a set of topic names, **or we get a timeout exception.**"*
#### 4.2 Check existence + validate + create — the real pattern [#42-check-existence--validate--create--the-real-pattern]
**Why not just list and search?**
> *"One way to check if a specific topic exists is to get a list of all topics and check if the topic you need is in the list. **On a large cluster, this can be inefficient.** In addition, **sometimes you want to check for more than just whether the topic exists — you want to make sure the topic has the right number of partitions and replicas.**"*
**The real-world example — and it's instructive because each requirement has a reason:**
> *"**Kafka Connect and Confluent Schema Registry use a Kafka topic to store configuration.** When they start up, they check:*
>
> * *the configuration topic exists,*
> * *that it has **only one partition** — **to guarantee that configuration changes will arrive in strict order**,*
> * *that it has **three replicas** — **to guarantee availability**,*
> * *and that the topic is **compacted** — **so the old configuration will be retained indefinitely.**"*
```java
DescribeTopicsResult demoTopic = admin.describeTopics(TOPIC_LIST); // ①
try {
topicDescription = demoTopic.values().get(TOPIC_NAME).get(); // ②
System.out.println("Description of demo topic:" + topicDescription); // ③
if (topicDescription.partitions().size() != NUM_PARTITIONS) {
System.out.println("Topic has wrong number of partitions. Exiting.");
System.exit(-1);
}
} catch (ExecutionException e) { // ④
// exit early for almost all exceptions
if (! (e.getCause() instanceof UnknownTopicOrPartitionException)) {
e.printStackTrace();
throw e;
}
// if we are here, topic doesn't exist
System.out.println("Topic " + TOPIC_NAME +
" does not exist. Going to create it now");
// Note that number of partitions and replicas is optional. If they are
// not specified, the defaults configured on the Kafka brokers will be used
CreateTopicsResult newTopic = admin.createTopics(Collections.singletonList(
new NewTopic(TOPIC_NAME, NUM_PARTITIONS, REP_FACTOR))); // ⑤
// Check that the topic was created correctly:
if (newTopic.numPartitions(TOPIC_NAME).get() != NUM_PARTITIONS) { // ⑥
System.out.println("Topic has wrong number of partitions.");
System.exit(-1);
}
}
```
① `describeTopics()` with a list of names → `DescribeTopicsResult`, *"which wraps **a map of topic names to Future descriptions.**"*
② `get()` on the Future gives a `TopicDescription`… **or throws.**
#### ⚠️ ③④ The `ExecutionException` rule — this bites everyone once [#️--the-executionexception-rule--this-bites-everyone-once]
> *"if the topic does not exist, the server can't respond with its description. In this case, **the server will send back an error, and the Future will complete by throwing an `ExecutionException`. The actual error sent by the server will be the CAUSE of the exception.**"*
>
> **"Note that ALL AdminClient result objects throw `ExecutionException` when Kafka responds with an error. This is because AdminClient results are wrapped `Future` objects, and those wrap exceptions. YOU ALWAYS NEED TO EXAMINE THE `CAUSE` of `ExecutionException` to get the error that Kafka returned."**
```java
catch (ExecutionException e) {
Throwable actual = e.getCause(); // ← THE REAL ERROR IS HERE
if (actual instanceof UnknownTopicOrPartitionException) { ... }
}
```
**What a `TopicDescription` contains:** *"a list of all the partitions of the topic, and for each partition, **in which a broker is the leader, a list of replicas and a list of in-sync replicas.** Note that **this does NOT include the configuration of the topic**"* — configuration is a separate API (§5).
⑤ **Creating:** *"you can specify **just the name and use default values for all the details.** You can also specify the number of partitions, number of replicas, and the configuration."*
⑥ **Validating the creation:** *"Checking the result is **more common if you relied on broker defaults** when creating the topic."*
> ⚠️ *"since we are again calling `get()`... this method could throw an exception. **`TopicExistsException` is common in this scenario**, and you'll want to handle it (perhaps by describing the topic to check for the correct configuration)."*
`TopicExistsException` here is the classic **race**: two instances of your app start simultaneously, both describe (not found), both create, one wins.
#### 4.3 Delete topics [#43-delete-topics]
```java
admin.deleteTopics(TOPIC_LIST).all().get();
// Check that it is gone. Note that due to the async nature of deletes,
// it is possible that at this point the topic still exists
try {
topicDescription = demoTopic.values().get(TOPIC_NAME).get();
System.out.println("Topic " + TOPIC_NAME + " is still around");
} catch (ExecutionException e) {
System.out.println("Topic " + TOPIC_NAME + " is gone");
}
```
> ### ⚠️ WARNING — deletion is final [#️-warning--deletion-is-final]
>
> *"Although the code is simple, please remember that **in Kafka, deletion of topics is FINAL — there is no recycle bin or trash can** to help you rescue the deleted topic, and **no checks to validate that the topic is empty and that you really meant to delete it. Deleting the wrong topic could mean unrecoverable loss of data, so handle this method with extra care.**"*
Cross-reference Ch. 2: **`delete.topic.enable=false`** is the broker-side guardrail against exactly this. If you expose `deleteTopics` in any tooling, put a confirmation and an audit log in front of it.
#### 4.4 Non-blocking AdminClient — `KafkaFuture.whenComplete()` [#44-non-blocking-adminclient--kafkafuturewhencomplete]
**When blocking `get()` is wrong:**
> *"Most of the time, \[blocking] is all you need — **admin operations are rare**, and waiting until the operation succeeds or times out is usually acceptable. **There is one exception: if you are writing to a server that is expected to process a large number of admin requests.** In this case, **you don't want to block the server threads** while waiting for Kafka to respond. You want to continue accepting requests from your users and sending them to Kafka, and when Kafka responds, send the response to the client."*
```java
vertx.createHttpServer().requestHandler(request -> { // ① Vert.x
String topic = request.getParam("topic"); // ②
String timeout = request.getParam("timeout");
int timeoutMs = NumberUtils.toInt(timeout, 1000);
DescribeTopicsResult demoTopic = admin.describeTopics( // ③
Collections.singletonList(topic),
new DescribeTopicsOptions().timeoutMs(timeoutMs));
demoTopic.values().get(topic).whenComplete( // ④ NOT get()
new KafkaFuture.BiConsumer() {
@Override
public void accept(final TopicDescription topicDescription,
final Throwable throwable) {
if (throwable != null) {
request.response().end("Error trying to describe topic " // ⑤
+ topic + " due to " + throwable.getMessage());
} else {
request.response().end(topicDescription.toString()); // ⑥
}
}
});
}).listen(8080);
```
> *"**The key here is that we are not waiting for a response from Kafka.** `DescribeTopicResult` will send the response to the HTTP client when a response arrives from Kafka. **Meanwhile, the HTTP server can continue processing other requests.**"*
**A clever way to prove it works:**
> *"You can check this behavior by using **`SIGSTOP` to pause Kafka** (don't try this in production!) and send two HTTP requests to Vert.x: **one with a long timeout value and one with a short value. Even though you sent the second request after the first, it will respond earlier thanks to the lower timeout value, and not block behind the first request.**"*
***
# 5. Managing Apache Kafka Programmatically (AdminClient) (/docs/kafka/managing-apache-kafka-programmatically)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 5*
> **Learning goal:** `AdminClient` (added in **Kafka 0.11**) is the programmatic replacement for CLI admin. Two audiences: **app developers** who need topics-on-demand and config validation, and **SREs** who need to build ad-hoc recovery tooling. The book's own framing: *"SREs can think of it as a **Swiss Army knife** for Kafka operations."*
***
# 5.3 Lifecycle: create, configure, close (/docs/kafka/managing-apache-kafka-programmatically/lifecycle-create-configure-close)
```java
Properties props = new Properties();
props.put(AdminClientConfig.BOOTSTRAP_SERVERS_CONFIG, "localhost:9092");
AdminClient admin = AdminClient.create(props);
// TODO: Do something useful with AdminClient
admin.close(Duration.ofSeconds(30));
```
**Only mandatory config:** the cluster URI — a comma-separated broker list. *"As usual, in production environments, you want to specify **at least three brokers** just in case one is currently unavailable."*
*(Note the escalation: Ch. 3 said "at least two" for producers; here it's "at least three." Admin operations are rarer and more critical, so failing to bootstrap is worse.)*
#### `close()` semantics — read carefully [#close-semantics--read-carefully]
> *"AdminClient is much simpler \[than producer/consumer], and there is not much to configure."*
#### 3.1 `client.dns.lookup` — two distinct problems, two different values [#31-clientdnslookup--two-distinct-problems-two-different-values]
Introduced in **Kafka 2.1.0**. Default behavior: *"Kafka validates, resolves, and creates connections based on the hostname provided in the bootstrap server configuration (and later in the names returned by the brokers as specified in `advertised.listeners`)."*
> *"This simple model works most of the time but fails to cover two important use cases."* **They are mutually exclusive scenarios.**
##### Problem A: DNS alias + SASL → authentication failure [#problem-a-dns-alias--sasl--authentication-failure]
**Fix:**
```properties
client.dns.lookup=resolve_canonical_bootstrap_servers_only
```
> *"the client will **'expend' the DNS alias**, and the result will be **the same as if you included all the broker names the DNS alias connects to as brokers in the original bootstrap list.**"*
##### Problem B: one DNS name → multiple IPs (load balancers) → false unavailability [#problem-b-one-dns-name--multiple-ips-load-balancers--false-unavailability]
**Fix:**
```properties
client.dns.lookup=use_all_dns_ips
```
> *"**highly recommended** ... to make sure the client doesn't miss out on the benefits of a highly available load balancing layer."*
**Summary:**
| Scenario | Value | Symptom without it |
| ------------------------------------------------- | ------------------------------------------ | ---------------------------------------------------- |
| DNS **alias** for bootstrap + **SASL** | `resolve_canonical_bootstrap_servers_only` | SASL auth failure / connection refused |
| One DNS name → **multiple IPs** (LBs, Kubernetes) | `use_all_dns_ips` | Client can't connect even though brokers are healthy |
#### 3.2 `request.timeout.ms` (default **120 seconds**) [#32-requesttimeoutms-default-120-seconds]
> *"limits the time your application can spend waiting for AdminClient to respond. **This includes the time spent on retrying if the client receives a retriable error.**"*
**Why the default is so long:** *"quite long, but **some AdminClient operations, especially consumer group management commands, can take a while to respond.**"*
**Per-call override via `Options`** — and the book gives a very practical pattern:
> *"If an AdminClient operation is **on the critical path** for your application, you may want to use a **lower timeout** and handle a lack of timely response from Kafka **in a different way.** A common example is that services **try to validate the existence of specific topics when they first start**, but **if Kafka takes longer than 30 seconds to respond, you may want to continue starting the server and validate the existence of topics later (or skip this validation entirely).**"*
This is a genuinely good availability pattern: **don't couple your service's startup liveness to another system's responsiveness.**
***
# 5.1 What problem does AdminClient solve? (/docs/kafka/managing-apache-kafka-programmatically/problem-does-adminclient-solve)
#### The concrete motivating case: IoT device topics [#the-concrete-motivating-case-iot-device-topics]
> *"Creating new topics on demand based on user input or data is an especially common use case: **IoT apps often receive events from user devices, and write events to topics based on the device type.** If the manufacturer produces a new type of device, you either have to **remember, via some process, to also create a topic**, or the application can **dynamically create a new topic** if it receives events with an unrecognized device type."*
The book is honest about the trade-off: *"The second alternative has downsides, but **avoiding the dependency on an additional process to generate topics is an attractive feature in the right scenarios.**"*
#### The "does my topic exist?" problem — and the three bad answers that preceded AdminClient [#the-does-my-topic-exist-problem--and-the-three-bad-answers-that-preceded-adminclient]
Your app produces to a specific topic. **Before producing the first event, the topic has to exist.** Pre-AdminClient, your options were *"few\... none of them particularly user-friendly"*:
**With AdminClient:** *"use AdminClient to **check whether the topic exists, and if it does not, create it on the spot.**"*
#### What it covers [#what-it-covers]
> *"Apache Kafka added the AdminClient in version 0.11 to provide a programmatic API for administrative functionality that was previously done in the command line: **listing, creating, and deleting topics; describing the cluster; managing ACLs; and modifying configuration.**"*
This chapter focuses on the three most-used areas: **topics, consumer groups, and entity configuration.**
***
# 5.11 Self-test (/docs/kafka/managing-apache-kafka-programmatically/self-test)
Before AdminClient existed, what were the three ways to handle "the topic might not exist," and what was wrong with each?
AdminClient Futures complete when *what* is true — and what is *not* guaranteed at that moment?
Which broker handles mutations? Which handles reads? What debugging symptom does that asymmetry produce?
Why must you always inspect `e.getCause()` on an `ExecutionException`?
What does `close(Duration)` do differently from `close()`?
You use a DNS alias for bootstrap and SASL auth fails. Why, and what's the fix?
Your Kubernetes clients can't connect though brokers are healthy. Why, and what's the fix?
Why is `request.timeout.ms` defaulted to 120 s, and when should you lower it? Describe the startup-validation pattern.
Kafka Connect's config topic requires 1 partition, 3 replicas, and compaction. Give the reason for each.
What does `TopicDescription` contain — and what notable thing does it *not* contain?
What race produces `TopicExistsException`, and how should you handle it?
Why is topic deletion described as needing "extra care"? What broker config guards it?
When is blocking `.get()` the wrong choice, and what replaces it?
`isDefault()` returns false. Name the two different situations that can cause that.
What do `APPEND` and `SUBTRACT` exist for, and what hazard do they avoid?
Tell the config-recovery war story. What general principle does it establish?
Give the exact formula for consumer lag and the two API calls it requires.
What are the three `OffsetSpec` implementations, and what is `forTimestamp()` used for beyond seeking?
Why is *setting* offsets preferable to *deleting* them?
Why does Kafka refuse to modify offsets for an active group? What exception do you get, and what does it really mean?
Explain the shoe-counting double-count. State the general rule about offsets and state.
`.valid()` vs `.errors()` vs `.all()` on `listConsumerGroups` — when do you want each?
`createPartitions` takes what number, exactly? What must you do first?
Why does Kafka's retention config not satisfy privacy law, and what two APIs compose into a compliant deletion?
After `deleteRecords`, what's true immediately and what's only eventually true?
Contrast preferred and unclean leader election. What does each cost you?
In a reassignment replica list like `[1,0]`, what is the significance of the ordering?
How do you cancel an in-flight reassignment? What state does that restore?
What must you be careful about when reassigning replicas, and what tool does the book point you at?
Name two limitations of `MockAdminClient`, and the Maven detail people miss.
Why is "never use ZooKeeper directly" a forward-compatibility argument rather than a style preference?
**Previous:** [Chapter 4 — Kafka Consumers](04-kafka-consumers.md)
**Next:** [Chapter 6 — Kafka Internals](06-kafka-internals.md)
# 5.8 Testing with `MockAdminClient` (/docs/kafka/managing-apache-kafka-programmatically/testing-mockadminclient)
> *"Apache Kafka provides a test class, **`MockAdminClient`**, which you can initialize with any number of brokers and use to test that your applications behave correctly **without having to run an actual Kafka cluster** and really perform the admin operations on it."*
#### ⚠️ The stability caveat — stated honestly [#️-the-stability-caveat--stated-honestly]
> *"While `MockAdminClient` is **not part of the Kafka API and therefore subject to change without warning**, it mocks methods that are public, and **therefore the method signatures will remain compatible.** There is a bit of a trade-off on whether **the convenience of this class is worth the risk that it will change and break your tests**, so keep this in mind."*
#### What's good and what's missing [#whats-good-and-whats-missing]
#### The class under test [#the-class-under-test]
```java
public TopicCreator(AdminClient admin) {
this.admin = admin;
}
// Example of a method that will create a topic if its name starts with "test"
public void maybeCreateTopic(String topicName)
throws ExecutionException, InterruptedException {
Collection topics = new ArrayList<>();
topics.add(new NewTopic(topicName, 1, (short) 1));
if (topicName.toLowerCase().startsWith("test")) {
admin.createTopics(topics);
// alter configs just to demonstrate a point
ConfigResource configResource =
new ConfigResource(ConfigResource.Type.TOPIC, topicName);
ConfigEntry compaction =
new ConfigEntry(TopicConfig.CLEANUP_POLICY_CONFIG,
TopicConfig.CLEANUP_POLICY_COMPACT);
Collection configOp = new ArrayList();
configOp.add(new AlterConfigOp(compaction, AlterConfigOp.OpType.SET));
Map> alterConf = new HashMap<>();
alterConf.put(configResource, configOp);
admin.incrementalAlterConfigs(alterConf).all().get();
}
}
```
**Note the design that makes it testable:** `AdminClient` is **injected via the constructor**. That's the only reason `MockAdminClient` can be substituted.
> **NOTE:** *"We are using the **Mockito** testing framework to verify that the `MockAdminClient` methods are called as expected **and to fill in for the unimplemented methods.**"*
#### Setup [#setup]
```java
@Before
public void setUp() {
Node broker = new Node(0,"localhost",9092);
this.admin = spy(new MockAdminClient(Collections.singletonList(broker),
broker)); // ①
// without this, the tests will throw
// `java.lang.UnsupportedOperationException: Not implemented yet`
AlterConfigsResult emptyResult = mock(AlterConfigsResult.class); // ②
doReturn(KafkaFuture.completedFuture(null)).when(emptyResult).all();
doReturn(emptyResult).when(admin).incrementalAlterConfigs(any());
}
```
① *"`MockAdminClient` is instantiated with **a list of brokers** (here just one), and **one broker that will be our controller.** The brokers are just the broker ID, hostname, and port — **all fake, of course. No brokers will run while executing these tests.** We'll use Mockito's **`spy`** injection, so we can later check that `TopicCreator` executed correctly."*
② *"we use Mockito's `doReturn` methods to make sure the mock admin client doesn't throw exceptions. The method we are testing expects the `AlterConfigsResult` object with an `all()` method that returns a `KafkaFuture`. **We made sure that the fake `incrementalAlterConfigs` returns exactly that.**"*
#### The tests [#the-tests]
```java
@Test
public void testCreateTestTopic()
throws ExecutionException, InterruptedException {
TopicCreator tc = new TopicCreator(admin);
tc.maybeCreateTopic("test.is.a.test.topic");
verify(admin, times(1)).createTopics(any()); // name starts with "test"
}
@Test
public void testNotTopic() throws ExecutionException, InterruptedException {
TopicCreator tc = new TopicCreator(admin);
tc.maybeCreateTopic("not.a.test");
verify(admin, never()).createTopics(any()); // must NOT be called
}
```
#### The dependency you will forget [#the-dependency-you-will-forget]
> *"Apache Kafka published `MockAdminClient` in a **test jar**, so make sure your `pom.xml` includes a **test dependency**:"*
```xml
org.apache.kafkakafka-clients2.5.0testtest
```
***
# 1.6 What actually breaks in production (Ch. 1 foundations) (/docs/kafka/meet-kafka/actually-breaks-production-ch)
Chapter 1 is conceptual, but it plants several landmines that detonate later. Naming them now makes the rest of the book read as consequences rather than trivia.
#### 6.1 Assuming topic-wide ordering [#61-assuming-topic-wide-ordering]
**Symptom:** events processed out of order; state machines end up in impossible states; "the update arrived before the create."
**Cause:** ordering holds **per partition**, not per topic. Default producer behavior spreads messages **evenly across all partitions**.
**Fix:** key by the entity whose ordering you need (`user_id`, `account_id`, `trace_id`) so all its events land in one partition.
#### 6.2 Changing partition count on a keyed topic [#62-changing-partition-count-on-a-keyed-topic]
**Symptom:** after adding partitions, a key's messages start landing in a *different* partition than its history. Ordering silently breaks; stateful consumers see split history; compacted-topic semantics get weird.
**Cause:** `hash(key) mod N` — change `N`, change the mapping for most keys. The guarantee is explicitly conditional: same partition *"provided that the partition count does not change."*
**Also:** partition count can **only be increased, never decreased**.
**Fix:** size partitions for **expected future** throughput, not current. This is why Ch. 2 says to calculate based on future usage when keying.
#### 6.3 Assuming offsets are contiguous [#63-assuming-offsets-are-contiguous]
**Symptom:** gap-detection logic false-alarms; `expected == last + 1` assertions fail.
**Cause:** offsets increase but **"not necessarily monotonically greater"** — gaps are legal (compaction, transaction markers, aborted transactions).
**Fix:** treat offsets as opaque, increasing cursors. Never arithmetic.
#### 6.4 Schema-less topics → coupled deployments [#64-schema-less-topics--coupled-deployments]
**Symptom:** you cannot add a field without a coordinated multi-team deploy in a strict order; consumers crash on unknown fields; a producer rollback breaks consumers.
**Cause:** no schema contract; writing and reading are tightly coupled.
**Fix:** schema registry + a format with real compatibility rules (Avro). This is *the* lesson from LinkedIn's XML tracking system breaking "constantly due to changing schemas."
#### 6.5 Retention shorter than your recovery time [#65-retention-shorter-than-your-recovery-time]
**Symptom:** a consumer is down for maintenance longer than retention; on restart, its committed offset no longer exists; it either jumps to latest (**silent data loss**) or to earliest (**re-processes everything**).
**Cause:** retention defines a *minimum window*, and it is a **deletion policy**, not an archive.
**Fix:** retention ≥ worst-case consumer outage + recovery, with margin. Alert on consumer lag approaching the retention edge.
#### 6.6 Treating MirrorMaker as intra-cluster replication [#66-treating-mirrormaker-as-intra-cluster-replication]
**Symptom:** people expect cross-DC failover to be as seamless as broker failover. It isn't — offsets are not identical across clusters, and MirrorMaker is asynchronous.
**Cause:** Kafka's replication is explicitly **intra-cluster only**. MirrorMaker is a consumer+producer pair, i.e. an application with its own lag, its own failure modes, and its own offsets.
**Fix:** design cross-cluster topologies deliberately (Ch. 10), and never assume offset equivalence between clusters.
#### 6.7 The ActiveMQ lesson — coupling telemetry to serving [#67-the-activemq-lesson--coupling-telemetry-to-serving]
**Symptom (historical, and still a live risk):** the broker pauses → client connections back up → **the application can no longer serve user requests**.
**Cause:** a messaging system that applies backpressure into a synchronous serving path.
**Fix / design rule:** producers must be async and bounded; a telemetry path must be allowed to **drop** rather than **block** the request thread. Kafka's push-pull split exists precisely so slow consumers can't stall producers — but you can still recreate the bug yourself with a synchronous, unbounded, blocking producer call in a request handler.
#### 6.8 Colocating Kafka with other memory-hungry apps [#68-colocating-kafka-with-other-memory-hungry-apps]
Foreshadowed here, explicit in Ch. 2: Kafka's read performance comes from the **OS page cache**. Anything else on the box competing for page cache degrades consumer performance. (Details in Ch. 2.)
***
# 1.7 Deployment / monitoring / scaling / backup — what Ch. 1 establishes (/docs/kafka/meet-kafka/deployment-monitoring-scaling-backup)
These are covered properly in Ch. 2, 7, 10, 12, 13, but Ch. 1 sets the frame:
**Deployment shape**
**Scaling axes**
| Axis | Mechanism |
| -------------------------------- | --------------------------------------------------------- |
| Write/read throughput of a topic | **more partitions** (spread across more brokers) |
| Consumer processing throughput | **more consumers in a group** (capped at partition count) |
| Fault tolerance | **higher replication factor** |
| Cluster capacity | **add brokers online** |
| Geographic / isolation | **more clusters + MirrorMaker** |
**"Backup" in Kafka terms — reframe the question.** Kafka is not backed up like a database, and Ch. 1 explains why the question changes shape:
1. **Replication** is the intra-cluster durability mechanism (redundancy, not backup — it won't save you from a bad delete or a logic bug).
2. **Retention** is a *time-bounded* replay buffer — real, but it expires.
3. **Log compaction** is *indefinite* per-key state retention.
4. **MirrorMaker to another cluster/DC** is the DR story.
5. The genuinely durable archive is usually **a sink**: Connect → HDFS/S3, which is also where "replay from the beginning of time" lives.
Kafka's actual durability contract is *"the log is the source of truth for a configured window, and replicated within the cluster during that window."* Anything longer is a sink's job.
**Monitoring, framed by design**
* **Consumer lag** is the master metric — it exists *because* consumers pull and offsets are tracked. Lag vs. retention is the data-loss early warning.
* **Under-replicated partitions** — replication falling behind is the "you are one failure from data loss" signal. The chapter notes network saturation as a common cause: *"should the network interface become saturated, it is not uncommon for cluster replication to fall behind, which can leave the cluster in a vulnerable state."*
* **Offline partitions** — leaderless partitions = unavailable data.
* **Controller count** — exactly one controller should be active.
***
# 1.8 Historical / trivia worth knowing (/docs/kafka/meet-kafka/historical-trivia-worth-knowing)
* Created at LinkedIn; team led by **Jay Kreps** (who had previously built and open-sourced **Voldemort**, a distributed key-value store), with **Neha Narkhede** and later **Jun Rao**.
* Open source on GitHub **late 2010**; Apache incubator **July 2011**; graduated **October 2012**.
* Used at **Netflix, Uber**, and many others in the largest data pipelines in the world.
* LinkedIn maintains **Cruise Control**, **Kafka Monitor**, **Burrow**.
* **Confluent** (founded fall 2014 by Kreps, Narkhede, Rao) released **ksqlDB**, a **schema registry**, and a **REST proxy** under a community license — *note: not strictly open source, as it includes use restrictions.* Confluent also runs **Kafka Summit** (since 2016).
* **The name:** Jay Kreps — *"since Kafka was a system optimized for writing, using a writer's name would make sense. I had taken a lot of lit classes in college and liked Franz Kafka. Plus the name sounded cool for an open source project. So basically there is not much of a relationship."*
***
# 1.3 How does it work internally? (the core object model) (/docs/kafka/meet-kafka/how-does-work-internally)
#### 3.1 Message and batch [#31-message-and-batch]
A **message** is the unit of data — think a DB row or record. To Kafka it is **an opaque array of bytes**. Kafka does not know or care what's inside.
A message optionally has a **key** — also an opaque byte array. The key exists for **partition routing control**:
```txt
partition = hash(key) mod num_partitions
```
Consequence, and it's a big one: **messages with the same key always land in the same partition — *provided the partition count does not change*.** That parenthetical is a production landmine covered in §6.
**Batching.** Messages are written in **batches** — a batch is a collection of messages **all bound for the same topic *and* the same partition**. Why: a network round trip per message is unacceptable overhead.
**The tradeoff is explicit: latency vs throughput.**
* Larger batches → more messages/unit time, but **longer for any individual message to propagate**.
* Batches are typically **compressed** → cheaper transfer and storage, **at the cost of CPU**.
#### 3.2 Schemas — Kafka's deliberate omission [#32-schemas--kafkas-deliberate-omission]
Kafka treats payloads as bytes. The book's position is that you should *nonetheless* impose a schema, and it explains why in terms of **decoupling deploys**, not tidiness.
| Option | Pros | Cons |
| ------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------- |
| JSON / XML | Easy, human-readable | **No robust type handling; no cross-version compatibility** |
| **Avro** (favored) | Compact; **schema separate from payload**; **no codegen needed when schema changes**; strong typing; **backward *and* forward compatible evolution** | Needs a schema repository |
**Why this is an architecture concern, not a serialization preference:**
Without schemas, writing and reading are **tightly coupled**, which forces this deployment dance:
A lockstep, ordered, cross-team deploy for every field addition — exactly the failure that made LinkedIn's XML tracking system break "constantly." With well-defined schemas in a **common repository**, messages are understandable **without coordination**. Schema registry is how you buy back independent deployability.
#### 3.3 Topics and partitions [#33-topics-and-partitions]
**Topic** ≈ a database table or a filesystem folder. **Partition** = a single log.
**The single most-violated fact in Kafka:**
> **Ordering is guaranteed *within a partition only*. There is NO ordering guarantee across a topic.**
If your correctness depends on global ordering, you need either one partition (throughput ceiling of a single broker/consumer) or a key scheme that co-locates all causally-related events in one partition.
**Partitions are the mechanism for both scale and redundancy:**
* **Scale:** each partition can live on a *different server*, so one topic scales horizontally far beyond a single machine.
* **Redundancy:** partitions can be **replicated** across servers, so a copy survives a server failure.
**"Stream"** — usually means a single topic of data, *regardless of partition count* — a single flow producers→consumers. The term is most used in stream processing (Kafka Streams, Samza, Storm) which acts on messages **in real time**, versus offline frameworks (Hadoop) that act on **bulk data later**.
#### 3.4 Producers and consumers [#34-producers-and-consumers]
**Client tiers:**
The advanced clients are **built on** producers and consumers. There is no separate protocol underneath — a useful thing to remember when debugging Connect or Streams.
**Producers** (elsewhere: publishers/writers) create messages to a specific topic.
* Default: **balance messages evenly over all partitions**.
* With a key: hash the key → map to a partition → all messages with that key go to one partition.
* **Custom partitioner:** arbitrary business rules for message→partition mapping.
**Consumers** (elsewhere: subscribers/readers) subscribe to topics and read messages **in produced order per partition**.
**Offsets.** Kafka adds an **offset** to each message at produce time — an integer that **continually increases**. Each message in a partition has a unique offset; the next message has a greater offset — **though not necessarily monotonically greater** (i.e. gaps are legal; do not write code that assumes offset+1). By storing the *next* offset per partition — **typically inside Kafka itself** — a consumer can stop and restart **without losing its place**.
**Consumer groups.**
Two properties fall out:
1. **Horizontal consumer scaling** — add members to divide the partitions.
2. **Automatic failover** — if a member dies, the remaining members **reassign** its partitions.
The ceiling this implies: **group parallelism ≤ partition count.** Extra consumers beyond the partition count sit idle. This is why partition count is a capacity decision made *early* (see Ch. 2).
#### 3.5 Brokers and clusters [#35-brokers-and-clusters]
A **broker** is a single Kafka server. Its job:
1. Receive messages from producers
2. **Assign offsets** to them
3. Write messages to **storage on disk**
4. Service consumer **fetch** requests
Capacity of one broker (hardware-dependent): **thousands of partitions and millions of messages/second.**
**Cluster + controller.**
**Leaders and followers.**
* A partition is **owned by a single broker** → that broker is the **leader**.
* A replicated partition is also assigned to other brokers → **followers**.
* **All producers must connect to the leader to publish.**
* **Consumers may fetch from the leader OR a follower.**
* If the leader's broker fails, **a follower takes over leadership**.
#### 3.6 Retention — the feature that changes everything [#36-retention--the-feature-that-changes-everything]
**Retention** = durable storage of messages for a period of time. Brokers have a default; topics can override.
Two policies:
* **By time** — e.g. 7 days
* **By size** — e.g. 1 GB per partition
When a limit is hit, messages are **expired and deleted**. So retention configuration defines **a minimum amount of data available at any time**.
Per-topic tuning is the point: a **tracking topic** might keep several days; **application metrics** only a few hours.
**Log compaction** — a third mode: **retain only the last message produced with a specific key**. Purpose: **changelog-type data, where only the latest update matters**. This is what makes Kafka usable as a *state store* rather than only a *pipe* (see Ch. 14 on KTables).
#### 3.7 Multiple clusters and MirrorMaker [#37-multiple-clusters-and-mirrormaker]
Three reasons deployments grow to multiple clusters:
1. **Segregation of types of data**
2. **Isolation for security requirements**
3. **Multiple datacenters (disaster recovery)**
**Critical constraint:** *"The replication mechanisms within the Kafka clusters are designed only to work within a single cluster, not between multiple clusters."*
Cross-cluster copying uses **MirrorMaker**, which is architecturally almost insultingly simple: **a Kafka consumer and a Kafka producer, linked together with a queue.** Consume from cluster A, produce to cluster B.
Motivating example from the book: a user edits public profile info; that change must be visible **regardless of which datacenter serves the search results**. Or: collect monitoring data from many sites into one central place where analysis/alerting lives.
"The simple nature of the application belies its power in creating sophisticated data pipelines." Full treatment in Ch. 9/10.
***
# 1. Meet Kafka (/docs/kafka/meet-kafka)
> *Source: Kafka: The Definitive Guide, 2nd Ed. (Shapira, Palino, Sivaram, Petty), Ch. 1*
> **Learning goal for this chapter:** understand *why a log-shaped broker exists at all*. Everything in later chapters (replication, exactly-once, Streams) is a consequence of the design decisions made here.
***
# 1.4 Why Kafka? (the property-by-property case) (/docs/kafka/meet-kafka/kafka-property-property-case)
| Property | What it means | Why it matters |
| ------------------------ | --------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Multiple producers** | Seamlessly handles many producers, many topics or the same topic | Many microservices write page views to **one** topic in a **common format**; consumers get a single unified stream instead of N topics to correlate |
| **Multiple consumers** | Many consumers read the same stream **without interfering with each other** | Explicitly contrasted with queues where **once consumed, a message is gone**. Consumers can *also* form a group to process each message once |
| **Disk-based retention** | Messages on disk with configurable, per-topic retention | Consumers **need not work in real time**. A slow consumer or a traffic burst → **no data loss**. Take a consumer offline for maintenance → producers don't back up, nothing is lost, restart resumes where it left off |
| **Scalable** | 1 broker (PoC) → 3 (dev) → tens/hundreds (prod). **Expansions performed online with no availability impact** | Also: multi-broker clusters survive individual broker failure; raise replication factor to tolerate more simultaneous failures |
| **High performance** | Producers, consumers, and brokers all scale out | **Subsecond latency** from produce to consumer availability, under very large message streams |
| **Platform features** | Kafka **Connect** (source→Kafka, Kafka→sink) and Kafka **Streams** (scalable, fault-tolerant stream processing) | Deliberately **APIs and libraries, not a structured runtime** like YARN — "a solid foundation to build on and flexibility as to where they can be run" |
Note the framing on platform features: these are **not** full platforms. That's a design stance — Kafka gives you libraries you can run under whatever scheduler you already have (k8s, ECS, bare metal) instead of imposing a cluster manager.
#### The ecosystem framing [#the-ecosystem-framing]
Coupled with a **message-schema system**, producers and consumers need **no tight coupling and no direct connections of any sort**. Components can be added and removed as business cases come and go, and **producers need not know who consumes the data or how many consumers exist.**
***
# 1.1 What problem does Kafka solve? (/docs/kafka/meet-kafka/problem-does-kafka-solve)
The problem is **not** "we need a queue." The problem is **N×M point-to-point coupling in a data platform**.
#### The decay sequence (this is the actual origin story) [#the-decay-sequence-this-is-the-actual-origin-story]
**Stage 1 — one producer, one consumer.** You have an app emitting metrics and a dashboard. You open a direct TCP connection. This works. It is the correct engineering decision at this size.
**Stage 2 — requirements multiply.** You want long-term metric analysis, so you add a storage/analysis service. Now the app writes to two places. Then 3 more apps start emitting metrics. Then a coworker wants *pull-based* alerting, so every app also runs an HTTP metrics server. Then more consumers appear pulling from those servers.
The cost here is real and specific, not aesthetic:
* **O(N×M) connections.** Every new consumer requires touching every producer.
* **Every producer knows every consumer.** Adding a consumer is a *deployment of the producer fleet*.
* **Mixed push and pull.** Some paths push, some poll. Different latency, different failure modes, different code.
* **Format drift.** Each pair negotiated its own format. No two are the same.
**Stage 3 — you build a broker.** You put a single service in the middle that accepts metrics from everyone and serves them to anyone.
Congratulations — **you have built a publish/subscribe messaging system.** This is the point the book makes bluntly: pub/sub is not a technology you choose, it's the shape every sufficiently-grown data platform converges on. The only question is whether you build it deliberately or accidentally.
**Stage 4 — the *real* problem.** Meanwhile, a coworker independently built the same thing for *log messages*. Another built it for *user activity tracking*. You now have **three separate pub/sub systems**, each with its own bugs, its own operational runbook, its own scaling limits, and its own on-call rotation.
**This** is the problem Kafka solves: not "move messages," but **one general-purpose, horizontally scalable, durable data backbone that any type of data can flow through**, so the organization stops re-solving pub/sub per data type.
#### The mental model that follows [#the-mental-model-that-follows]
Kafka is described two ways in the book, and both matter:
1. **"A distributed commit log."** A DB or filesystem commit log is a durable, ordered record of every transaction, so state can be *rebuilt deterministically by replay*. Kafka is that, as a service.
2. **"A distributed streaming platform."** Same substrate, framed as continuous data-in-motion rather than storage.
The commit-log framing is the load-bearing one. It explains every property Kafka has:
| Because it's a log… | You get… |
| --------------------------------- | ------------------------------------------------------------- |
| Appends are sequential | Extremely high write throughput on cheap disks |
| Reads are positional (offset) | Many independent readers, no per-consumer state on the broker |
| The log is retained, not consumed | Replay, reprocessing, late consumers, backfill |
| Log is deterministic and ordered | State can be rebuilt; stream processing is well-defined |
| Log can be sharded and copied | Horizontal scale + redundancy |
***
# 1.9 Self-test — can you answer these without looking? (/docs/kafka/meet-kafka/self-test-can-answer)
Why is "we built a pub/sub system" the *inevitable* outcome of a growing data platform, and what specific cost does it eliminate?
LinkedIn had a monitoring system and a tracking system. Give three concrete reasons neither could absorb the other's workload.
ActiveMQ was rejected for two reasons. One was scale. What was the other, and why is it the more important lesson?
What exactly does a key guarantee, and under what condition does that guarantee evaporate?
Why does batching require messages to share *both* topic and partition?
Kafka treats payloads as opaque bytes. Explain the deploy-ordering problem that creates, and how a schema registry fixes it.
State the ordering guarantee precisely. What's the cost of getting global ordering?
Who must producers connect to? Who may consumers connect to? Why the asymmetry?
What's the difference between a delete-retention topic and a compacted topic, and which use case needs which?
Why is Kafka's replication *not* a cross-datacenter solution, and what fills that gap?
Why is "how do I back up Kafka?" the wrong question — and what are the five things that actually cover the intent?
Why are Connect and Streams deliberately libraries rather than a YARN-style platform?
**Next:** [Chapter 2 — Installing Kafka](02-installing-kafka.md) — hardware selection, broker config, OS tuning, and the production concerns that make or break a cluster.
# 1.5 Use cases (with the "why Kafka specifically" for each) (/docs/kafka/meet-kafka/use-cases-kafka-specifically)
**Activity tracking** — *the original LinkedIn use case.* Frontends generate messages about user actions: passive (page views, click tracking) or complex (profile updates). Published to topics; consumed by backends generating reports, feeding ML, updating search results.
**Messaging** (e.g. user notifications/emails) — the interesting part is *why* Kafka helps. Producing apps emit "notify this user" without knowing formatting or delivery. **One** application reads all of them and handles consistently:
* **Formatting/decorating** with a common look and feel
* **Collecting multiple messages into a single notification**
* **Applying the user's delivery preferences**
The win: avoids duplicating this logic in every app, **and enables aggregation that would not otherwise be possible** (you can't batch a user's notifications if each app sends its own emails).
**Metrics and logging** — where "multiple applications producing the same type of message" shines. Apps publish metrics to a topic; consumed by monitoring/alerting **and** by an offline system like Hadoop for long-term analysis (growth projections). Logs route to Elasticsearch or security analysis. **Key benefit: when the destination system changes (time to replace the log store), you do not alter the frontend applications or the aggregation mechanism.** The producers are insulated from sink churn.
**Commit log** — publish DB changes to Kafka; apps monitor the stream for live updates. Uses:
* Replicate DB updates to a remote system
* Consolidate changes from multiple applications into a single database view
* **Durable retention buffers the changelog** → replay after a consumer-side failure
* **Log-compacted topics** give longer retention by keeping only one change per key
(This is CDC — see Ch. 9 for Connect-based implementations.)
**Stream processing** — applications giving map/reduce-like functionality but **on data in real time, as quickly as messages are produced**, versus Hadoop's hours/days aggregation windows. Tasks: counting metrics, repartitioning messages for efficient downstream processing, transforming messages using data from multiple sources.
***
# 1.2 Why wasn't something else enough? (/docs/kafka/meet-kafka/wasn-t-something-else)
This is answered *concretely* by LinkedIn's history, which is the honest version of the story.
#### What LinkedIn had before Kafka [#what-linkedin-had-before-kafka]
**System A — monitoring/metrics.** Custom collectors + open-source storage/presentation. Its faults:
* **Poll-based collection** → you get data at the poller's convenience, not the event's.
* **Large intervals between metrics** → low resolution, blind spots.
* **No self-service** — application owners could not manage their own metrics.
* **High-touch** — human intervention for routine tasks.
* **Inconsistent** — the same measurement had different metric names in different systems.
**System B — user activity tracking.** Frontends periodically POSTed **batches of XML** to an HTTP service; batches were moved to offline platforms and parsed there. Its faults:
* **XML parsing was computationally expensive**, and the formatting was inconsistent.
* **Schema changes broke it constantly.**
* **Changing what you tracked required coordinated frontend + offline work** — a cross-team project for a new field.
* **Hourly batching** → *structurally* incapable of real time.
#### Why they couldn't just merge the two [#why-they-couldnt-just-merge-the-two]
This is the crux, and it's the best "why not something else" argument in the book:
| | Monitoring system | Tracking system |
| ---------- | -------------------------------- | -------------------------- |
| Model | **Pull** (polling) | **Push** (frontends post) |
| Cadence | Interval-based | Hourly batches |
| Data shape | Metrics-oriented | Activity/event-oriented |
| Robustness | Clunky but tolerable for metrics | Too fragile for metrics |
| Real-time? | Sort of | No — batch by construction |
The monitoring backend couldn't take tracking data (wrong data model, pull model incompatible with push). The tracking backend couldn't take metrics (too fragile, batch-oriented — useless for alerting).
**And they desperately needed them joined.** The high-value question was *"how does this specific type of user activity affect application performance?"* — a correlation across both datasets. With hours of batch delay, a drop in a user activity type (a strong signal that the serving app is broken) surfaced hours late. The business cost of the architecture was **slow incident response**.
#### Why not off-the-shelf? [#why-not-off-the-shelf]
They did evaluate existing open source properly. **ActiveMQ was prototyped and rejected**, for two specific reasons:
1. **It could not handle the scale** at the time.
2. **The brokers would pause.** This is the fatal one. When an ActiveMQ broker paused, it **backed up the client connections**, which **interfered with the applications' ability to serve user requests**.
Read that failure mode carefully, because it's the single most important operational lesson in the chapter: **a telemetry system that can stall its producers has coupled your observability pipeline to your revenue path.** Metrics/tracking must be a system that *cannot* apply backpressure into the serving tier. That requirement — "the broker must absorb, never stall the producer" — drives Kafka's whole design: sequential appends, batching, page-cache reads, disk retention as a buffer instead of memory-bound queues.
#### The design goals they wrote down [#the-design-goals-they-wrote-down]
1. **Decouple producers and consumers using a push-pull model.** Producers *push*; consumers *pull* at their own rate. This is the direct fix for the ActiveMQ pause problem — a slow consumer cannot stall a producer, because consumption is pull-based off a durable log.
2. **Persist message data inside the messaging system, to allow multiple consumers.** Persistence is not for durability alone; it's what makes *many independent consumers* possible.
3. **Optimize for high throughput.**
4. **Allow horizontal scaling as data streams grow.**
The result: "an interface typical of messaging systems, but **a storage layer more like a log-aggregation system**." That sentence is the whole product.
#### Did it work? [#did-it-work]
* Combined with **Apache Avro** for serialization, it handled metrics *and* activity tracking at **billions of messages/day**.
* By **February 2020**, LinkedIn: **>7 trillion messages produced** and **>5 PB consumed daily**.
* Open sourced on GitHub **late 2010** → Apache incubator **July 2011** → graduated **October 2012**.
***
# 13.10 What actually breaks in production — Ch. 13 consolidated (/docs/kafka/monitoring-kafka/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | --------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------------- |
| 1 | **Kafka is down and no alert fired** | The monitoring pipeline **runs on Kafka** | Separate monitoring system for Kafka, or cross-DC metric shipping |
| 2 | **A JMX port became an RCE vector** | Remote JMX enabled without security — *"JMX not only allows a view into the state, it also allows CODE EXECUTION"* | In-process agent (Jolokia/MX4J); keep remote JMX disabled |
| 3 | **All broker metrics look fine but clients can't connect** | Broker metrics are **subjective** | Synthetic clients (Xinfra Monitor) and client-side metrics |
| 4 | **Team ignores alerts** | **Alert fatigue** from too many metric-based alerts with hand-tuned thresholds | SLO/burn-rate alerting; "Check Engine light" principle |
| 5 | **URP alerts fire constantly for benign reasons** | *"the URP metric can frequently be NONZERO for benign reasons"* | **The book retracts its own prior advice** — don't alert on URP; use it for diagnosis |
| 6 | **SLO breached before anyone noticed** | Alerted on the SLO itself, which is a lagging, week-scale indicator | **Burn-rate** alerting (0.1% → ticket at 0.4% → page at 2%) |
| 7 | **Can't tell how many requests breached the latency target** | Used **quantile** metrics as SLIs | Counters/buckets of events inside vs outside the threshold |
| 8 | **Took on an SLO you can't control (data freshness/correctness)** | Scope creep from customers | *"DO NOT agree to support SLOs that you are not responsible for"* |
| 9 | **Monitoring system overwhelmed by tens of thousands of series** | Collecting *debugging* metrics continuously | Debugging metrics only need to be **available**, not collected |
| 10 | **Alert can't distinguish "broker down" from "monitoring down"** | Relied on **stale metrics** as the health check | Add an external port-connect health check |
| 11 | **One broker's slow disk destroyed cluster-wide producer throughput** | **"One bad egg"** — every producer talks to every broker, so one slow broker back-pressures all of them | Treat single-broker outliers as urgent; SMART + IPMI monitoring |
| 12 | **Write latency tripled with nothing "broken"** | **RAID controller BBU failed** → onboard cache silently disabled | Monitor the disk controller and BBU, *"whether you are using hardware RAID or not"* |
| 13 | **Broker running on reduced memory/CPU with no failure** | **Soft hardware failures** — bad memory segment or CPU bypassed by the OS | IPMI health, `dmesg` kernel ring buffer |
| 14 | **Disk filled and the broker died abruptly** | *"brokers operate properly RIGHT UP UNTIL THE DISK IS FILLED, and then this disk WILL FAIL ABRUPTLY"* | Alert on free space **and inodes** as capacity, not performance |
| 15 | **Intermittent network problems** | Bad cable/connector, speed/duplex mismatch, undersized OS buffers | *"the number of ERRORS on the network interfaces — if the error count is increasing, there is probably an unaddressed issue"* |
| 16 | **One broker behaves differently from the rest** | Config drift | Chef/Puppet-style configuration management |
| 17 | **A monitoring agent starved the broker** | Another process consuming CPU/memory — *"a process that is SUPPOSED to be running, such as a monitoring agent, but is having problems"* | `top`; don't colocate (Ch. 2) |
| 18 | **Cluster imbalanced even though partitions are evenly spread** | *"Kafka does not detect issues such as HOT PARTITIONS"*, and leadership doesn't return after a restart | Preferred replica election **first**, then Cruise Control / `kafka-assigner` |
| 19 | **Two brokers both claim to be controller** | *"a controller thread that should have exited has become STUCK"* | Restart both; expect **controlled shutdown to fail** — force stop |
| 20 | **No broker is the controller; topic creation hangs** | e.g. **network partition from ZooKeeper** | Fix the root cause, then **restart all brokers to reset controller state** |
| 21 | **Controller queue size grows and never drops** | Controller stuck | Move the controller — but *"there will often be problems performing a controlled shutdown of ANY broker"*; Ch. 12's znode deletion is the alternative |
| 22 | **Broker latency high; requests queueing** | **Request handler idle ratio \< 10%** | Threads = CPU count (incl. hyperthreads); then reduce load or add brokers |
| 23 | **Request handler threads burning CPU unnecessarily** | Old clients force **decompress → validate → recompress**, historically behind a **synchronous lock** | *"ONE OF THE SINGLE LARGEST PERFORMANCE IMPROVEMENTS"* — get all clients and brokers to message format 0.10+ |
| 24 | **"Bytes out equals bytes in with no consumers"** | Bytes-out **includes replica fetch traffic** | Expected at RF 2. Formula: `in × ((RF−1) + consumer_groups)` |
| 25 | **Looking for a "messages out" metric** | Doesn't exist — the broker never expands batches on the read path | Use fetch **request** rate |
| 26 | **Alerted on `MeanRate` and never saw a spike** | `MeanRate` is averaged since broker start | `OneMinuteRate` for spikes; 5/15-minute for balance |
| 27 | **A broker leads 0 partitions after recovering** | Leadership isn't reclaimed automatically | Alert on **leader count** / `LeaderCount ÷ PartitionCount ≈ 1/RF` |
| 28 | **Offline partitions metric reads 0 everywhere** | *"ONLY provided by the broker that is the controller"* | Aggregate with `max`, or scrape the controller specifically |
| 29 | **Latency is high; no idea which phase** | Only collected `TotalTimeMs` | Collect all seven timings; map to queue → local → remote → throttle → response |
| 30 | **Fetch-latency alerts flap** | Governed by `fetch.min.bytes` / `fetch.max.wait.ms` | Baseline **Produce** p99.9 instead |
| 31 | **One partition is far bigger than its siblings** | **Hot key** — uneven key distribution | Per-partition `Size`; then a custom partitioner (Ch. 3 §9.5) |
| 32 | **`LogEndOffset − LogStartOffset` ≠ message count** | **Log compaction** creates offset gaps | Don't compute counts from offsets |
| 33 | **Broker hit "too many open files"** | FDs for every log segment **and** every network connection; *"a problem closing network connections properly could cause the broker to RAPIDLY EXHAUST"* | Monitor `OpenFileDescriptorCount` vs `MaxFileDescriptorCount` |
| 34 | **Broker dropped out of the cluster with no restart** | **Long GC pause** → ZK session expiry (Ch. 6 §1) | GC `CollectionTime`/`CollectionCount` + `LastGcInfo.duration` |
| 35 | **Misread system load as a percentage** | Load average is a **count of runnable + uninterruptible-sleep threads**; 100% == number of CPUs | Divide by CPU count |
| 36 | **Compacted topics grew without bound; tombstones never processed** | *"failure in compaction of a SINGLE partition can HALT the log compaction threads ENTIRELY, AND SILENTLY"* — **and no metric exposes it** | Enable `kafka.log.LogCleaner`, `kafka.log.Cleaner`, `kafka.log.LogCleanerManager` at **DEBUG by default** |
| 37 | **Broker filled its disk with its own logs** | Request logger left at DEBUG/TRACE | Enable only while debugging |
| 38 | **Producer silently dropping messages** | Retries exhausted | **`record-error-rate` should ALWAYS be zero — alert on it** |
| 39 | **Producer back-pressuring the application** | Rising `request-latency-avg` | Baseline it and alert above |
| 40 | **Can't tell which broker a producer is struggling with** | Only looked at overall metrics | **Per-broker `request-latency-avg`** — the only stable per-broker metric |
| 41 | **Consumer lag alert missed a stalled consumer** | `records-lag-max` shows **one partition** and **requires a working consumer** | External lag monitoring — **Burrow** |
| 42 | **Lag thresholds unmaintainable at scale** | 100,000 partitions × per-partition thresholds | Burrow's threshold-free, progress-based status |
| 43 | **False alerts on low consumer throughput** | Alerting on `records-consumed-rate` minimums assumes the **producer** is healthy | Don't; Kafka deliberately decouples the two |
| 44 | **Consumer group pauses repeatedly** | Rebalance storms | **`sync-rate` should be \~0**; check `sync-time-avg` |
| 45 | **Offset commits slow, consumer throughput drops** | Commits are produce requests to one broker (Ch. 7 §5.2) | Baseline and alert on **`commit-latency-avg`** |
| 46 | **Load uneven inside a consumer group** | Assignor imbalance (`RangeAssignor`, Ch. 4 §6.5) | Compare **`assigned-partitions`** across instances |
| 47 | **Client mysteriously slow; broker healthy; no errors** | **Quota throttling — the broker returns NO error code** | `produce-throttle-time-avg` / `fetch-throttle-time-avg` — monitor them **even before enabling quotas** |
| 48 | **Can't tell whether it's the client, the network, or Kafka** | No end-to-end view | **Xinfra Monitor** — produce+consume across every broker, measuring availability and total latency |
***
# 13.4 The broker metrics reference (/docs/kafka/monitoring-kafka/broker-metrics-reference)
#### 4.1 Active controller count [#41-active-controller-count]
> *"**At all times, ONLY ONE broker should be the controller, and ONE broker MUST ALWAYS be the controller.**"*
#### 4.2 Controller queue size [#42-controller-queue-size]
> *"indicates **how many requests the controller is currently waiting to process** for the brokers. ... **SPIKES IN THE METRIC ARE TO BE EXPECTED, BUT if this value CONTINUOUSLY INCREASES, OR STAYS STEADY AT A HIGH VALUE AND DOES NOT DROP, IT INDICATES THAT THE CONTROLLER MAY BE STUCK.**"*
>
> ► **FIX:** *"you will need to **MOVE THE CONTROLLER to a different broker, which requires SHUTTING DOWN the broker that is currently the controller. HOWEVER, WHEN THE CONTROLLER IS STUCK, THERE WILL OFTEN BE PROBLEMS PERFORMING A CONTROLLED SHUTDOWN OF ANY BROKER.**"*
*(Ch. 12 §9.1's `/admin/controller` znode deletion is the alternative that doesn't require a shutdown.)*
#### 4.3 💡 Request handler idle ratio — the load metric [#43--request-handler-idle-ratio--the-load-metric]
**Why this pool and not the network threads:**
> *"**The NETWORK threads** are responsible for reading and writing data to the clients across the network. **This does not require significant processing, which means that EXHAUSTION OF THE NETWORK THREADS IS LESS OF A CONCERN.** The **REQUEST HANDLER threads**, however, are responsible for **servicing the client request itself, WHICH INCLUDES READING OR WRITING THE MESSAGES TO DISK.** As such, **as the brokers get more heavily loaded, THERE IS A SIGNIFICANT IMPACT ON THIS THREAD POOL.**"*
**The thresholds — memorize these:**
> ### INTELLIGENT THREAD USAGE [#intelligent-thread-usage]
>
> *"While it may seem like you will need hundreds of request handler threads, **in reality YOU DO NOT NEED TO CONFIGURE ANY MORE THREADS THAN YOU HAVE CPUs in the broker.** Apache Kafka is **very smart about the way it uses the request handlers, making sure to OFFLOAD TO PURGATORY those requests that will take a long time to process.** This is used, for example, **when requests are being quoted or when MORE THAN ONE ACKNOWLEDGMENT of produce requests is required.**"*
*(Ch. 6 §5.1's purgatory, and Ch. 6 §5.3's `acks=all` path. Because slow waits go to purgatory, the handler threads are never blocked waiting — so CPU count is the right sizing.)*
**Two causes of high utilization besides an undersized cluster:**
*(This is the flip side of Ch. 6 §6.5's down-conversion problem — the relative-offset design in the v2 format is precisely what eliminates broker-side recompression.)*
#### 4.4 The rate metrics — and how to read their seven attributes [#44-the-rate-metrics--and-how-to-read-their-seven-attributes]
**Every rate metric has seven attributes:**
> ⚠️ *"**Make sure to use the metrics APPROPRIATELY, or you will end up with a FLAWED VIEW of the broker.**"*
##### Bytes out — and why it can be 6× bytes in [#bytes-out--and-why-it-can-be-6-bytes-in]
> *"The outbound bytes rate **may scale DIFFERENTLY than the inbound bytes rate**, thanks to Kafka's capacity to handle multiple consumers with ease. **There are many deployments of Kafka where the outbound rate can easily be SIX TIMES the inbound rate!** This is why it is important to **observe and trend the outbound bytes rate SEPARATELY.**"*
> ### ⚠️ REPLICA FETCHERS INCLUDED [#️-replica-fetchers-included]
>
> *"The outbound bytes rate **ALSO INCLUDES THE REPLICA TRAFFIC.** This means that **if all topics are configured with a replication factor of 2, YOU WILL SEE A BYTES OUT RATE EQUAL TO THE BYTES IN RATE WHEN THERE ARE NO CONSUMER CLIENTS.** If you have ONE consumer client reading all the messages, **the bytes out rate will be TWICE the bytes in rate. THIS CAN BE CONFUSING when looking at the metrics if you're not aware of what is counted.**"*
##### Messages in — and why there's no "messages out" [#messages-in--and-why-theres-no-messages-out]
> ### WHY NO MESSAGES OUT? [#why-no-messages-out]
>
> *"when messages are consumed, **the broker just sends the NEXT BATCH to the consumer WITHOUT EXPANDING IT to find out how many messages are inside. Therefore, THE BROKER DOESN'T REALLY KNOW HOW MANY MESSAGES WERE SENT OUT.** The only metric that can be provided is **the number of FETCHES per second, which is a REQUEST RATE, not a messages count.**"*
*(This is zero-copy showing up in the metrics: the broker never parses the batch on the read path, so it can't count records. Ch. 6 §5.4.)*
**Messages-in uses:** *"a growth metric as a **different measure of producer traffic.** It can also be used **in conjunction with the bytes in rate to determine an AVERAGE MESSAGE SIZE.**"*
#### 4.5 Partition count and leader count [#45-partition-count-and-leader-count]
```txt
Partition count kafka.server:type=ReplicaManager,name=PartitionCount
Leader count kafka.server:type=ReplicaManager,name=LeaderCount
```
**Partition count:** *"the total number of partitions assigned to that broker. **This includes EVERY replica the broker has, regardless of whether it is a leader or follower.** Monitoring this is often more interesting in a cluster that has **automatic topic creation enabled**, as that can leave the creation of topics **outside of the control of the person running the cluster.**"*
**Leader count — and why it deserves an alert:**
> *"**It is MUCH MORE IMPORTANT to check the leader count on a regular basis, POSSIBLY ALERTING ON IT, as it will indicate when the cluster is IMBALANCED EVEN IF THE NUMBER OF REPLICAS ARE PERFECTLY BALANCED in count and size.** This is because **a broker can drop leadership for many reasons, such as a ZOOKEEPER SESSION EXPIRATION, and IT WILL NOT AUTOMATICALLY TAKE LEADERSHIP BACK once it recovers** (except with automatic leader rebalancing). In these cases, this metric will show **fewer leaders, or often ZERO**, which indicates that **you need to run a preferred replica election.**"*
> 💡 **The derived metric worth building:**
> *"use it **along with the partition count to show A PERCENTAGE of partitions that the broker is the leader for.** In a well-balanced cluster using **replication factor 2, all brokers should be leaders for approximately 50%.** If the replication factor is **3, this percentage drops to 33%.**"*
#### 4.6 Offline partitions — the "site down" metric [#46-offline-partitions--the-site-down-metric]
> ⚠️ *"**This measurement is ONLY provided by the broker that is THE CONTROLLER for the cluster (all other brokers will report 0)**"* — so you must aggregate with `max`, not `sum` or `avg`, or scrape only the controller.
**Two causes:**
> *"In a production Kafka cluster, an offline partition **may be impacting the producer clients, LOSING MESSAGES or CAUSING BACK PRESSURE in the application. THIS IS MOST OFTEN A 'SITE DOWN' TYPE OF PROBLEM AND WILL NEED TO BE ADDRESSED IMMEDIATELY.**"*
#### 4.7 Request metrics — the per-phase latency breakdown [#47-request-metrics--the-per-phase-latency-breakdown]
> As of **version 2.5.0**, metrics exist for **\~50 request types**, including: `AddOffsetsToTxn`, `AddPartitionsToTxn`, `AlterConfigs`, `AlterPartitionReassignments`, `ApiVersions`, `ControlledShutdown`, `CreateAcls`, `CreatePartitions`, `CreateTopics`, `DeleteRecords`, `DeleteTopics`, `DescribeGroups`, `ElectLeaders`, `EndTxn`, **`Fetch`**, **`FetchConsumer`**, **`FetchFollower`**, `FindCoordinator`, `Heartbeat`, `InitProducerId`, `JoinGroup`, `LeaderAndIsr`, `ListOffsets`, `Metadata`, `OffsetCommit`, `OffsetFetch`, **`Produce`**, `SaslAuthenticate`, `StopReplica`, `SyncGroup`, `TxnOffsetCommit`, `UpdateMetadata`, `WriteTxnMarkers`, and more.
**Eight metrics per request type** — seven timings plus a rate. For `Fetch`:
**💡 The seven phases map exactly onto Ch. 6 §5.1's threading model — this is how you localize a latency problem:**
**Reading it diagnostically:**
**Attributes per timing metric:** `Count`, `Min`, `Max`, `Mean`, `StdDev`, and percentiles `50th`, `75th`, `95th`, `98th`, `99th`, `999th`.
> ⚠️ *"The metrics are **ALL CALCULATED SINCE THE BROKER WAS STARTED**, so keep that in mind when looking at metrics that do not change for long periods; **the LONGER your broker has been running, the MORE STABLE the numbers will be.**"*
> ### WHAT IS A PERCENTILE? [#what-is-a-percentile]
>
> *"A 99th percentile measurement tells us that **99% of all values in the sample group are LESS THAN the value of the metric.** This means that **1% of the values are GREATER.** A common pattern is to **view the AVERAGE value AND the 99% or 99.9% value.** In this way, you can understand **how the average request performs AND what the OUTLIERS are.**"*
**What to actually collect:**
**On alert thresholds:**
> \*"**the timing metrics can be DIFFICULT.** The timing for a **Fetch** request can **vary wildly** depending on **settings on the client for how long it will wait for messages, how busy the particular topic is, and the speed of the network connection.**
>
> 💡 *"It can be very useful, however, to **develop a BASELINE value for the 99.9th percentile for at least the TOTAL TIME, ESPECIALLY FOR PRODUCE REQUESTS, and alert on this. Much like the URP metric, A SHARP INCREASE IN THE 99.9TH PERCENTILE FOR PRODUCE REQUESTS CAN ALERT YOU TO A WIDE RANGE OF PERFORMANCE PROBLEMS.**"*
**Produce, not Fetch** — because Fetch latency is dominated by client config (`fetch.max.wait.ms`), while Produce latency reflects the broker's actual health.
#### 4.8 Topic and partition metrics [#48-topic-and-partition-metrics]
> *"In larger clusters these can be **numerous, and it may not be possible to collect all of them** as a matter of normal operations. **However, they are quite useful for debugging specific issues with a client.** For example, the topic metrics can be used to **identify a specific topic that is causing a large increase in traffic.** It also may be important to provide these metrics **so that USERS of Kafka are able to access them.**"*
**Per-topic (Table 13-16)** — all `kafka.server:type=BrokerTopicMetrics,...,topic=TOPICNAME`:
> *"almost certainly metrics that you will **NOT want to set up monitoring and alerts for. They are useful to PROVIDE TO CLIENTS, however, so that they can evaluate and debug their own usage of Kafka.**"*
**Per-partition (Table 13-17)** — `kafka.log:type=Log,...,topic=TOPICNAME,partition=0`:
| Metric | Use |
| ------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`Size`** | *"the amount of data (in bytes) currently being retained on disk for the partition."* Combined → *"the amount of data retained for a single topic, which can be useful in **ALLOCATING COSTS for Kafka to individual clients.**"* 💡 *"**A DISCREPANCY BETWEEN THE SIZE OF TWO PARTITIONS FOR THE SAME TOPIC CAN INDICATE A PROBLEM WHERE THE MESSAGES ARE NOT EVENLY DISTRIBUTED ACROSS THE KEY** that is being used when producing."* |
| **`NumLogSegments`** | *"the number of log-segment files on disk. Useful along with partition size for resource tracking."* |
| **`LogEndOffset` / `LogStartOffset`** | *"the highest and lowest offsets for messages in that partition."* ⚠️ *"**the difference between these two numbers DOES NOT NECESSARILY INDICATE THE NUMBER OF MESSAGES, as LOG COMPACTION can result in 'MISSING' OFFSETS.**"* Use case: *"a more granular mapping of timestamp to offset, allowing consumers to roll back to a specific time (though this is **less important with time-based index searching, introduced in Kafka 0.10.1**)."* |
**Partition `Size` skew is the hot-key detector** — this is how you find Ch. 3 §9.5's "Banana" problem in production.
> ### UNDER-REPLICATED PARTITION METRICS (per-partition) [#under-replicated-partition-metrics-per-partition]
>
> *"there is a per-partition metric... **In general, this is NOT VERY USEFUL in day-to-day operations, as there are TOO MANY METRICS to gather and watch. IT IS MUCH EASIER TO MONITOR THE BROKER-WIDE URP COUNT AND THEN USE THE COMMAND-LINE TOOLS to determine the specific partitions.**"*
***
# 13.7 Client monitoring (/docs/kafka/monitoring-kafka/client-monitoring)
#### 7.1 Producer metrics [#71-producer-metrics]
> *"The Kafka producer client has **greatly compacted the metrics** available by making them available as attributes on **a small number of JMX MBeans.** In contrast, the previous (unsupported) version used **a larger number of MBeans but had MORE DETAIL** (a greater number of percentile measurements and different moving averages). As a result, **the overall number of metrics covers a WIDER SURFACE AREA, but IT CAN BE MORE DIFFICULT TO TRACK OUTLIERS.**"*
**Three beans (Table 13-19):**
> *"while we will discuss several metrics that are **averages (ending in `-avg`)**, there are also **maximum values for each metric (ending in `-max`) that have LIMITED USEFULNESS.**"*
##### The two producer metrics to ALERT on [#the-two-producer-metrics-to-alert-on]
*(Ch. 7 §6.3 named exactly these two: *"the two metrics most important for reliability are error-rate and retry-rate per record."* Note `record-error-rate` also counts benign idempotent-duplicate rejections — Ch. 8 §2.3.)*
##### The three views of traffic volume [#the-three-views-of-traffic-volume]
> *"**A single REQUEST contains one or more BATCHES. A single BATCH contains one or more MESSAGES. And, of course, each MESSAGE is made up of some number of BYTES.** These metrics are all useful to have on an application dashboard."*
##### The size metrics [#the-size-metrics]
##### 💡 `record-queue-time-avg` — the linger.ms tuning metric [#-record-queue-time-avg--the-lingerms-tuning-metric]
> *"the **average amount of time, in milliseconds, that a single message WAITS IN THE PRODUCER, AFTER THE APPLICATION SENDS IT, BEFORE IT IS ACTUALLY PRODUCED TO KAFKA.**"*
This is the metric that tells you whether `linger.ms` is actually costing you what you think (Ch. 3 §5).
##### Per-broker and per-topic producer metrics [#per-broker-and-per-topic-producer-metrics]
> *"useful for **debugging problems in some cases, but they are NOT metrics that you are going to want to review on an ongoing basis.** All of the attributes... are **the same** as the overall producer beans."*
**Per-broker — one metric stands out:**
> *"The **most useful** metric provided by the per-broker producer metrics is **`request-latency-avg`.** This is because **this metric will be MOSTLY STABLE (given stable batching) and CAN STILL SHOW A PROBLEM WITH CONNECTIONS TO A SPECIFIC BROKER.** The other attributes... **tend to VARY depending on WHAT PARTITIONS EACH BROKER IS LEADING.** This means that what these measurements 'should' be **can quickly change, depending on the state of the Kafka cluster.**"*
**Per-broker `request-latency-avg` is how you find the "one bad egg" from the client side.**
**Per-topic:**
> *"only useful for producers working with **MORE than one topic**... only usable on a regular basis if the producer is **NOT working with a LOT of topics.** For example, **a MirrorMaker could be producing HUNDREDS, OR THOUSANDS, of topics. It is difficult to review all of those metrics, and NEARLY IMPOSSIBLE TO SET REASONABLE ALERT THRESHOLDS on them.**"*
> *"**`record-send-rate` and `record-error-rate` can be used to ISOLATE DROPPED MESSAGES TO A SPECIFIC TOPIC** (or validated to be across all topics)."* Plus a **`byte-rate`** per topic.
#### 7.2 Consumer metrics [#72-consumer-metrics]
**Five beans (Table 13-20):**
> *"the **overall consumer metric bean is LESS USEFUL for us** because the metrics of interest are located in **the FETCH MANAGER beans instead.** The overall consumer bean has metrics regarding **lower-level network operations**, but the fetch manager bean has metrics regarding **bytes, request, and record rates.**"*
> 💡 *"Unlike the producer client, **the metrics provided by the CONSUMER are USEFUL TO LOOK AT BUT NOT USEFUL FOR SETTING UP ALERTS ON.**"*
##### `fetch-latency-avg` — and why it's a poor alert [#fetch-latency-avg--and-why-its-a-poor-alert]
> *"As with `request-latency-avg` in the producer, this tells us how long fetch requests take. **THE PROBLEM WITH ALERTING ON THIS METRIC IS THAT THE LATENCY IS GOVERNED BY THE CONSUMER CONFIGURATIONS `fetch.min.bytes` AND `fetch.max.wait.ms`. A SLOW TOPIC WILL HAVE ERRATIC LATENCIES, as sometimes the broker will respond quickly (when there are messages available), and sometimes it will not respond for `fetch.max.wait.ms` (when there are none).** When consuming topics that have **more regular, and abundant, message traffic**, this metric may be more useful."*
##### ⚠️ WAIT! NO LAG? [#️-wait-no-lag]
> \*"The best advice for all consumers is that **you MUST monitor the consumer lag.** So why do we **not recommend monitoring the `records-lag-max` attribute** on the fetch manager bean? This metric shows **the current lag for THE PARTITION THAT IS THE MOST BEHIND.**
>
> **The problem is TWOFOLD:**
>
> 1. **IT ONLY SHOWS THE LAG FOR ONE PARTITION**, and
> 2. **IT RELIES ON PROPER FUNCTIONING OF THE CONSUMER.**
>
> **If you have NO OTHER OPTION, use this attribute for lag and set up alerting for it. BUT THE BEST PRACTICE IS TO USE EXTERNAL LAG MONITORING.**"\* (→ §8)
##### Traffic metrics — and a warning about alerting on them [#traffic-metrics--and-a-warning-about-alerting-on-them]
```txt
bytes-consumed-rate bytes per second consumed by this client instance
records-consumed-rate messages per second
```
> ⚠️ *"Some users set **MINIMUM thresholds** on these metrics for alerting so they are notified **if the consumer is NOT DOING ENOUGH WORK. YOU SHOULD BE CAREFUL WHEN DOING THIS, HOWEVER. Kafka is intended to DECOUPLE the consumer and producer clients... The rate at which the consumer is able to consume messages IS OFTEN DEPENDENT ON WHETHER OR NOT THE PRODUCER IS WORKING CORRECTLY, so monitoring these metrics ON THE CONSUMER MAKES ASSUMPTIONS ABOUT THE STATE OF THE PRODUCER. THIS CAN LEAD TO FALSE ALERTS.**"*
**Relationship metrics:**
##### 💡 Consumer coordinator metrics — the rebalance detector [#-consumer-coordinator-metrics--the-rebalance-detector]
> *"The **BIGGEST problem that consumers can run into due to coordinator activities is A PAUSE IN CONSUMPTION WHILE THE CONSUMER GROUP SYNCHRONIZES.** This is when the consumer instances negotiate which partitions will be consumed by which client. **Depending on the number of partitions, THIS CAN TAKE SOME TIME.**"*
| Metric | Meaning and use |
| ------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`sync-time-avg`** | *"the average amount of time, in milliseconds, that the sync activity takes"* |
| **`sync-rate`** | *"the number of group syncs that happen every second. **FOR A STABLE CONSUMER GROUP, THIS NUMBER SHOULD BE ZERO MOST OF THE TIME.**"* |
| **`commit-latency-avg`** | *"the average amount of time that offset commits take. **You should monitor this value JUST AS YOU WOULD THE REQUEST LATENCY IN THE PRODUCER.** It should be possible to establish a baseline and **set reasonable thresholds for alerting.**"* — because *"these commits are essentially just **produce requests** (though they have their own request type), in that the offset commit is **a message produced to a special topic**"* |
| **`assigned-partitions`** | *"a count of the number of partitions that the consumer client (as a single instance in the group) has been assigned. **This is helpful because, WHEN COMPARED TO THIS METRIC FROM OTHER CONSUMER CLIENTS IN THE GROUP, IT IS POSSIBLE TO SEE THE BALANCE OF LOAD ACROSS THE ENTIRE CONSUMER GROUP.** We can use this to identify **imbalances that might be caused by problems in the algorithm used by the consumer coordinator.**"* |
**`sync-rate` is the single best rebalance-storm detector** (Ch. 4 §2), and **`assigned-partitions` across instances** is how you detect an unbalanced assignor like `RangeAssignor` (Ch. 4 §6.5).
#### 7.3 Quota throttling metrics [#73-quota-throttling-metrics]
> *"**The Kafka broker DOES NOT USE ERROR CODES in the response to indicate that the client is being throttled. This means that IT IS NOT OBVIOUS TO THE APPLICATION THAT THROTTLING IS HAPPENING WITHOUT MONITORING THE METRICS.**"*
> 💡 *"Quotas are **not enabled by default**, but **IT IS SAFE TO MONITOR THESE METRICS IRRESPECTIVE of whether you are currently using quotas. Monitoring them is a good practice as THEY MAY BE ENABLED AT SOME POINT IN THE FUTURE, and IT'S EASIER TO START WITH MONITORING THEM AS OPPOSED TO ADDING METRICS LATER.**"*
**This is the "my client is mysteriously slow and the broker looks fine" diagnosis.** Throttling is invisible without these two metrics.
***
# 13.9 End-to-end monitoring (/docs/kafka/monitoring-kafka/end-end-monitoring)
> *"Consumer and producer clients **have metrics that CAN INDICATE that there might be a problem with the Kafka cluster, BUT THIS CAN BE A GUESSING GAME as to whether increased latency is due to A PROBLEM WITH THE CLIENT, THE NETWORK, OR KAFKA ITSELF.** In addition, it means that **if you are responsible for running the Kafka cluster, AND NOT THE CLIENTS, YOU WOULD NOW HAVE TO MONITOR ALL OF THE CLIENTS AS WELL.**"*
**The two questions you actually need answered:**
```txt
① Can I PRODUCE messages to the Kafka cluster?
② Can I CONSUME messages from the Kafka cluster?
```
**Xinfra Monitor (formerly Kafka Monitor):**
> *"open sourced by the Kafka team at LinkedIn, \[it] **continually produces and consumes data from a topic that is SPREAD ACROSS ALL BROKERS in a cluster.** It measures **the AVAILABILITY of both produce and consume requests ON EACH BROKER, as well as the TOTAL PRODUCE-TO-CONSUME LATENCY.**"*
> *"In an ideal world, you would be able to monitor this **for every topic individually. However, in most situations IT IS NOT REASONABLE TO INJECT SYNTHETIC TRAFFIC INTO EVERY TOPIC.** We can, however, at least provide those answers **for every BROKER in the cluster.**"*
> *"This type of monitoring is **INVALUABLE to be able to EXTERNALLY VERIFY that the Kafka cluster is operating as intended, since — JUST LIKE CONSUMER LAG MONITORING — THE KAFKA BROKER CANNOT REPORT WHETHER OR NOT CLIENTS ARE ABLE TO USE THE CLUSTER PROPERLY.**"*
***
# 13. Monitoring Kafka (/docs/kafka/monitoring-kafka)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 13*
> **The problem this chapter solves:** *"The Apache Kafka applications have numerous measurements — **so many, in fact, that it can easily become confusing as to what is important to watch and what can be set aside.** ... They provide a detailed view into every operation in the broker, **but they can also make you the bane of whoever is responsible for managing your monitoring system.**"*
**Note:** this edition **reverses** the previous edition's central advice. The old guidance ("alert on under-replicated partitions") is explicitly retracted in favor of **SLO-based alerting**. See §3.2.
***
# 13.5 JVM and OS monitoring (/docs/kafka/monitoring-kafka/jvm-os-monitoring)
#### 5.1 Garbage collection [#51-garbage-collection]
**For Oracle Java 1.8 with G1 (Table 13-18):**
**Two attributes per metric:**
**`LastGcInfo`** — a composite of five fields:
> \*"The important value to look at is **`duration`**, as this tells you **how long, in milliseconds, the LAST GC cycle took.** The other values (`GcThreadCount`, `id`, `startTime`, `endTime`) are informational and not very useful.
>
> ⚠️ *"**you will NOT be able to see the timing of EVERY GC cycle using this attribute, as YOUNG GC cycles in particular can happen FREQUENTLY.**"*
*(Why GC matters so much here: Ch. 6 §1 — a long GC pause makes a broker's ZooKeeper ephemeral node vanish, which is indistinguishable from the broker being dead.)*
#### 5.2 Java OS monitoring — two attributes worth having [#52-java-os-monitoring--two-attributes-worth-having]
Via `java.lang:type=OperatingSystem`:
> *"this information is **LIMITED and does not represent everything you need to know.** The two attributes that are of use, **which are DIFFICULT TO COLLECT IN THE OS**, are:*
#### 5.3 OS monitoring [#53-os-monitoring]
**Five areas:** *"CPU usage, memory usage, disk usage, disk I/O, and network usage."*
**CPU:** *"you will want to look at the **system load average** at the very least."* Plus the percentage breakdown:
**`wa` and `st` are the two to watch on Kafka:** `wa` means disk is the bottleneck; `st` means your cloud neighbor is stealing your CPU.
> ### 💡 WHAT IS SYSTEM LOAD? — most people get this wrong [#-what-is-system-load--most-people-get-this-wrong]
>
> \*"While many know that system load is a measure of CPU usage, **MOST PEOPLE MISUNDERSTAND HOW IT IS MEASURED. The load average is A COUNT OF THE NUMBER OF PROCESSES THAT ARE RUNNABLE AND ARE WAITING FOR A PROCESSOR TO EXECUTE ON. LINUX ALSO INCLUDES THREADS THAT ARE IN AN UNINTERRUPTABLE SLEEP STATE, SUCH AS WAITING FOR THE DISK.**
>
> ... In a **single CPU** system, a value of **1** would mean the system is **100% loaded.** This means that on a **multiple CPU** system, **THE LOAD AVERAGE NUMBER THAT INDICATES 100% IS EQUAL TO THE NUMBER OF CPUs. For example, if there are 24 processors, 100% would be a load average of 24.**"\*
**Memory:** *"**LESS IMPORTANT to track for the broker itself**, as Kafka will normally be run with a **relatively SMALL JVM heap size.** It will use a small amount of memory outside of the heap **for compression functions**, but **most of the system memory will be left to be used for CACHE.** All the same, you should keep track of memory utilization **to make sure OTHER APPLICATIONS DO NOT INFRINGE on the broker.** You will also want to make sure that **SWAP MEMORY IS NOT BEING USED.**"*
**Disk — the most important:**
> *"**Disk is BY FAR the most important subsystem when it comes to Kafka.** All messages are persisted to disk, so **the performance of Kafka depends HEAVILY on the performance of the disks.**"*
**Network:**
> *"**Keep in mind that EVERY BIT INBOUND to the Kafka broker WILL BE A NUMBER OF BITS OUTBOUND EQUAL TO THE REPLICATION FACTOR of the topics, NOT INCLUDING CONSUMERS.** Depending on the number of consumers, **outbound network traffic could easily be AN ORDER OF MAGNITUDE LARGER than inbound. KEEP THIS IN MIND WHEN SETTING THRESHOLDS FOR ALERTS.**"*
***
# 13.3 Kafka broker metrics (/docs/kafka/monitoring-kafka/kafka-broker-metrics)
> ### ⚠️ WHO WATCHES THE WATCHERS? [#️-who-watches-the-watchers]
>
> \*"Many organizations use Kafka for **collecting application metrics, system metrics, and logs** for consumption by a central monitoring system. **This is an excellent way to decouple the applications from the monitoring system, BUT IT PRESENTS A SPECIFIC CONCERN FOR KAFKA ITSELF. If you use this same system for monitoring Kafka itself, IT IS VERY LIKELY THAT YOU WILL NEVER KNOW WHEN KAFKA IS BROKEN BECAUSE THE DATA FLOW FOR YOUR MONITORING SYSTEM WILL BE BROKEN AS WELL.**
>
> Two fixes:
>
> * *"use a **SEPARATE monitoring system for Kafka** that does not have a dependency on Kafka"*
> * *"if you have multiple datacenters, **make sure that the metrics for the Kafka cluster in datacenter A are produced to datacenter B, and VICE VERSA**"*
>
> **"However you decide to handle it, MAKE SURE THAT THE MONITORING AND ALERTING FOR KAFKA DOES NOT DEPEND ON KAFKA WORKING."**
#### 3.1 The three categories of cluster problem [#31-the-three-categories-of-cluster-problem]
> ### 💡 PREFERRED REPLICA ELECTIONS — always do this first [#-preferred-replica-elections--always-do-this-first]
>
> *"**THE FIRST STEP before trying to diagnose a problem further is to ensure that you have RUN A PREFERRED REPLICA ELECTION recently.** Kafka brokers **do not automatically take partition leadership back** (unless auto leader rebalance is enabled) after they have released leadership... **THE PREFERRED REPLICA ELECTION IS SAFE AND EASY TO RUN, SO IT'S A GOOD IDEA TO DO THAT FIRST AND SEE IF THE PROBLEM GOES AWAY.**"*
#### 3.2 ⚠️ Under-replicated partitions — and the retracted advice [#32-️-under-replicated-partitions--and-the-retracted-advice]
> *"gives a count of the number of partitions **for which the broker is the LEADER replica, where the FOLLOWER replicas are NOT CAUGHT UP.**"*
> ### ⚠️ THE URP ALERTING TRAP — the book retracting its own advice [#️-the-urp-alerting-trap--the-book-retracting-its-own-advice]
>
> \*"**In the PREVIOUS EDITION of this book, as well as in many conference talks, the authors have spoken at length about the fact that the URP metric SHOULD BE YOUR PRIMARY ALERTING METRIC** because of how many problems it describes.
>
> **THIS APPROACH HAS A SIGNIFICANT NUMBER OF PROBLEMS**, not the least of which is that **the URP metric CAN FREQUENTLY BE NONZERO FOR BENIGN REASONS.** This means that **you will receive FALSE ALERTS, WHICH LEAD TO THE ALERT BEING IGNORED. It also requires A SIGNIFICANT AMOUNT OF KNOWLEDGE to be able to understand what the metric is telling you.**
>
> **FOR THIS REASON, WE NO LONGER RECOMMEND THE USE OF URP FOR ALERTING. Instead, you should depend on SLO-BASED ALERTING to detect unknown problems.**"\*
**URP remains excellent for *diagnosis* — the interpretation guide:**
**The common-broker technique, worked:**
```bash
kafka-topics.sh --bootstrap-server kafka1.example.com:9092/kafka-cluster \
--describe --under-replicated
Topic: topicOne Partition: 5 Leader: 1 Replicas: 1,2 Isr: 1
Topic: topicOne Partition: 6 Leader: 3 Replicas: 2,3 Isr: 3
Topic: topicTwo Partition: 3 Leader: 4 Replicas: 2,4 Isr: 4
Topic: topicTwo Partition: 7 Leader: 5 Replicas: 5,2 Isr: 5
Topic: topicSix Partition: 1 Leader: 3 Replicas: 2,3 Isr: 3
... ▲
BROKER 2 APPEARS IN EVERY ROW
but is never in the Isr
```
> *"In this example, **the common broker is number 2.** This indicates that **this broker is having a problem with message replication** and will lead us to focus our investigation on that one broker. **If there is NO common broker, there is likely a CLUSTER-WIDE problem.**"*
#### 3.3 Cluster-level problems [#33-cluster-level-problems]
##### Problem A: unbalanced load — *"the easiest to FIND even though FIXING it can be an involved process"* [#problem-a-unbalanced-load--the-easiest-to-find-even-though-fixing-it-can-be-an-involved-process]
**Five metrics to compare across brokers:**
**What "balanced" looks like (Table 13-4):**
> *"**Assuming you have ALREADY RUN A PREFERRED REPLICA ELECTION**, a large deviation indicates that the traffic is not balanced. To resolve this, you will need to **move partitions from the heavily loaded brokers to the less heavily loaded brokers** using `kafka-reassign-partitions.sh` (Ch. 12)."*
> ### HELPERS FOR BALANCING CLUSTERS [#helpers-for-balancing-clusters]
>
> *"**The Kafka broker itself does NOT provide for automatic reassignment of partitions.** This means that balancing traffic can be **a MIND-NUMBING PROCESS of manually reviewing long lists of metrics and trying to come up with a replica assignment that works.**"* Tools: **`kafka-assigner`** in LinkedIn's open source **kafka-tools** repo; *"Some enterprise offerings for Kafka support also provide this feature."*
##### Problem B: resource exhaustion [#problem-b-resource-exhaustion]
**The bottlenecks:** *"**CPU, disk IO, and network throughput** are a few of the most common."*
> ⚠️ **The disk-utilization exception:** *"**DISK UTILIZATION is NOT one of them, as the brokers will operate properly RIGHT UP UNTIL THE DISK IS FILLED, AND THEN THIS DISK WILL FAIL ABRUPTLY.**"*
That's an important nuance: disk *space* isn't a gradual-degradation signal — it's a cliff. Alert on free space as a **capacity** metric, not a performance one.
**OS-level metrics to track:**
> 💡 **The key insight:** *"**Exhausting ANY of these resources will typically show up as THE SAME PROBLEM: under-replicated partitions.** It's critical to remember that **THE BROKER REPLICATION PROCESS OPERATES IN EXACTLY THE SAME WAY THAT OTHER KAFKA CLIENTS DO. IF YOUR CLUSTER IS HAVING PROBLEMS WITH REPLICATION, THEN YOUR CUSTOMERS ARE HAVING PROBLEMS WITH PRODUCING AND CONSUMING MESSAGES AS WELL.**"*
*(Ch. 6 §4.3: followers use the same Fetch requests consumers use. Replication health **is** a proxy for client health.)*
> *"It makes sense to **develop a BASELINE for these metrics when your cluster is operating correctly and then SET THRESHOLDS THAT INDICATE A DEVELOPING PROBLEM LONG BEFORE YOU RUN OUT OF CAPACITY.** ... **All Topics Bytes In Rate is a good guideline to show cluster usage.**"*
#### 3.4 Host-level problems [#34-host-level-problems]
**Four categories:** *"Hardware failures · Networking · Conflicts with another process · Local configuration differences"*
##### Hardware — the soft failures are the dangerous ones [#hardware--the-soft-failures-are-the-dangerous-ones]
> *"Hardware failures are **sometimes obvious**, like when the server just stops working, **but it's the LESS OBVIOUS problems that cause performance issues. These are usually SOFT FAILURES that ALLOW THE SYSTEM TO KEEP RUNNING BUT DEGRADE OPERATION.** This could be **a bad bit of memory, where the system has detected the problem and BYPASSED THAT SEGMENT (reducing the overall available memory). The same can happen with a CPU failure.**"*
**Tools:** *"the facilities that your hardware provides, such as an **intelligent platform management interface (IPMI)**"*; *"looking at the **kernel ring buffer using `dmesg`** will help you to see log messages that are getting thrown to the system console."*
##### 💡 Disk failure — and why one bad disk ruins everything [#-disk-failure--and-why-one-bad-disk-ruins-everything]
> *"**The MORE COMMON type of hardware failure that leads to a performance degradation in Kafka is A DISK FAILURE.** Kafka is dependent on the disk for persistence, and **PRODUCER PERFORMANCE IS DIRECTLY TIED TO HOW FAST YOUR DISKS COMMIT THOSE WRITES.** Any deviation will show up as **problems with the performance of the producers AND the replica fetchers. The latter is what leads to under-replicated partitions.**"*
> ### ⚠️ ONE BAD EGG [#️-one-bad-egg]
>
> *"**A SINGLE DISK FAILURE ON A SINGLE BROKER CAN DESTROY THE PERFORMANCE OF AN ENTIRE CLUSTER.** This is because **producer clients will connect to ALL brokers that lead partitions for a topic**, and if you have followed best practices, **those partitions will be EVENLY SPREAD over the entire cluster. If ONE broker starts performing poorly and slowing down produce requests, THIS WILL CAUSE BACK PRESSURE IN THE PRODUCERS, SLOWING DOWN REQUESTS TO ALL BROKERS.**"*
**Disk monitoring checklist:**
**The BBU failure mode is a classic:** nothing is "broken," no alert fires, and write latency quietly triples because the controller silently switched from write-back to write-through.
##### Networking [#networking]
##### Process conflicts and config drift [#process-conflicts-and-config-drift]
> *"another common problem to look for is **another application running on the system that is consuming resources.** This could be **something that was installed in error**, or **a process that is SUPPOSED to be running, such as A MONITORING AGENT, but is having problems.** Use tools such as **`top`.**"*
> *"If the other options have been exhausted... **a CONFIGURATION DIFFERENCE has likely crept in.** ... **This is why it is CRUCIAL that you utilize a CONFIGURATION MANAGEMENT SYSTEM, such as Chef or Puppet, in order to maintain consistent configurations across your OSes and applications (including Kafka).**"*
***
# 13.8 Lag monitoring (/docs/kafka/monitoring-kafka/lag-monitoring)
> *"For Kafka consumers, **the most important thing to monitor is the consumer lag.** ... **this is ONE OF THE CASES WHERE EXTERNAL MONITORING FAR SURPASSES WHAT IS AVAILABLE FROM THE CLIENT ITSELF.**"*
**Why the client metric is inadequate (restated):**
**The correct architecture:**
> \*"an **EXTERNAL PROCESS** that can watch **BOTH the state of the partition on the broker, tracking the offset of the most recently produced message, AND the state of the consumer, tracking the last offset the consumer group has committed.** This provides **an OBJECTIVE view that can be updated REGARDLESS OF THE STATUS OF THE CONSUMER ITSELF.**
>
> **This checking must be performed for EVERY PARTITION that the consumer group consumes. For a large consumer, like MirrorMaker, this may mean TENS OF THOUSANDS OF PARTITIONS.**"\*
#### ⚠️ Why the CLI approach doesn't scale [#️-why-the-cli-approach-doesnt-scale]
> \*"Monitoring lag like this, however, **presents its own problems:**
>
> 1. *"you must understand **FOR EACH PARTITION what is A REASONABLE AMOUNT OF LAG.** A topic that receives **100 messages an hour** will need a different threshold than a topic that receives **100,000 messages per second.**"*
> 2. *"you must be able to **consume ALL of the lag metrics into a monitoring system and set alerts on them.** If you have a consumer group that consumes **100,000 partitions over 1,500 topics, YOU MAY FIND THIS TO BE A DAUNTING TASK.**"*
#### 💡 Burrow — the recommended answer [#-burrow--the-recommended-answer]
> \*"an **open source application, originally developed by LinkedIn**, that provides **consumer STATUS monitoring** by gathering lag information for **ALL consumer groups in a cluster** and **calculating A SINGLE STATUS for each group saying whether the consumer group is WORKING PROPERLY, FALLING BEHIND, or is STALLED OR STOPPED ENTIRELY.**
>
> **IT DOES THIS WITHOUT REQUIRING THRESHOLDS by MONITORING THE PROGRESS that the consumer group is making on processing messages**, though you can also get the message lag as an absolute number."\*
> *"Deploying Burrow can be **an easy way to provide monitoring for ALL consumers in a cluster, as well as in MULTIPLE clusters**, and it can be easily integrated with your existing monitoring and alerting system."*
> *"**If there is NO other option**, the `records-lag-max` metric from the consumer client will provide **at least a PARTIAL view.** It is **strongly suggested**, however, that you utilize an external monitoring system like Burrow."*
***
# 13.6 Logging (/docs/kafka/monitoring-kafka/logging)
> *"Like many applications, **the Kafka broker will FILL DISKS with log messages IN MINUTES if you let it.** In order to get useful information from logging, it is important to **enable the right loggers at the right levels.**"*
**Baseline:** *"By simply logging all messages at the **INFO** level, you will capture a significant amount of important information."*
**Two loggers to separate into their own files, both at INFO:**
#### 💡 The log-compaction loggers — covering a genuine monitoring blind spot [#-the-log-compaction-loggers--covering-a-genuine-monitoring-blind-spot]
> \*"It is also helpful to log information regarding the status of the **log compaction threads. THERE IS NO SINGLE METRIC TO SHOW THE HEALTH OF THESE THREADS, AND IT IS POSSIBLE FOR FAILURE IN COMPACTION OF A SINGLE PARTITION TO HALT THE LOG COMPACTION THREADS ENTIRELY, AND SILENTLY.**
>
> Enabling **`kafka.log.LogCleaner`, `kafka.log.Cleaner`, and `kafka.log.LogCleanerManager`** at the **DEBUG** level will output information about the status of these threads, **including information about each partition being compacted, including the size and number of messages in each.**
>
> **Under normal operations, THIS IS NOT A LOT OF LOGGING, WHICH MEANS THAT IT CAN BE ENABLED BY DEFAULT WITHOUT OVERWHELMING YOU.**"\*
**For debugging only:**
*(Ch. 11 §7 showed what one of these lines contains — it's simultaneously an audit record and a full latency trace.)*
***
# 13.1 Metric basics (/docs/kafka/monitoring-kafka/metric-basics)
#### 1.1 Where the metrics live [#11-where-the-metrics-live]
**All Kafka metrics are exposed via JMX.** Three collection approaches:
> ### ⚠️ FINDING THE JMX PORT — and why you probably shouldn't open it [#️-finding-the-jmx-port--and-why-you-probably-shouldnt-open-it]
>
> \*"the broker sets the configured JMX port in the broker information stored in ZooKeeper. The **`/brokers/ids/`** znode contains JSON-formatted data including **`hostname`** and **`jmx_port`** keys.
>
> **However, REMOTE JMX IS DISABLED BY DEFAULT in Kafka FOR SECURITY REASONS. If you are going to enable it, you must properly configure security for the port. THIS IS BECAUSE JMX NOT ONLY ALLOWS A VIEW INTO THE STATE OF THE APPLICATION, IT ALSO ALLOWS CODE EXECUTION.**
>
> **It is HIGHLY RECOMMENDED that you use a JMX metrics agent that is LOADED INTO THE APPLICATION.**"\*
**JMX is a remote-code-execution surface.** That single sentence should settle the architecture decision: use an in-process agent (option ②), not an open remote JMX port.
#### 1.2 💡 The five metric sources — ordered by objectivity [#12--the-five-metric-sources--ordered-by-objectivity]
> *"the **LOWER in the list, the more OBJECTIVE a view of Kafka they provide.**"*
| Category | Description |
| -------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Application metrics** | *"from Kafka itself, from the JMX interface"* |
| **Logs** | *"Also from Kafka itself. Because it is some form of **text or structured data, and not just a number**, it requires a little more processing"* |
| **Infrastructure metrics** | *"from systems that you have **in front of** Kafka but are still **within the request path and under your control.** An example is a **load balancer**"* |
| **Synthetic clients** | *"**external to your Kafka deployment, just like a client**, but under your direct control and **typically not performing the same work as your clients.** An external monitor like **Kafka Monitor**"* |
| **Client metrics** | *"exposed by the Kafka clients that connect to your cluster"* |
**The website analogy that makes the point:**
> *"The web server is running properly, and **all of the metrics IT is reporting say that it is working. HOWEVER, there is a problem with the NETWORK between your web server and your external users, which means that NONE OF YOUR USERS CAN REACH THE WEB SERVER.** A synthetic client running **outside your network** would detect this."*
> *"relying on metrics from your brokers **will SUFFICE AT THE START, but later on you will want a more objective view.**"*
#### 1.3 Alerting vs debugging vs historical — three different data lifecycles [#13-alerting-vs-debugging-vs-historical--three-different-data-lifecycles]
| Purpose | Retention | Character | Consumer |
| -------------- | --------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------ |
| **Alerting** | *"a very short period... **hours, or maybe days**"* | *"important for these metrics to be **more OBJECTIVE**, as a problem that does not impact clients is far less critical than one that does"* | *"automation that responds to known problems, as well as the human operators"* |
| **Debugging** | *"**days or weeks** past when it is collected"* | *"more **SUBJECTIVE** measurements, or data from the Kafka application itself"* | Humans, diagnosing |
| **Historical** | *"**measured in YEARS**"* | *"resources used, including compute, storage, and network"* + *"**additional METADATA to put the metrics into context, such as when brokers were added to or removed from the cluster**"* | Capacity management |
> 💡 **You don't have to collect debugging data continuously:** *"**it is NOT ALWAYS NECESSARY to collect this data into a monitoring system.** If the metrics are used for debugging problems **in place**, it is sufficient that the metrics are **available when needed. YOU DO NOT NEED TO OVERWHELM THE MONITORING SYSTEM BY COLLECTING TENS OF THOUSANDS OF VALUES ON AN ONGOING BASIS.**"*
#### 1.4 💡 Automation vs humans — the "Check Engine light" principle [#14--automation-vs-humans--the-check-engine-light-principle]
**The car analogy:**
> *"To properly adjust the ratio of air to fuel, the computer needs a number of measurements of air density, fuel, exhaust... **These measurements would be OVERWHELMING to the human operator.** Instead, we have a **'CHECK ENGINE' LIGHT. A SINGLE INDICATOR tells you that there is a problem, and there is a way to find out more detailed information to tell you exactly what the problem is.**"*
#### 1.5 Application health checks [#15-application-health-checks]
***
# 13.11 The recommended monitoring stack (/docs/kafka/monitoring-kafka/recommended-monitoring-stack)
#### The diagnostic decision tree [#the-diagnostic-decision-tree]
***
# 13.12 Self-test (/docs/kafka/monitoring-kafka/self-test)
Why is an in-process JMX agent recommended over an open remote JMX port?
List the five metric sources in order of objectivity. Which are best for SLIs, and which for debugging?
Give the website analogy for why broker metrics alone are insufficient.
Contrast the retention and character of alerting, debugging, and historical metrics. What extra metadata does the last one need?
Why is "collect everything continuously" wrong for debugging metrics?
Explain the "Check Engine light" principle and the specific failure mode it prevents.
Give both application health-check methods and the drawback of the second.
Define SLI, SLO, and SLA precisely. What's an OLA?
Why does the book say most engineers should be setting SLOs, not SLAs — and why bother if you have no external customers?
What makes a good SLI? Why do quantile metrics fail as SLIs, and what's the compromise?
Name the five SLI types.
Why shouldn't you accept an SLO on data freshness?
Why can't you alert on an SLO directly? Explain the burn-rate technique with the 0.1% / 0.4% / 2% example.
"Who watches the watchers?" — state the problem and both solutions.
Name the three categories of cluster problem. Which produces "that's really weird," and which two metrics cover it?
What single action should you always take *before* diagnosing further, and why is it safe?
Why does the book **retract** its previous advice about URP alerting? What is URP still good for?
A steady URP count reported by many brokers means what? A fluctuating one?
Describe the "common broker" technique. What does *no* common broker tell you?
List the five metrics for detecting cluster imbalance, and what balanced looks like.
Why is disk *utilization* not a performance bottleneck metric for Kafka?
Why does exhausting *any* resource show up as under-replicated partitions? What does that imply about your clients?
Explain the "one bad egg" cascade in full. Why does *good* partition balance make it worse?
What is a BBU, and what's its silent failure mode?
Name three soft hardware failures that degrade rather than break.
What does `ActiveControllerCount` summing to 2 mean? To 0? What's the fix for each, and what complication should you expect?
What are the two thresholds for request handler idle ratio? Why is this pool more critical than the network threads?
How many request handler threads should you configure, and why is that enough given how many requests exist?
Describe the pre-0.10 recompression tax and what the v2 format changed. Why is fixing it "one of the single largest performance improvements"?
Give the seven attributes of a rate metric. Which is best for spikes? Which should you not alert on?
Why can bytes-out equal bytes-in with zero consumers? Give the formula.
Why is there no "messages out" metric?
Why is leader count more alert-worthy than partition count? What derived percentage should you compute, and what's the expected value?
Why does `OfflinePartitionsCount` read 0 on most brokers, and how must you aggregate it?
Name all seven request timing phases and what a spike in each one points to.
Which request type's p99.9 total time makes the best broker-side latency alert, and why not Fetch?
What does a size discrepancy between partitions of the same topic indicate?
Why isn't `LogEndOffset − LogStartOffset` the message count?
Which two JVM `OperatingSystem` attributes matter, and what consumes file descriptors?
What exactly does Linux system load average count? What value means 100% on a 24-CPU box?
Why is heap memory relatively unimportant for a broker, and what should you still watch?
Which three loggers should you enable at DEBUG by default, and what invisible failure do they cover?
Which two loggers should get their own files, and why?
Name the two producer metrics that deserve alerts and what each one means.
What are the three views of producer traffic volume, and how do requests, batches, messages, and bytes nest?
What does `record-queue-time-avg` measure, and which two configs does it help you tune?
Why is per-broker `request-latency-avg` the most useful per-broker producer metric?
Why are consumer metrics "useful to look at but not useful for alerting"?
Give both reasons `records-lag-max` is a poor lag metric.
Why is alerting on a *minimum* `records-consumed-rate` risky?
Name the four coordinator metrics and what each detects. What should `sync-rate` normally be?
Why is quota throttling invisible without specific metrics, and which two are they?
Describe the correct architecture for lag monitoring, and the two reasons the CLI approach doesn't scale.
How does Burrow avoid needing thresholds?
What two questions does end-to-end monitoring answer, and why can't the broker answer them itself?
State the recurring theme: which three critical facts about Kafka cannot come from broker metrics?
**Previous:** [Chapter 12 — Administering Kafka](12-administering-kafka.md)
**Next:** [Chapter 14 — Stream Processing](14-stream-processing.md)
# 13.2 Service-level objectives (/docs/kafka/monitoring-kafka/service-level-objectives)
> *"One area of monitoring that is **especially critical for INFRASTRUCTURE SERVICES, such as Kafka**... This is how we **communicate to our clients what level of service they can expect.** The clients **want to be able to treat services like Kafka as an OPAQUE SYSTEM: they do not want or need to understand the internals** — only the interface they are using and knowing it will do what they need it to do."*
#### 2.1 The definitions — used incorrectly constantly [#21-the-definitions--used-incorrectly-constantly]
> *"Frequently, you will hear engineers, managers, executives, and everyone else **use terms in the 'service-level' space INCORRECTLY, which leads to confusion about what is actually being talked about.**"*
> ### OPERATIONAL-LEVEL AGREEMENT (OLA) [#operational-level-agreement-ola]
>
> *"**Less frequently used.** It describes **agreements between MULTIPLE INTERNAL SERVICES or support providers in the overall delivery of an SLA.** The goal is to assure that the multiple activities necessary to fulfill the SLA are properly described and accounted for in day-to-day operations."*
> 💡 **The practical correction:** *"It is very common to hear people talk about SLAs when **they really mean SLOs.** ... **it is RARE that the engineers running the applications are responsible for anything more than the performance of that service WITHIN THE SLOs.** ... those who only have internal clients generally **do not have SLAs** with those internal customers. **THIS SHOULD NOT PREVENT YOU FROM SETTING AND COMMUNICATING SLOs, HOWEVER, as doing that will lead to FEWER ASSUMPTIONS BY CUSTOMERS as to how they think Kafka should be performing.**"*
#### 2.2 What makes a good SLI [#22-what-makes-a-good-sli]
> \*"In general, the metrics for your SLIs **should be gathered using something EXTERNAL to the Kafka brokers.** ... **YOUR CLIENTS DO NOT CARE IF YOU THINK YOUR SERVICE IS RUNNING CORRECTLY; IT IS THEIR EXPERIENCE (IN AGGREGATE) THAT MATTERS.**
>
> This means: **infrastructure metrics are OK, synthetic clients are GOOD, and CLIENT-SIDE METRICS ARE PROBABLY THE BEST for most of your SLIs.**"\*
**The five common SLI types (Table 13-2):**
| Type | Question |
| ---------------- | ------------------------------------------------------------------------------------------------------ |
| **Availability** | *"Is the client able to make a request and get a response?"* |
| **Latency** | *"How quickly is the response returned?"* |
| **Quality** | *"Does the response include a proper response?"* |
| **Security** | *"Are the request and response appropriately protected, whether that is authorization or encryption?"* |
| **Throughput** | *"Can the client get enough data, fast enough?"* |
#### ⚠️ 2.3 Why quantiles make bad SLIs [#️-23-why-quantiles-make-bad-slis]
> \*"it is usually **better for your SLIs to be based on A COUNTER OF EVENTS THAT FALL INSIDE THE THRESHOLDS of the SLO.** This means that ideally, **each event would be individually checked** to see if it meets the threshold.
>
> **THIS RULES OUT QUANTILE METRICS AS GOOD SLIs, as those will only tell you that 90% of your events were below a given value WITHOUT ALLOWING YOU TO CONTROL WHAT THAT VALUE IS.**"\*
> 💡 **The bucket compromise:** *"aggregating values into **BUCKETS** (e.g., 'less than 10 ms,' '10–50 ms,' '50–100 ms') can be useful, **ESPECIALLY WHEN YOU ARE NOT YET SURE WHAT A GOOD THRESHOLD IS.** This will give you **a view into the DISTRIBUTION of events within the range of the SLO**, and you can configure the buckets so that **the boundaries are reasonable values for the SLO threshold.**"*
> ### CUSTOMERS ALWAYS WANT MORE [#customers-always-want-more]
>
> \*"There are some SLOs that your customers may be interested in that are important to them **but NOT WITHIN YOUR CONTROL.** For example, they may be concerned about **the CORRECTNESS or FRESHNESS of the data produced to Kafka.**
>
> **DO NOT AGREE TO SUPPORT SLOs THAT YOU ARE NOT RESPONSIBLE FOR, as that will only lead to taking on work that DILUTES THE CORE JOB of keeping Kafka running properly.** Make sure to **connect them with the proper group.**"\*
#### 💡 2.4 Using SLOs for alerting — the burn-rate technique [#-24-using-slos-for-alerting--the-burn-rate-technique]
> \*"**SLOs should inform your PRIMARY ALERTS.** ... **Generally speaking, IF A PROBLEM DOES NOT IMPACT YOUR CLIENTS, IT DOES NOT NEED TO WAKE YOU UP AT NIGHT.**
>
> **SLOs will also tell you about the problems that YOU DON'T KNOW HOW TO DETECT because you've never seen them before. THEY WON'T TELL YOU WHAT THOSE PROBLEMS ARE, BUT THEY WILL TELL YOU THAT THEY EXIST.**"\*
**The problem with alerting on the SLO directly:**
> \*"**SLOs are best for LONG TIMESCALES**, such as a week, as we want to report them to management and customers... In addition, **BY THE TIME THE SLO ALERT FIRES, IT'S TOO LATE — YOU'RE ALREADY OPERATING OUTSIDE OF THE SLO.**
>
> **The BEST way to approach using SLOs for alerting is to OBSERVE THE RATE AT WHICH YOU ARE BURNING THROUGH YOUR SLO over its timeframe.**"\*
**The worked example — follow the arithmetic:**
**Further reading the book recommends:** *Site Reliability Engineering* and *The Site Reliability Workbook*, both ed. Betsy Beyer et al. (O'Reilly).
***
# 7.7 What actually breaks in production — Ch. 7 consolidated (/docs/kafka/reliable-data-delivery/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ---------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------- |
| 1 | **Producer got success; message gone. Brokers configured "perfectly."** | `acks=1` + leader crashed before replicating; followers were **still nominally in-sync** because `replica.lag.time.max.ms` hadn't elapsed | `acks=all` **and** `min.insync.replicas=2` |
| 2 | **`acks=all` set, message still lost** | `min.insync.replicas=1` and both followers were down → "all in-sync replicas" meant **one** | `min.insync.replicas=2` — this is the *only* thing that makes `acks=all` mean what you think |
| 3 | **`acks=all` set, still lost the message on "Leader not Available"** | Producer **didn't retry**; the broker never received it | Default (infinite) retries + `delivery.timeout.ms` > measured recovery time |
| 4 | **Total silent loss during a cluster outage** | `acks=0` — *"we won't get any error if the partition is offline, a leader election is in progress, or even if the entire Kafka cluster is unavailable"* | Never `acks=0` for data you care about; don't trust `acks=0` benchmarks |
| 5 | **Under-replicated partitions climbing while producers are happy** | `acks=1` lets you *"write to the leader faster than it can replicate"* | `acks=all`; fix replication throughput (NIC — Ch. 2) |
| 6 | **Latency mysteriously improved mid-incident** | A limping replica **fell out of the ISR** — the latency drag vanished, and so did a third of your durability | Alert on **UnderReplicatedPartitions**, not just latency. Latency recovery ≠ resolution |
| 7 | **A replica flaps in and out of the ISR** | Long **GC pauses** disconnecting the broker from ZooKeeper; historically worsened by large max request size + large heap | Kafka 2.5.0+ defaults; JVM 8+ with **G1**; tune for large messages |
| 8 | **A follower fetches continuously yet is declared out of sync** | Condition ③: it must have had **zero lag at least once** in the window — steady lag is still out-of-sync | Fix throughput; understand the definition before "fixing" the config |
| 9 | **RF=3 and a partition still went offline when one switch failed** | All three replicas in **one rack** | `broker.rack` everywhere; AZs as racks in cloud; audit after reassignments (Ch. 2) |
| 10 | **Partition offline for hours after multiple broker failures** | No in-sync replica; `unclean.leader.election.enable=false` (correct default) | Accept the outage, or **deliberately** enable unclean election — **and turn it back off after recovery** |
| 11 | **Same offset returns different messages to different consumers; downstream reports disagree** | **Unclean leader election** — offsets 100–200 were rewritten with new data; the old leader later **deleted** its divergent messages | This is the *known cost* of unclean election. Prefer `false` + `min.insync.replicas=2` so it can't be needed |
| 12 | **Producers suddenly get `NotEnoughReplicasException`** | Working as designed: in-sync count fell below `min.insync.replicas` → the partition went **read-only** | Restore a replica and let it catch up. This is durability protecting you |
| 13 | **Acknowledged data lost in a correlated power event** | No fsync by default — Kafka bets on **independent failure domains** | Replicas in separate racks/AZs; only consider `flush.messages`/`flush.ms` after reading the throughput implications |
| 14 | **Cluster destabilizes in a cloud environment with variable latency** | `zookeeper.session.timeout.ms` too low for the environment | 2.5.0+ default of 18 s; tune high enough to avoid flapping, low enough to catch frozen brokers |
| 15 | **Consume latency worse after upgrading to 2.5.0** | `replica.lag.time.max.ms` 10 s → 30 s raises the ceiling on *"until a message arrives to all replicas and consumers are allowed to consume it"* | Understand the trade; lower it only if you accept more ISR flapping |
| 16 | **A single consumer only sees a fraction of messages** | It shares `group.id` with another application | *"it will need a unique `group.id`"* |
| 17 | **Consumer restarts and silently skips everything that arrived while it was down** | `auto.offset.reset=latest` with an invalid/expired committed offset | `earliest` (accept duplicates) or `none` (fail loudly) |
| 18 | **Records read but never processed after handing them to a thread pool** | **Autocommit** committed offsets for records that had only been *read* | *"there is no choice but to use manual offset commit"* once processing leaves the poll loop |
| 19 | **Offsets committed for the batch even though processing threw mid-batch** | Autocommit doesn't know what you processed | Manual commits after processing (Ch. 4) |
| 20 | **One broker is overloaded purely by offset commits** | *"all offset commits of a single consumer group are produced to the same broker"*, and each commit ≈ produce with `acks=all` | Commit less often; *"committing after every message should only ever be done on very low-throughput topics"* |
| 21 | **A failed record is silently skipped** | Committed offset 31 after record 30 failed — the watermark marks 30 processed too | Pattern A (`pause()` + buffer + retry) or Pattern B (**retry topic / DLQ**) |
| 22 | **Retry topic reorders events** | Pattern B trades ordering for liveness | Choose deliberately; Pattern A preserves order but head-of-line blocks |
| 23 | **Stateful consumer resumes at the right offset with the wrong aggregate** | Offsets and application state were committed **separately** | Write results + offsets atomically (Ch. 8 transactions), or use **Kafka Streams / Flink** |
| 24 | **Application "reliable" in test, loses data in production** | Never tested under **hanging disk / high latency / disk full / rolling restarts** | Trogdor or equivalent; write the expected behavior *first*, then measure |
| 25 | **`delivery.timeout.ms` too short; producer gives up during failover** | Never measured how long leader election actually takes in *your* cluster | Run the leader-election test with VerifiableProducer/Consumer and use the measured number |
| 26 | **Lag alerts flap constantly; team ignores them** | Lag oscillates by construction — static thresholds don't work | Trend/status-based checking (**Burrow**) |
| 27 | **Messages "lost somewhere" and nobody can prove where** | No end-to-end reconciliation of produced vs consumed counts | Build produce/consume counters + timestamp-based latency; note there's **no open-source implementation** |
| 28 | **Produce-to-consume latency looks impossibly low** | Broker configured for **append-time** timestamps, overriding create-time | Know which timestamp type your topics use |
| 29 | **Failed request metrics rising** | Could be benign (`NOT_LEADER_FOR_PARTITION` during maintenance) or serious | *"Unexplained increases should always be investigated"* — the metrics are **tagged with the specific error** |
| 30 | **Hand-rolled retry loop caused duplicates and reordering** | Application-level retries layered on producer retries | *"if all the error handler is doing is retrying, we'll be better off relying on the producer's retry functionality"* |
***
# 7.3 Broker configuration — three knobs (/docs/kafka/reliable-data-delivery/broker-configuration-three-knobs)
> *"these can apply at the **broker level**, controlling configuration for all topics, **and at the topic level**, controlling behavior for a specific topic."*
**Why per-topic control matters — the bank example:**
> *"at a bank, the administrator will probably want to set **very reliable defaults for the entire cluster** but **make an exception to the topic that stores customer complaints where some data loss is acceptable.**"*
#### 3.1 Replication factor [#31-replication-factor]
`replication.factor` (topic) / `default.replication.factor` (broker, for auto-created topics).
> *"Even after a topic exists, we can choose to add or remove replicas and thereby modify the replication factor using Kafka's replica assignment tool."*
##### The five considerations for choosing N [#the-five-considerations-for-choosing-n]
| Consideration | The argument |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **Availability** | *"A partition with just **one replica will become unavailable even during a ROUTINE RESTART of a single broker.**"* |
| **Durability** | *"If a partition has a single replica and the disk becomes unusable... we've lost all the data. With more copies, **especially on different storage devices**, the probability of losing all of them is reduced."* |
| **Throughput** | Replication traffic multiplies — see the table below |
| **End-to-end latency** | *"Each produced record has to be replicated to all in-sync replicas before it is available for consumers. **In theory**, with more replicas, there is higher probability that one of these replicas is a bit slow\... **In practice, if one broker becomes slow for any reason, it will slow down every client that tries using it, REGARDLESS OF REPLICATION FACTOR.**"* |
| **Cost** | *"the most common reason for using a replication factor lower than 3 for noncritical data"* |
**The throughput arithmetic — worth internalizing:**
*(This is the term Ch. 2 said people forget when sizing NICs: replication is an additional consumer of every byte.)*
**The RF=2 cost argument, and its honest caveat:**
> *"Since many storage systems already replicate each block 3 times, it sometimes makes sense to reduce costs by configuring Kafka with a replication factor of 2. **Note that this will still REDUCE AVAILABILITY compared to a replication factor of 3, but DURABILITY will be guaranteed by the storage device.**"*
That's a precise distinction: underlying storage replication protects your **bytes**; it does nothing for **partition availability** during a broker restart or failure.
##### Replica placement [#replica-placement]
> *"Kafka will always make sure each replica for a partition is on a **separate broker**. **In some cases, this is not safe enough.** If all replicas for a partition are placed on brokers that are on the same rack, **and the top-of-rack switch misbehaves, we will lose availability of the partition REGARDLESS OF THE REPLICATION FACTOR.**"*
*(Mechanism: Ch. 6 §6.3's rack-alternating broker list. Caveat: Ch. 2 — rack awareness applies to **newly created** partitions only, and nothing monitors it after a reassignment.)*
#### 3.2 Unclean leader election [#32-unclean-leader-election]
`unclean.leader.election.enable` — **broker-level only (and in practice cluster-wide). Default: `false`.**
> *"leader election is 'clean' in the sense that it **guarantees no loss of committed data — by definition, committed data exists on all in-sync replicas.** But **what do we do when no in-sync replica exists except for the leader that just became unavailable?**"*
##### The two scenarios that produce this state [#the-two-scenarios-that-produce-this-state]
**Scenario A is the one that should change how you configure Kafka.** With `min.insync.replicas=1` (the default), `acks=all` degenerates into `acks=1` the moment your followers die — and it does so *silently*. This is precisely the hole `min.insync.replicas=2` plugs (§3.3).
##### The choice, and the consistency damage spelled out [#the-choice-and-the-consistency-damage-spelled-out]
**The book's worked example of the inconsistency — this is the part people underestimate:**
**The operational recipe:**
> *"By default it is set to `false`, which will not allow out-of-sync replicas to become leaders. This is the safest option... It is always possible for an administrator to look at the situation, **decide to accept the data loss** in order to make the partitions available, and **switch this configuration to `true` before starting the cluster. JUST DON'T FORGET TO TURN IT BACK TO `false` AFTER THE CLUSTER RECOVERED.**"*
*(Ch. 5's `electLeaders(ElectionType.UNCLEAN)` is the per-partition, no-restart alternative.)*
#### 3.3 `min.insync.replicas` — the fix for the silent-degradation hole [#33-mininsyncreplicas--the-fix-for-the-silent-degradation-hole]
Available at **both** topic and broker level.
**The precise statement of the problem:**
> *"**part of the problem is that, per Kafka reliability guarantees, data is considered committed when it is written to all in-sync replicas — EVEN WHEN 'ALL' MEANS JUST ONE REPLICA and the data could be lost if that replica is unavailable.**"*
**Why this is the right behavior:**
> *"This **prevents the undesirable situation where data is produced and consumed, only to disappear when unclean election occurs.**"*
**Recovery:** *"we must make one of the two unavailable partitions available again (maybe restart the broker) and wait for it to catch up and get in sync."*
#### 3.4 Keeping replicas in sync [#34-keeping-replicas-in-sync]
Two configs, matching the two ways a replica goes out of sync:
| Config | Guards against | Default history | Tuning guidance |
| ---------------------------------- | -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **`zookeeper.session.timeout.ms`** | Losing connectivity to ZooKeeper | **6 s → 18 s in 2.5.0**, *"in order to increase the stability of Kafka clusters in cloud environments where network latencies show higher variance"* | *"high enough to avoid random flapping caused by **garbage collection or network conditions**, but still low enough to make sure brokers that are **actually frozen** will be detected in a timely manner"* |
| **`replica.lag.time.max.ms`** | Falling behind the leader | **10 s → 30 s in 2.5.0**, *"to improve resilience of the cluster and avoid unnecessary flapping"* | ⚠️ *"**this higher value also impacts MAXIMUM LATENCY FOR THE CONSUMER — with the higher value it can take up to 30 seconds until a message arrives to all replicas and the consumers are allowed to consume it.**"* |
**That last caveat is the hidden cost of the 2.5.0 default change.** Raising `replica.lag.time.max.ms` buys cluster stability and pays with a *worse worst-case consume latency* — because visibility is gated on ISR replication (Ch. 6 §5.6).
#### 3.5 Persisting to disk — and why Kafka mostly doesn't [#35-persisting-to-disk--and-why-kafka-mostly-doesnt]
> *"Kafka will **acknowledge messages that were not persisted to disk**, depending just on the number of replicas that received the message. Kafka will flush messages to disk **when rotating segments (by default 1 GB in size) and before restarts** but will otherwise **rely on Linux page cache to flush messages when it becomes full.**"*
**The reasoning:**
> *"**having three machines in separate racks or availability zones, each with a copy of the data, is SAFER than writing the messages to disk on the leader, because simultaneous failures on two different racks or zones are so unlikely.**"*
**If you want fsync anyway:**
> *"Before using this feature, it is worth reading **how fsync impacts Kafka's throughput and how to mitigate its drawbacks.**"*
***
# 7. Reliable Data Delivery (/docs/kafka/reliable-data-delivery)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 7*
> **The thesis, stated in the first sentence:** *"**Reliability is a property of a system — not of a single component**... the systems that integrate with Kafka are as important as Kafka itself. And because reliability is a system concern, **it cannot be the responsibility of just one person. Everyone — Kafka administrators, Linux administrators, network and storage administrators, and the application developers — must work together to build a reliable system.**"*
**And the warning that makes this chapter necessary:**
> *"Kafka was written to be configurable enough, and its client API flexible enough, to allow all kinds of reliability trade-offs. **Because of its flexibility, it is also easy to accidentally shoot ourselves in the foot when using Kafka — believing that our system is reliable when in fact it is not.**"*
***
# 7.1 Kafka's reliability guarantees — the exact contract (/docs/kafka/reliable-data-delivery/kafka-s-reliability-guarantees)
The book frames this by analogy to **ACID**: *"Those guarantees are the reason people trust relational databases with their most critical applications — **they know exactly what the system promises and how it will behave in different conditions.** They understand the guarantees and can write safe applications by relying on them."*
#### The four guarantees, verbatim [#the-four-guarantees-verbatim]
**Read the fine print in each:**
* **①** requires *"the same producer in the same partition."* Two producers writing to one partition have no relative ordering guarantee. Neither do two partitions.
* **②** "**not necessarily flushed to disk**" — the durability unit is *replicas holding it*, not *bytes on platters* (Ch. 6 §5.3).
* **④** "as long as **at least one** replica remains alive" — lose all of them simultaneously and committed data is gone. This is why rack/AZ placement is a durability control, not a nicety.
> *"These basic guarantees can be used while building a reliable system, but **in themselves, they don't make the system fully reliable.**"*
**The trade-off space, named explicitly:**
***
# 7.8 The reliability configuration matrix (/docs/kafka/reliable-data-delivery/reliability-configuration-matrix)
#### The dependency graph — why these settings only work together [#the-dependency-graph--why-these-settings-only-work-together]
#### Monitoring summary [#monitoring-summary]
| Layer | Signal | Why |
| ---------- | --------------------------------------------------------------------------------- | -------------------------------------------------------------------------- |
| Broker | **UnderReplicatedPartitions** | Effective RF dropped — the #6 trap |
| Broker | **OfflinePartitions** | Unavailable data |
| Broker | `FailedProduceRequestsPerSec` / `FailedFetchRequestsPerSec` (**tagged by error**) | Distinguishes benign `NOT_LEADER_FOR_PARTITION` from `NOT_ENOUGH_REPLICAS` |
| Broker | ISR shrink/expand rate | Flapping = GC/network trouble |
| Producer | **error-rate, retry-rate per record** | *"the two metrics most important for reliability"* |
| Producer | WARN logs with **"0 attempts left"** | Out of retries |
| Producer | ERROR logs | Complete failure: non-retriable, exhausted, or timeout |
| Consumer | **Consumer lag** (trend, not threshold) | Use Burrow |
| End-to-end | produced/sec vs consumed/sec reconciliation | The only way to *prove* nothing was lost |
| End-to-end | produce-timestamp → consume-timestamp latency | Verifies "timely" against business requirements |
***
# 7.2 Replication recap — and what "in sync" precisely means (/docs/kafka/reliable-data-delivery/replication-recap-sync-precisely)
> *"Kafka's replication mechanism, with its multiple replicas per partition, is **at the core of all of Kafka's reliability guarantees.**"*
**Foundations:** a partition is stored **on a single disk**; ordering is guaranteed within it; a partition is either **online (available)** or **offline (unavailable)**; all events are produced to the leader and *usually* consumed from it; if the leader becomes unavailable, **one of the in-sync replicas** becomes the new leader (with one exception — unclean election, §3.2).
#### The three conditions for being in-sync [#the-three-conditions-for-being-in-sync]
A replica is in sync if it is **the leader**, or if it is a follower that:
**Condition ③ is the subtle one.** A follower that steadily fetches but is *permanently 5 seconds behind* is **out of sync** — it never achieves zero lag. Continuous progress is not sufficient; **momentary full catch-up is required.**
**Getting back in:** *"An out-of-sync replica gets back into sync when it connects to ZooKeeper again **and catches up to the most recent message written to the leader.** This usually happens quickly after a temporary network glitch is healed **but can take a while if the broker the replica is stored on was down for a longer period.**"*
#### 💡 The counterintuitive performance effect of falling out of sync [#-the-counterintuitive-performance-effect-of-falling-out-of-sync]
> *"**An in-sync replica that is slightly behind can slow down producers and consumers** — since they wait for all the in-sync replicas to get the message before it is committed. **Once a replica falls out of sync, we no longer wait for it to get messages. It is still behind, but now THERE IS NO PERFORMANCE IMPACT.** The catch is that with fewer in-sync replicas, **the effective replication factor of the partition is lower, and therefore there is a higher risk for downtime or data loss.**"*
> ### OUT-OF-SYNC REPLICA FLAPPING — historical context [#out-of-sync-replica-flapping--historical-context]
>
> \*"In older versions of Kafka, it was **not uncommon to see one or more replicas rapidly flip between in-sync and out-of-sync status. This was a sure sign that something was wrong with the cluster.** A relatively common cause was **a large maximum request size and large JVM heap that required tuning to prevent long garbage collection pauses that would cause the broker to temporarily disconnect from ZooKeeper.**
>
> These days the problem is **very rare**, especially with **Kafka 2.5.0+** and its default ZooKeeper connection timeout and maximum replica lag. **JVM 8+** (now the minimum supported) **with G1** helped curb this — *"although tuning may still be required for large messages."*
>
> *"Generally speaking, **Kafka's replication protocol became significantly more reliable in the years since the first edition.**"* References: Jason Gustafson, *"Hardening Apache Kafka Replication"*; Gwen Shapira, *"Please Upgrade Apache Kafka Now."*
***
# 7.9 Self-test (/docs/kafka/reliable-data-delivery/self-test)
State Kafka's reliability guarantees precisely. What are the qualifying clauses in the ordering guarantee and the durability guarantee?
What exactly does "committed" mean — and what does it explicitly *not* mean?
Give all three conditions for a replica to be in sync. Which one surprises people, and why?
A replica falls out of the ISR and your p99 latency *improves*. Explain what just happened and why it's dangerous.
Give the formula for replication traffic. Why is this the number Ch. 2 said people forget?
Someone proposes RF=2 because the SAN already triple-replicates blocks. What's right and what's wrong about that reasoning?
RF=3 and a single top-of-rack switch failure took a partition offline. How, and what's the fix?
Describe both scenarios that leave a partition with no in-sync replica. Which one should change your default configuration, and how?
Walk through the offset 100–200 inconsistency caused by unclean leader election. What happens to the old leader's data when it returns?
Why does `acks=all` provide almost no protection with `min.insync.replicas=1`?
What exactly happens when in-sync replicas drop below `min.insync.replicas`? What can consumers still do?
Explain why "a produce request that fails loudly" is better than "an acknowledged write you'll lose."
Kafka doesn't fsync on every write. State the bet it's making, and the condition under which that bet is bad.
`replica.lag.time.max.ms` went from 10 s to 30 s in 2.5.0. Name the benefit and the hidden cost.
With perfect broker configuration, describe the two ways a producer can still lose data.
Why is `acks=0` popular in benchmarks and misleading as a latency result?
`LEADER_NOT_AVAILABLE` vs `INVALID_CONFIG`: which is retriable, and what determines the category?
Producer retries give you which delivery guarantee? What upgrades it, and how?
Name four error categories the producer will *not* handle for you.
When is autocommit actually safe, and what single change makes it unsafe?
Why are offset commits more expensive than they look, and why does the cost concentrate on one broker?
Record 30 failed, 31 succeeded. Why can't you commit 31? Give both remediation patterns and their trade-offs.
Your consumer restarts at the correct offset but produces wrong aggregates. What's missing, and what are the two proper solutions?
What do `VerifiableProducer`/`VerifiableConsumer` do, and what's the pass criterion for a failure test?
List the four configuration-validation scenarios. Which one produces a number you must feed into a producer config?
What is a "brown out," and why is it harder to handle than an outright failure?
Why is *"no more than 1,000 duplicate values"* a better test expectation than *"no duplicates"*?
Why are static thresholds on consumer lag a bad alert design, and what's the alternative?
What must you instrument to *prove* no messages were lost end to end? What's the state of open-source tooling for it?
`FailedProduceRequestsPerSec` is rising. Name one benign cause and one serious one, and how you'd tell them apart.
Draw the reliability dependency chain from `replication.factor` through to transactions. Pick any one link and explain what breaks if you omit it.
**Previous:** [Chapter 6 — Kafka Internals](06-kafka-internals.md)
**Next:** [Chapter 8 — Exactly-Once Semantics](08-exactly-once-semantics.md)
# 7.5 Using consumers reliably (/docs/kafka/reliable-data-delivery/using-consumers-reliably)
**The clean division of responsibility:**
> *"data is only available to consumers **after it has been committed** to Kafka... This means that **consumers get data that is guaranteed to be consistent. The only thing consumers are left to do is make sure they keep track of which messages they've read and which messages they haven't.**"*
**How consumption works mechanically:** *"a consumer is fetching a batch of messages, checking the last offset in the batch, and then requesting another batch starting from the last offset received. This guarantees that a Kafka consumer will **always get new data in correct order without missing any messages.**"*
#### ⚠️ The single way consumers lose messages [#️-the-single-way-consumers-lose-messages]
> *"**The main way consumers can lose messages is when committing offsets for events they've read but haven't COMPLETELY PROCESSED yet.** This way, when another consumer picks up the work, **it will SKIP those messages and they will NEVER GET PROCESSED.** This is why **paying careful attention to when and how offsets get committed is critical.**"*
> ### COMMITTED MESSAGES vs COMMITTED OFFSETS — don't conflate them [#committed-messages-vs-committed-offsets--dont-conflate-them]
>
> | Term | Meaning |
> | --------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
> | **Committed message** | *"a message that was **written to all in-sync replicas** and is available to consumers"* |
> | **Committed offset** | *"offsets **the consumer sent to Kafka** to acknowledge that it **received and processed** all the messages in a partition up to this specific offset"* |
#### 5.1 The four consumer configs that determine reliability [#51-the-four-consumer-configs-that-determine-reliability]
##### `group.id` [#groupid]
> *"if two consumers have the same group ID and subscribe to the same topic, each will be assigned a subset of the partitions... **If we need a consumer to see, on its own, EVERY SINGLE MESSAGE in the topics it is subscribed to, it will need a UNIQUE `group.id`.**"*
##### `auto.offset.reset` [#autooffsetreset]
> *"There are only **two** options here"* (in the reliability framing):
The reliability-first default is **`earliest`**. (Ch. 4 adds a third option, `none`, which throws — often the best choice when neither loss nor mass reprocessing is acceptable and you want a human decision.)
##### `enable.auto.commit` — the real decision [#enableautocommit--the-real-decision]
> *"This is a **big decision**: are we going to let the consumer commit offsets for us based on schedule, or are we planning on committing offsets manually?"*
**This is a genuinely useful clarification.** Autocommit is not sloppy *per se* — it's correct for the strictly-in-poll-loop case and **unsafe the instant processing escapes the poll loop.** Handing records to a thread pool or an async client silently converts autocommit into a data-loss mechanism.
##### `auto.commit.interval.ms` (default **5 s**) [#autocommitintervalms-default-5-s]
> *"committing more frequently adds overhead but reduces the number of duplicates that can occur when a consumer stops."*
##### And a fifth, indirect factor [#and-a-fifth-indirect-factor]
> *"While not directly related to reliable data processing, **it is difficult to consider a consumer reliable if it frequently stops consuming in order to rebalance.**"* (→ Ch. 4: cooperative sticky assignor, static membership, `max.poll.records`.)
#### 5.2 Explicit offset commits — five rules [#52-explicit-offset-commits--five-rules]
##### Rule 1: Always commit offsets AFTER messages were processed [#rule-1-always-commit-offsets-after-messages-were-processed]
> *"If we do all the processing within the poll loop and don't maintain state between poll loops (e.g., for aggregation), this should be easy... **If there are additional threads or stateful processing involved, this becomes more complex, especially since THE CONSUMER OBJECT IS NOT THREAD SAFE.**"*
##### Rule 2: Commit frequency trades performance against duplicates [#rule-2-commit-frequency-trades-performance-against-duplicates]
> *"**Committing has significant performance overhead. It is similar to produce with `acks=all`, BUT ALL OFFSET COMMITS OF A SINGLE CONSUMER GROUP ARE PRODUCED TO THE SAME BROKER, WHICH CAN BECOME OVERLOADED.**"*
>
> *"**Committing after every message should only ever be done on very low-throughput topics.**"*
##### Rule 3: Commit the RIGHT offsets at the RIGHT time [#rule-3-commit-the-right-offsets-at-the-right-time]
> *"A common pitfall when committing in the middle of the poll loop is **accidentally committing the last offset READ when polling and not the offset AFTER the last offset PROCESSED.** Remember that it is **critical to always commit offsets for messages AFTER they were processed — committing offsets for messages read but not processed can lead to the consumer MISSING messages.**"*
##### Rule 4: Handle rebalances [#rule-4-handle-rebalances]
> *"consumer rebalances **will** happen, and we need to handle them properly. This usually involves **committing offsets before partitions are revoked** and **cleaning any state the application maintains when it is assigned new partitions.**"* (→ Ch. 4's `ConsumerRebalanceListener`.)
##### Rule 5: Consumers may need to retry — and Kafka has no per-message ack [#rule-5-consumers-may-need-to-retry--and-kafka-has-no-per-message-ack]
**The fundamental constraint:**
> *"unlike traditional pub/sub messaging systems, **Kafka consumers commit offsets and do NOT 'ack' individual messages.** This means that **if we failed to process record #30 and succeeded in processing record #31, we should NOT commit offset #31 — this would result in marking as processed all the records up to #31 INCLUDING #30**, which is usually not what we want."*
**The two patterns:**
**Trade-off:** Pattern A preserves ordering but **head-of-line blocks** (nothing progresses while you retry). Pattern B keeps the main stream flowing but **abandons ordering** for the failed records.
##### Rule 6: Consumers may need to maintain state [#rule-6-consumers-may-need-to-maintain-state]
**The moving-average example:**
> *"if we want to calculate a moving average, we'll want to update the average every time we poll Kafka for new messages. **If our process is restarted, we will need to not just start consuming from the last offset, but we'll also need to RECOVER THE MATCHING MOVING AVERAGE.**"*
**The DIY approach and its limitation:**
> *"One way to do this is to **write the latest accumulated value to a 'results' topic AT THE SAME TIME the application is committing the offset.**"* → so a starting thread picks up the latest accumulated value and resumes.
>
> *"In Chapter 8, we discuss how an application can **write results and commit offsets in a SINGLE TRANSACTION.** **In general, this is a rather complex problem to solve, and we recommend looking at a library like KAFKA STREAMS or FLINK**, which provides high-level DSL-like APIs for aggregation, joins, windows, and other complex analytics."*
*(This is the same "offsets are only half your position" point Ch. 5 made with the shoe-counting example — and it's the motivation for Ch. 8's transactions and Ch. 14's Streams.)*
***
# 7.4 Using producers reliably (/docs/kafka/reliable-data-delivery/using-producers-reliably)
> *"**Even if we configure the brokers in the most reliable configuration possible, the system as a whole can still potentially lose data if we don't configure the producers to be reliable as well.**"*
#### 4.1 The two loss scenarios — with the *perfect* broker configuration [#41-the-two-loss-scenarios--with-the-perfect-broker-configuration]
**The two rules that follow:**
```txt
① Use the correct `acks` configuration to match reliability requirements
② HANDLE ERRORS CORRECTLY — both in configuration AND in code
```
#### 4.2 The three ack modes, restated with reliability precision [#42-the-three-ack-modes-restated-with-reliability-precision]
| | What it means | What you lose |
| -------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **`acks=0`** | *"a message is considered written successfully **if the producer managed to send it over the network**"* | *"We will still get errors if the object cannot be **serialized** or if the **network card failed**, but **we won't get any error if the partition is offline, a leader election is in progress, or even if the ENTIRE KAFKA CLUSTER IS UNAVAILABLE.**"* |
| **`acks=1`** | *"the leader will send either an acknowledgment or an error **the moment it gets the message and writes it to the partition data file (but not necessarily synced to disk)**"* | Scenario 1 above. **Plus:** *"it is also possible to **write to the leader faster than it can replicate** messages and end up with **under-replicated partitions**, since the leader will acknowledge messages from the producer before replicating them."* |
| **`acks=all`** | *"the leader will wait until **all in-sync replicas** get the message before sending back an acknowledgment or an error. **In conjunction with `min.insync.replicas`**, this lets us control how many replicas get the message before it is acknowledged."* | *"the option with the **longest producer latency**"* |
> ### 💡 The `acks=0` benchmark warning [#-the-acks0-benchmark-warning]
>
> *"Running with `acks=0` has low produce latency (**which is why we see a lot of benchmarks with this configuration**), but **it will NOT improve end-to-end latency** (remember that consumers will not see messages until they are replicated to all available replicas)."*
That parenthetical is worth quoting to anyone waving a Kafka benchmark at you. And it restates Ch. 3's core insight: *lower acks buys producer-side latency only, never end-to-end latency.*
#### 4.3 Producer retries [#43-producer-retries]
**Retriable vs non-retriable, with concrete error codes:**
**The recommended configuration:**
> *"when our goal is to never lose a message, our best approach is to configure the producer to **keep trying** to send the messages when it encounters a retriable error. And the best approach to retries... is to **leave the number of retries at its current default (`MAX_INT`, or effectively infinite)** and use **`delivery.timeout.ms`** to configure the maximum amount of time we are willing to wait until giving up — the producer will retry sending the message as many times as possible within this time interval."*
#### ⚠️ Retries guarantee at-least-once, NOT exactly-once [#️-retries-guarantee-at-least-once-not-exactly-once]
> *"Retrying to send a failed message includes a risk that **both messages were successfully written to the broker, leading to duplicates.** Retries and careful error handling can guarantee that **each message will be stored AT LEAST ONCE, but not EXACTLY ONCE.** Using **`enable.idempotence=true`** will cause the producer to include additional information in its records, which brokers will use to **skip duplicate messages caused by retries.**"*
#### 4.4 Errors YOU must handle [#44-errors-you-must-handle]
*"Using the built-in producer retries is an easy way to correctly handle a large variety of errors without loss of messages, but as developers, we must still be able to handle other types of errors:"*
**The design questions the book poses — note it refuses to answer them for you:**
> *"do we **throw away 'bad messages'**? **Log errors**? **Stop reading messages from the source system**? **Apply back pressure to the source system** to stop sending messages for a while? **Store these messages in a directory on the local disk**? These decisions **depend on the architecture and the product requirements.**"*
> **The one clear rule:** *"**Just note that if all the error handler is doing is retrying to send the message, then we'll be better off relying on the producer's retry functionality.**"*
Hand-rolled retry loops on top of the producer's retry machinery are strictly worse: they duplicate delivery, break the `delivery.timeout.ms` accounting, and can reorder messages.
***
# 7.6 Validating system reliability (/docs/kafka/reliable-data-delivery/validating-system-reliability)
> *"Once we have gone through the process of figuring out our reliability requirements, configuring the brokers, configuring the clients, and using the APIs in the best way for our use case, **we can just relax and run everything in production, confident that no event will ever be missed, right?**"*
**Three layers of validation:**
#### 6.1 Validating configuration [#61-validating-configuration]
**Two reasons to test config in isolation from application logic:**
**The tools:** `org.apache.kafka.tools` includes **`VerifiableProducer`** and **`VerifiableConsumer`** — *"These can run as command-line tools or be embedded in an automated testing framework."*
**The four scenarios the book says to test:**
| Test | The question to answer |
| --------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Leader election** | *"what happens if we **kill the leader**? **How long** does it take the producer and consumer to start working as usual again?"* |
| **Controller election** | *"how long does it take the system to resume after a **restart of the controller**?"* |
| **Rolling restart** | *"can we restart the brokers **one by one without losing any messages**?"* |
| **Unclean leader election** | *"what happens when we **kill all the replicas for a partition one by one (to make sure each goes out of sync)** and then **start a broker that was out of sync**? **What needs to happen in order to resume operations? IS THIS ACCEPTABLE?**"* |
**The leader-election test is not optional trivia** — its answer is the number you plug into `delivery.timeout.ms` (Ch. 3: *"it typically takes leader election 30 seconds... let's keep retrying for 120 seconds"*). You are supposed to **measure** it, not guess it.
> *"The Apache Kafka source repository includes an **extensive test suite**. Many of the tests are based on the same principle and use the verifiable producer and consumer to make sure **rolling upgrades** work."*
#### 6.2 Validating applications [#62-validating-applications]
> *"This will check things like **custom error-handling code, offset commits, and rebalance listeners** and similar places where the application logic interacts with Kafka's client libraries."*
**The eight failure conditions to test under:**
**Fault injection:** *"Apache Kafka itself includes the **Trogdor test framework** for fault injection."*
**The methodology — write the expectation first:**
> *"**For each scenario, we will have EXPECTED BEHAVIOR, which is what we planned on seeing when we developed the application. Then we run the test to see what ACTUALLY happens.**"*
>
> **Example:** *"when planning for a rolling restart of consumers, we planned for **a short pause as consumers rebalance and then continue consumption with NO MORE THAN 1,000 DUPLICATE VALUES.** Our test will show whether the way the application commits offsets and handles rebalances actually works this way."*
Note the shape of that expectation: **a quantified duplicate budget.** "No duplicates" is not a testable expectation for an at-least-once system; "≤1,000 duplicates" is.
#### 6.3 Monitoring reliability in production [#63-monitoring-reliability-in-production]
> *"Testing the application is important, **but it does not replace the need to continuously monitor production systems to make sure data is flowing as expected.**"*
##### Producer-side metrics [#producer-side-metrics]
> *"For the producers, the **two metrics most important for reliability are ERROR-RATE and RETRY-RATE PER RECORD (aggregated).** Keep an eye on those, since error or retry rates going up can indicate an issue."*
**The log lines to watch, and how to read them:**
*(Note the **correlation id** in the log line — that's the request header field from Ch. 6 §5, and it's how you match this WARN to a broker-side log entry.)*
> *"Of course, **it is always better to solve the problem that caused the errors in the first place**"* — rather than just extending `delivery.timeout.ms`.
##### Consumer-side: lag [#consumer-side-lag]
> *"the most important metric is **CONSUMER LAG**... **Ideally, the lag would always be zero.** In practice, because calling `poll()` returns multiple messages and then the consumer spends time processing them before fetching more, **the lag will always fluctuate a bit.** **What is important is to make sure consumers do EVENTUALLY CATCH UP rather than fall further and further behind.**"*
>
> *"**Because of the expected fluctuation in consumer lag, setting traditional alerts on the metric can be challenging.** **Burrow** is a consumer lag checker by LinkedIn and can make this easier."*
##### End-to-end flow monitoring [#end-to-end-flow-monitoring]
> *"Monitoring flow of data also means **making sure all produced data is consumed in a timely manner** ('timely manner' is usually based on **business requirements**). In order to make sure data is consumed in a timely manner, **we need to know when the data was produced.**"*
**Kafka helps:** *"starting with version 0.10.0, **all messages include a timestamp** that indicates when the event was produced (**although note that this can be overridden either by the application that is sending the events or by the brokers themselves if they are configured to do so**)."*
*(That's the create-time vs append-time distinction from Ch. 6's batch header. If brokers are set to append-time, your "produce latency" measurement is measuring the wrong thing.)*
**What you must build:**
> **The honest caveat:** *"This type of end-to-end monitoring system can be **challenging and time-consuming to implement. To the best of our knowledge, THERE IS NO OPEN SOURCE IMPLEMENTATION of this type of system**, but Confluent provides a commercial implementation as part of the **Confluent Control Center.**"*
##### Broker-side error metrics [#broker-side-error-metrics]
```txt
kafka.server:type=BrokerTopicMetrics,name=FailedProduceRequestsPerSec
kafka.server:type=BrokerTopicMetrics,name=FailedFetchRequestsPerSec
```
> *"**At times, some level of error responses is expected** — for example, if we shut down a broker for maintenance and new leaders are elected on another broker, it is expected that producers will receive a **`NOT_LEADER_FOR_PARTITION`** error, which will cause them to request updated metadata before continuing as usual. **Unexplained increases in failed requests should ALWAYS be investigated.** To assist in such investigations, **the failed requests metrics are TAGGED WITH THE SPECIFIC ERROR RESPONSE that the broker sent.**"*
The tagging is the important operational detail: `NOT_LEADER_FOR_PARTITION` during a rolling restart is *expected*; the same counter rising with `NOT_ENOUGH_REPLICAS` means your `min.insync.replicas` floor is being hit and producers are being rejected.
***
# 11.10 What actually breaks in production — Ch. 11 consolidated (/docs/kafka/securing-kafka/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | ------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------- |
| 1 | **Anyone can read/write; no identity in logs** | PLAINTEXT listener → principal is `User:ANONYMOUS` | SSL or SASL\_SSL listener |
| 2 | **Clients on an SSL listener show as `User:ANONYMOUS`** | `ssl.client.auth=requested` (not `required`) and the client has no key store | Use `required` if you need identity |
| 3 | **Man-in-the-middle possible** | Hostname verification disabled to "fix" a cert problem | *"Should NOT be disabled in production."* Fix the SAN/CN, or use `client.dns.lookup` (Ch. 5) |
| 4 | **TLS handshake failures appear overnight** | **Certificates expired** | Rotate before expiry; broker stores are **dynamically updatable** via configs tool / Admin API |
| 5 | **Inter-broker TLS fails after enabling client auth** | Broker trust store lacks the **client** CA (or the broker CA) | *"Broker trust stores should include the CA of the broker certificates AS WELL AS the CA of the client certificates"* |
| 6 | **Throughput drops 20–30% after enabling SSL** | **Zero-copy is not supported for SSL** | Expected. Consider where encryption is truly needed (Ch. 10 §7.2) |
| 7 | **Broker network threads saturated; clients can't connect** | **TLS handshake DoS** — handshakes run on network threads | Connection quotas/limits + `connection.failed.authentication.delay.ms` |
| 8 | **Private keys readable by other users** | Key stores are plain files on disk | **Filesystem permissions** on all key/trust stores and keytabs |
| 9 | **Kerberos auth fails intermittently** | Forward/reverse DNS mismatch | `rdns=false` in client `krb5.conf`; secure DNS is a **requirement** |
| 10 | **All clients fail to authenticate at once** | **KDC or DNS outage** (or DoS) | *"It is NECESSARY to monitor the availability of these services"* |
| 11 | **Kerberos replay detection breaks / auth fails after clock drift** | Clock skew beyond configured variability | Secure, monitored NTP — **clock sync is part of the security perimeter** |
| 12 | **Adding one user requires restarting every broker** | SASL/PLAIN's built-in JAAS password store | Custom **server callback handler** → external password server |
| 13 | **Passwords visible in logs** | Config not declared as `PASSWORD` type | Use `ConfigDef` `PASSWORD` type; externalize/encrypt |
| 14 | **Password rotation causes an outage** | Server accepts only one password at a time | Callback that accepts **old and new** for an overlap window + reauthentication |
| 15 | **Credentials stolen off the wire** | SASL/PLAIN or SASL/SCRAM over **SASL\_PLAINTEXT** | **Always SASL\_SSL** — PLAIN sends clear text; SCRAM exposes hashed keys during handshake |
| 16 | **SCRAM credentials stolen from ZooKeeper** | ZooKeeper not SSL-enabled / disk not encrypted | Both are stated **requirements** for production SCRAM |
| 17 | **A deleted user keeps working** | Existing connections survive user deletion | **`connections.max.reauth.ms`**; Deny ACL for immediate effect |
| 18 | **Compromised user still active after removal** | No reauth interval; **SSL renegotiation is not supported** so SSL connections never re-verify | **Deny ACL** — *"the quickest way to disable access"* — plus reauth config |
| 19 | **Compromised super user can't be revoked quickly** | `super.users` **cannot be denied** and requires a **broker restart** to change | Don't use `super.users` in production; grant explicit ACLs |
| 20 | **`super.users` list parsed wrongly** | Used commas; DNs contain commas | **Semicolon**-separated |
| 21 | **Adding an ACL unexpectedly revoked others' access** | `allow.everyone.if.no.acl.found=true`, and a new prefix/wildcard ACL made `no.acl.found` false | Don't use it in production |
| 22 | **New topics silently world-accessible** | Same config | Same fix |
| 23 | **OAuth "works" in staging but is insecure** | Built-in OAUTHBEARER uses **unsecured JWTs** and **does not validate tokens** | Custom login + **server validator** callbacks against a real OAuth server |
| 24 | **Connections outlive their OAuth tokens** | No reauthentication | `connections.max.reauth.ms` + token revocation |
| 25 | **A delegation token was used to mint more tokens** | It can't be | *"Clients authenticated using delegation tokens CANNOT create other delegation tokens"* |
| 26 | **All delegation tokens broke** | **Master key rotated** — requires restarting all brokers and deleting existing tokens | Plan the rotation: delete tokens → update key on all brokers → restart → recreate |
| 27 | **Idempotent producer fails authorization** | Missing **`Cluster:IdempotentWrite`** (non-transactional only) | Grant it |
| 28 | **Transactional producer fails authorization** | Missing **`TransactionalId:Write`** and/or **`Group:Read`** | Grant both (Ch. 8) |
| 29 | **Consumer can fetch but not join a group** | Has `Topic:Read` but not **`Group:Read`** | Grant `Group:Read` |
| 30 | **A client was granted unintended broker powers** | `Cluster:ClusterAction` granted to a non-broker | *"Should ONLY be granted to brokers"* |
| 31 | **Unmanageable ACL sprawl** | Per-resource literal ACLs at scale | **Prefixed** ACLs by department + group/role principals via a custom authorizer |
| 32 | **Departing employee's credentials still power a service** | Application used a **personal** principal | *"Long-running applications can be configured with SERVICE credentials"* |
| 33 | **A reused principal name inherits old access** | Principal reuse | *"Reuse of principals must be AVOIDED"* |
| 34 | **No record of who accessed what** | Grants log at **DEBUG**, only denials at **INFO** | Enable DEBUG on `kafka.authorizer.logger` if you need a full trail |
| 35 | **Sensitive data found in a broker heap dump** | TLS + disk encryption don't cover **broker memory** | **End-to-end encryption** (serializer/deserializer + KMS) |
| 36 | **Cloud provider / platform admin could read customer data** | Broker sees plaintext | End-to-end encryption — *"brokers never see the unencrypted contents"* |
| 37 | **Compression gives no benefit and adds CPU** | Compressing **after** encryption (high-entropy data) | Compress **before** encrypting; **disable Kafka compression** |
| 38 | **Partitioning and compaction break after encrypting keys** | Encrypted keys aren't **hash-stable** | Message key = **secure hash** of the original; encrypted key in header/payload via a **producer interceptor** |
| 39 | **Key rotation needs a maintenance window** | Compacted topics retain old-key messages indefinitely; re-encryption requires **producers and consumers offline** | Plan it; keep old keys available for the retention period |
| 40 | **ZooKeeper ACLs written by one broker exclude others** | Full Kerberos principals differ per broker | `kerberos.removeHostFromPrincipal=true` + `kerberos.removeRealmFromPrincipal=true` |
| 41 | **Unexpected ZooKeeper access granted** | ZK with SASL **and** SSL client auth associates **multiple principals**; **any** may grant | Understand the model; audit ZK ACLs |
| 42 | **Anyone can read Kafka metadata from ZooKeeper** | `zookeeper.set.acl` not enabled | Enable it — metadata becomes broker-writable only; sensitive paths (SCRAM) are not world-readable |
| 43 | **DIGEST-MD5 used in production** | It has *"known security vulnerabilities"* | Kerberos or TLS mutual auth |
| 44 | **Flag-day protocol migration caused an outage** | Changed the listener protocol in place | Add a **new listener on a new port**, migrate clients, then remove the old |
***
# 11.7 Auditing (/docs/kafka/securing-kafka/auditing)
> *"Kafka brokers can be configured to generate comprehensive **log4j** logs for auditing and debugging."*
**Two independently configurable loggers:**
#### ⚠️ The log-level asymmetry — this is the important part [#️-the-log-level-asymmetry--this-is-the-important-part]
**Example authorizer log lines:**
```txt
DEBUG Principal = User:Alice is Allowed Operation = Write from host = 127.0.0.1
on resource = Topic:LITERAL:customerOrders for request = Produce
with resourceRefCount = 1 (kafka.authorizer.logger)
INFO Principal = User:Mallory is Denied Operation = Describe from host = 10.0.0.13
on resource = Topic:LITERAL:customerOrders for request = Metadata
with resourceRefCount = 1 (kafka.authorizer.logger)
```
**Request logging:**
```txt
DEBUG Completed request:RequestHeader(apiKey=PRODUCE, apiVersion=8,
clientId=producer-1, correlationId=6) -- {acks=-1,timeout=30000,
partitionSizes=[customerOrders-0=15514]},response:{...base_offset=13...},
... totalTime:2.42,requestQueueTime:0.112,localTime:2.15,remoteTime:0.0,
throttleTime:0,responseQueueTime:0.04,sendTime:0.118,
securityProtocol:SASL_SSL,principal:User:Alice,listener:SASL_SSL,
clientInformation:ClientInformation(softwareName=apache-kafka-java,
softwareVersion=2.7.0-SNAPSHOT) (kafka.request.logger)
```
**Notice what that single line gives you:** the principal, the listener, the security protocol, the client software *and version*, the correlation ID (Ch. 6 §5), **and the full request-pipeline timing breakdown** (`requestQueueTime`, `localTime`, `remoteTime`, `throttleTime`, `responseQueueTime`, `sendTime` — the exact stages from Ch. 6's threading diagram). It's simultaneously an audit record and a latency trace.
> \*"**Authorizer and request logs can be analyzed to DETECT SUSPICIOUS ACTIVITIES. METRICS THAT TRACK AUTHENTICATION FAILURES, as well as AUTHORIZATION FAILURE LOGS, can be extremely useful for auditing and provide valuable information IN THE EVENT OF AN ATTACK OR UNAUTHORIZED ACCESS.**
>
> For **end-to-end auditability and traceability** of messages, **audit metadata can be included in MESSAGE HEADERS when messages are produced. END-TO-END ENCRYPTION CAN BE USED TO PROTECT THE INTEGRITY OF THIS METADATA.**"\*
***
# 11.3 Authentication (/docs/kafka/securing-kafka/authentication)
> *"Once authenticated, **Alice's identity is associated with the connection throughout the LIFETIME of the connection.** Kafka uses an instance of **`KafkaPrincipal`** to represent client identity and uses this principal to **grant access to resources and allocate quotas.**"*
`KafkaPrincipal` is established during authentication based on the protocol (e.g. `User:Alice`) and can be customized via **`principal.builder.class`**.
> ### ANONYMOUS CONNECTIONS [#anonymous-connections]
>
> *"The principal **`User:ANONYMOUS`** is used for unauthenticated connections. This includes **clients on PLAINTEXT listeners** as well as **unauthenticated clients on SSL listeners.**"*
#### 3.1 SSL authentication [#31-ssl-authentication]
> *"When a connection is established over TLS, the TLS handshake process **performs authentication, negotiates cryptographic parameters, and generates shared keys for encryption.** The server's digital certificate is verified by the client to establish the identity of the server. If client authentication using SSL is enabled, the server also verifies the client's digital certificate."*
> ### ⚠️ SSL PERFORMANCE [#️-ssl-performance]
>
> *"SSL channels are encrypted and hence introduce a **noticeable overhead in terms of CPU usage. ZERO-COPY TRANSFER IS CURRENTLY NOT SUPPORTED FOR SSL.** Depending on the traffic pattern, **the overhead may be up to 20–30%.**"*
*(This is the same zero-copy loss from Ch. 6 §5.4 — and it's exactly why Ch. 10 §7.2 says to consider consuming locally and producing remotely when only the WAN hop needs encryption.)*
##### What goes where [#what-goes-where]
> *"Broker certificates should contain the **broker hostname** as a **Subject Alternative Name (SAN)** extension or as the **Common Name (CN)** to enable clients to verify the server hostname. **Wildcard certificates can be used to simplify administration** by using the same key store for all brokers in a domain."*
> ### ⚠️ SERVER HOSTNAME VERIFICATION [#️-server-hostname-verification]
>
> *"By default, Kafka clients verify that the hostname of the server stored in the server certificate **matches the host that the client is connecting to.** The connection hostname may be **a bootstrap server** the client is configured with **or an ADVERTISED LISTENER hostname returned by a broker in a metadata response. HOSTNAME VERIFICATION IS A CRITICAL PART OF SERVER AUTHENTICATION THAT PROTECTS AGAINST MAN-IN-THE-MIDDLE ATTACKS AND HENCE SHOULD NOT BE DISABLED IN PRODUCTION SYSTEMS.**"*
*(This is the flip side of Ch. 5 §3.1's DNS-alias problem: hostname verification is *why* a DNS alias breaks SASL, and disabling verification is the wrong fix.)*
##### Client authentication modes [#client-authentication-modes]
> *"By default, the **distinguished name (DN) of the client certificate is used as the `KafkaPrincipal`** for authorization and quotas. The configuration option **`ssl.principal.mapping.rules`** can be used to provide a list of rules to customize the principal."*
> ⚠️ *"If SSL is used for inter-broker communication, **broker trust stores should include the CA of the BROKER certificates AS WELL AS the CA of the CLIENT certificates.**"*
##### Certificate generation — the full sequence [#certificate-generation--the-full-sequence]
**Step 1 — self-signed CA for brokers:**
```bash
# Create the CA key-pair; "We use this for signing certificates."
keytool -genkeypair -keyalg RSA -keysize 2048 -keystore server.ca.p12 \
-storetype PKCS12 -storepass server-ca-password -keypass server-ca-password \
-alias ca -dname "CN=BrokerCA" -ext bc=ca:true -validity 365
# Export the CA's public certificate; goes into trust stores + cert chains
keytool -export -file server.ca.crt -keystore server.ca.p12 \
-storetype PKCS12 -storepass server-ca-password -alias ca -rfc
```
**Step 2 — broker key store with a CA-signed certificate:**
```bash
# ① generate the broker's private key
keytool -genkey -keyalg RSA -keysize 2048 -keystore server.ks.p12 \
-storepass server-ks-password -keypass server-ks-password -alias server \
-storetype PKCS12 -dname "CN=Kafka,O=Confluent,C=GB" -validity 365
# ② certificate signing request
keytool -certreq -file server.csr -keystore server.ks.p12 -storetype PKCS12 \
-storepass server-ks-password -keypass server-ks-password -alias server
# ③ sign it with the CA — NOTE THE SAN EXTENSION (hostname verification!)
keytool -gencert -infile server.csr -outfile server.crt \
-keystore server.ca.p12 -storetype PKCS12 -storepass server-ca-password \
-alias ca -ext SAN=DNS:broker1.example.com -validity 365
# ④ import the certificate CHAIN back into the broker key store
cat server.crt server.ca.crt > serverchain.crt
keytool -importcert -file serverchain.crt -keystore server.ks.p12 \
-storepass server-ks-password -keypass server-ks-password -alias server \
-storetype PKCS12 -noprompt
```
> *"If using wildcard hostnames, the same key store can be used for all brokers. **Otherwise, create a key store for EACH broker with its fully qualified domain name (FQDN).**"*
**Step 3 — trust stores:**
```bash
# broker trust store (for inter-broker TLS): contains the BROKER CA
keytool -import -file server.ca.crt -keystore server.ts.p12 \
-storetype PKCS12 -storepass server-ts-password -alias server -noprompt
# client trust store: contains the BROKER CA
keytool -import -file server.ca.crt -keystore client.ts.p12 \
-storetype PKCS12 -storepass client-ts-password -alias ca -noprompt
```
**Step 4 — client CA + client key store (only if `ssl.client.auth` is on):**
```bash
# a SEPARATE CA for clients
keytool -genkeypair ... -keystore client.ca.p12 -dname CN=ClientCA -ext bc=ca:true
keytool -export -file client.ca.crt -keystore client.ca.p12 ...
# client key store; the DN becomes the principal:
# User:CN=Metrics App,O=Confluent,C=GB
keytool -genkey ... -keystore client.ks.p12 \
-dname "CN=Metrics App,O=Confluent,C=GB" -validity 365
keytool -certreq ... ; keytool -gencert ... ;
cat client.crt client.ca.crt > clientchain.crt
keytool -importcert -file clientchain.crt -keystore client.ks.p12 ...
# ⚠ "The broker's trust store should contain the CAs of ALL clients."
keytool -import -file client.ca.crt -keystore server.ts.p12 -alias client ...
```
##### Broker and client TLS configuration [#broker-and-client-tls-configuration]
```properties
# BROKER
ssl.keystore.location=/path/to/server.ks.p12
ssl.keystore.password=server-ks-password
ssl.key.password=server-ks-password
ssl.keystore.type=PKCS12
ssl.truststore.location=/path/to/server.ts.p12
ssl.truststore.password=server-ts-password
ssl.truststore.type=PKCS12
ssl.client.auth=required
```
```properties
# CLIENT
ssl.truststore.location=/path/to/client.ts.p12
ssl.truststore.password=client-ts-password
ssl.truststore.type=PKCS12
ssl.keystore.location=/path/to/client.ks.p12 # only if client auth
ssl.keystore.password=client-ks-password
ssl.key.password=client-ks-password
ssl.keystore.type=PKCS12
```
> ### TRUST STORES — you may not need them [#trust-stores--you-may-not-need-them]
>
> *"Trust store configuration **can be omitted in brokers as well as clients when using certificates signed by WELL-KNOWN TRUSTED AUTHORITIES. The default trust stores in the Java installation will be sufficient** to establish trust in this case."*
##### 💡 Rotating certificates without a restart [#-rotating-certificates-without-a-restart]
> *"**Key stores and trust stores must be updated periodically BEFORE CERTIFICATES EXPIRE to avoid TLS handshake failures.** Broker SSL stores **can be DYNAMICALLY UPDATED** by modifying the same file or setting the configuration option to a new versioned file. In both cases, **the Admin API or the Kafka configs tool can be used to trigger the update.**"*
```bash
bin/kafka-configs.sh --bootstrap-server localhost:9092 \
--command-config admin.props \
--entity-type brokers --entity-name 0 --alter --add-config \
'listener.name.external.ssl.keystore.location=/path/to/server.ks.p12'
```
##### TLS security considerations [#tls-security-considerations]
*(The DoS point connects to Ch. 6 §5.1: TLS handshakes burn **network thread** time — the same finite pool that moves every request onto the request queue.)*
***
#### 3.2 SASL — the framework and four mechanisms [#32-sasl--the-framework-and-four-mechanisms]
> *"SASL authentication is performed through **a sequence of server challenges and client responses** where the SASL mechanism defines **the sequence and wire format** of challenges and responses."*
| Mechanism | What it is | Production readiness |
| --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------ |
| **GSSAPI** | *"**Kerberos** authentication... can be used to integrate with Kerberos servers like **Active Directory or OpenLDAP**"* | Production-ready |
| **PLAIN** | *"Username/password authentication that is **typically used with a CUSTOM SERVER-SIDE CALLBACK** to verify passwords from an external password store"* | ⚠️ Built-in store is insecure |
| **SCRAM-SHA-256 / SCRAM-SHA-512** | *"Username/password authentication **available out of the box WITHOUT the need for additional password stores**"* | ✅ *if ZooKeeper is secure* |
| **OAUTHBEARER** | *"Authentication using **OAuth bearer tokens**, typically used with custom callbacks to acquire and validate tokens"* | ⚠️ Built-in impl is **not for production** |
**Enabling and selecting:**
```properties
# broker, per listener
sasl.enabled.mechanisms=
# client
sasl.mechanism=
```
**JAAS configuration:**
> *"Kafka uses the **Java Authentication and Authorization Service (JAAS)** for configuring SASL. The configuration option **`sasl.jaas.config`** contains a single JAAS configuration entry that specifies **a login module and its options.** Brokers use **the listener AND mechanism prefixes**"* — e.g. `listener.name.external.gssapi.sasl.jaas.config`.
> ### JAAS CONFIGURATION FILE vs `sasl.jaas.config` [#jaas-configuration-file-vs-sasljaasconfig]
>
> *"JAAS configuration may also be specified in configuration files using the Java system property `java.security.auth.login.config`. **However, the Kafka option `sasl.jaas.config` is RECOMMENDED since it supports PASSWORD PROTECTION and SEPARATE CONFIGURATION FOR EACH SASL MECHANISM when multiple mechanisms are enabled on a listener.**"*
**The three customization hooks — memorize what each is for:**
***
#### 3.3 SASL/GSSAPI (Kerberos) [#33-saslgssapi-kerberos]
**Broker configuration:**
```properties
sasl.enabled.mechanisms=GSSAPI
listener.name.external.gssapi.sasl.jaas.config=\
com.sun.security.auth.module.Krb5LoginModule required \
useKeyTab=true storeKey=true \
keyTab="/path/to/broker1.keytab" \
principal="kafka/broker1.example.com@EXAMPLE.COM";
```
* *"Keytab files **must be readable by the broker process**."*
* ⚠️ *"**Service principal for brokers SHOULD INCLUDE THE BROKER HOSTNAME**"* — *"Broker hostnames are verified by clients to ensure server authenticity and prevent man-in-the-middle attacks."*
**Inter-broker:**
```properties
sasl.mechanism.inter.broker.protocol=GSSAPI
sasl.kerberos.service.name=kafka
```
**Client:**
```properties
sasl.mechanism=GSSAPI
sasl.kerberos.service.name=kafka # the name of the service you connect TO
sasl.jaas.config=com.sun.security.auth.module.Krb5LoginModule required \
useKeyTab=true storeKey=true \
keyTab="/path/to/alice.keytab" \
principal="Alice@EXAMPLE.COM"; # clients MAY omit the hostname
```
**Principal derivation:** *"The **SHORT NAME** of the principal is used as the client identity by default. For example, **`User:Alice`** is the client principal and **`User:kafka`** is the broker principal."* Customize with **`sasl.kerberos.principal.to.local.rules`**.
> ⚠️ **DNS requirement:** *"Kerberos requires a **SECURE DNS SERVICE** for hostname lookup during authentication. **In deployments where forward and reverse lookup DO NOT MATCH**, the Kerberos configuration file `krb5.conf` on clients can be configured to set **`rdns=false`** to disable reverse lookup."*
##### Kerberos security considerations [#kerberos-security-considerations]
That last one is subtle and worth flagging: **NTP becomes part of your security perimeter.** An attacker who can skew clocks can defeat replay detection.
***
#### 3.4 SASL/PLAIN [#34-saslplain]
**The default (insecure) implementation uses the broker's JAAS config as the password store:**
```properties
sasl.enabled.mechanisms=PLAIN
sasl.mechanism.inter.broker.protocol=PLAIN
listener.name.external.plain.sasl.jaas.config=\
org.apache.kafka.common.security.plain.PlainLoginModule required \
username="kafka" password="kafka-password" \ # ① broker's own creds
user_kafka="kafka-password" \ # ② the password STORE
user_Alice="Alice-password";
```
```properties
# CLIENT
sasl.mechanism=PLAIN
sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule \
required username="Alice" password="Alice-password";
```
> ⚠️ *"The built-in implementation that **stores ALL passwords in EVERY broker's JAAS configuration is INSECURE AND NOT VERY FLEXIBLE since ALL BROKERS WILL NEED TO BE RESTARTED TO ADD OR REMOVE A USER.**"*
##### The production pattern: a custom server callback handler [#the-production-pattern-a-custom-server-callback-handler]
**Two jobs:** integrate with a secure third-party password server, **and support password rotation.**
> 💡 *"On the server side, a server callback handler **should support BOTH OLD AND NEW PASSWORDS for an OVERLAPPING PERIOD until all clients switch to the new password.**"*
```java
public class PasswordVerifier extends PlainServerCallbackHandler {
private final List passwdFiles = new ArrayList<>(); // ① multiple
// files →
@Override // rotation
public void configure(Map configs, String mechanism,
List jaasEntries) {
Map loginOptions = jaasEntries.get(0).getOptions();
String files = (String) loginOptions.get("password.files"); // ② JAAS option
Collections.addAll(passwdFiles, files.split(","));
}
@Override
protected boolean authenticate(String user, char[] password) {
return passwdFiles.stream() // ③ match ANY
.anyMatch(file -> authenticate(file, user, password)); // file
}
private boolean authenticate(String file, String user, char[] password) {
try {
String cmd = String.format("htpasswd -vb %s %s %s", // ④ htpasswd
file, user, new String(password)); // for
return Runtime.getRuntime().exec(cmd).waitFor() == 0; // simplicity
} catch (Exception e) {
return false;
}
}
}
```
④ *"We use `htpasswd` for simplicity. **A secure database can be used for production deployments.**"*
```properties
listener.name.external.plain.sasl.jaas.config=\
org.apache.kafka.common.security.plain.PlainLoginModule required \
password.files="/path/to/htpassword.props,/path/to/oldhtpassword.props";
listener.name.external.plain.sasl.server.callback.handler.class=\
com.example.PasswordVerifier
```
##### Client-side callback: load passwords at connection time, not startup [#client-side-callback-load-passwords-at-connection-time-not-startup]
> *"a client callback handler that implements `org.apache.kafka.common.security.auth.AuthenticateCallbackHandler` can be used to **load passwords DYNAMICALLY AT RUNTIME when a connection is established** instead of loading statically from the JAAS configuration during startup."*
```java
@Override
public void handle(Callback[] callbacks) throws IOException {
Properties props = Utils.loadProps(passwdFile); // ① reload EVERY time
PasswordConfig config = new PasswordConfig(props); // → supports rotation
String user = config.getString("username");
String password = config.getPassword("password").value();// ② returns the real
for (Callback callback: callbacks) { // value even if
if (callback instanceof NameCallback) // EXTERNALIZED
((NameCallback) callback).setName(user);
else if (callback instanceof PasswordCallback) {
((PasswordCallback) callback).setPassword(password.toCharArray());
}
}
}
private static class PasswordConfig extends AbstractConfig {
static ConfigDef CONFIG = new ConfigDef()
.define("username", STRING, HIGH, "User name")
.define("password", PASSWORD, HIGH, "User password"); // ③ PASSWORD type
PasswordConfig(Properties props) { super(CONFIG, props, false); }
}
```
③ 💡 *"We define password configs with the **`PASSWORD` type to ensure that passwords are NOT INCLUDED IN LOG ENTRIES.**"* — a cheap, high-value trick.
```properties
sasl.jaas.config=org.apache.kafka.common.security.plain.PlainLoginModule \
required file="/path/to/credentials.props";
sasl.client.callback.handler.class=com.example.PasswordProvider
```
##### PLAIN security considerations [#plain-security-considerations]
***
#### 3.5 SASL/SCRAM [#35-saslscram]
> *"**RFC-5802** introduces a secure username/password authentication mechanism that **addresses the security concerns with password authentication mechanisms like SASL/PLAIN, which send passwords over the wire.** The **Salted Challenge Response Authentication Mechanism (SCRAM)** **avoids transmitting clear-text passwords** and **stores passwords in a format that makes it IMPRACTICAL TO IMPERSONATE CLIENTS. Salting combines passwords with some random data before applying a one-way cryptographic hash function.**"*
**Creating users — note this uses `--zookeeper`, before brokers start:**
```bash
bin/kafka-configs.sh --zookeeper localhost:2181 --alter --add-config \
'SCRAM-SHA-512=[iterations=8192,password=Alice-password]' \
--entity-type users --entity-name Alice
```
> *"An initial set of users can be created **after starting ZooKeeper PRIOR TO STARTING BROKERS. Brokers load SCRAM user metadata into an IN-MEMORY CACHE during startup**, ensuring that all users, **including the broker user for inter-broker communication**, can authenticate successfully. **Users can be added or deleted AT ANY TIME. Brokers keep the cache up-to-date using notifications based on a ZOOKEEPER WATCHER.**"*
```properties
# BROKER
sasl.enabled.mechanisms=SCRAM-SHA-512
sasl.mechanism.inter.broker.protocol=SCRAM-SHA-512
listener.name.external.scram-sha-512.sasl.jaas.config=\
org.apache.kafka.common.security.scram.ScramLoginModule required \
username="kafka" password="kafka-password";
```
```properties
# CLIENT
sasl.mechanism=SCRAM-SHA-512
sasl.jaas.config=org.apache.kafka.common.security.scram.ScramLoginModule \
required username="Alice" password="Alice-password";
```
**Deleting a user:**
```bash
bin/kafka-configs.sh --zookeeper localhost:2181 --alter --delete-config \
'SCRAM-SHA-512' --entity-type users --entity-name Alice
```
> ⚠️ *"When an existing user is deleted, **new connections cannot be established for that user, BUT EXISTING CONNECTIONS OF THE USER WILL CONTINUE TO WORK. A REAUTHENTICATION INTERVAL can be configured for the broker to LIMIT THE AMOUNT OF TIME existing connections may continue to operate after a user is deleted.**"*
##### SCRAM security considerations [#scram-security-considerations]
**The dependency to remember: built-in SCRAM's security is *bounded by ZooKeeper's* security.** That's the price of not running a password server.
***
#### 3.6 SASL/OAUTHBEARER [#36-sasloauthbearer]
> *"**RFC-7628** defines the OAUTHBEARER SASL mechanism that enables credentials obtained using **OAuth 2.0** to access protected resources in **non-HTTP protocols.** OAUTHBEARER **avoids security vulnerabilities in mechanisms that use LONG-TERM PASSWORDS by using OAuth 2.0 bearer tokens with a SHORTER LIFETIME and LIMITED RESOURCE ACCESS.**"*
> ### ⚠️ *"The built-in implementation of OAUTHBEARER uses **unsecured JSON Web Tokens (JWTs)** and **IS NOT SUITABLE FOR PRODUCTION USE.** Custom callbacks can be added to integrate with standard OAuth servers."* [#️-the-built-in-implementation-of-oauthbearer-uses-unsecured-json-web-tokens-jwts-and-is-not-suitable-for-production-use-custom-callbacks-can-be-added-to-integrate-with-standard-oauth-servers]
**Built-in (dev only) — note it does not validate tokens:**
```properties
# BROKER
sasl.enabled.mechanisms=OAUTHBEARER
sasl.mechanism.inter.broker.protocol=OAUTHBEARER
listener.name.external.oauthbearer.sasl.jaas.config=\
org.apache.kafka.common.security.oauthbearer.OAuthBearerLoginModule \
required unsecuredLoginStringClaim_sub="kafka";
```
```properties
# CLIENT — User:Alice is the resulting default KafkaPrincipal
sasl.mechanism=OAUTHBEARER
sasl.jaas.config=\
org.apache.kafka.common.security.oauthbearer.OAuthBearerLoginModule \
required unsecuredLoginStringClaim_sub="Alice";
```
> *"The option **`unsecuredLoginStringClaim_sub` is the SUBJECT CLAIM that determines the `KafkaPrincipal`** for the connection by default."*
**Production — two callbacks required:**
**① Client `sasl.login.callback.handler.class`** — *"to acquire tokens from the OAuth server using the long-term password or a refresh token"*:
```java
@Override
public void handle(Callback[] callbacks) throws UnsupportedCallbackException {
OAuthBearerToken token = null;
for (Callback callback : callbacks) {
if (callback instanceof OAuthBearerTokenCallback) {
token = acquireToken(); // ① from OAuth server
((OAuthBearerTokenCallback) callback).token(token);
} else if (callback instanceof SaslExtensionsCallback) {// ② optional
((SaslExtensionsCallback) callback).extensions(processExtensions(token));
} else
throw new UnsupportedCallbackException(callback);
}
}
```
> ⚠️ *"**If OAUTHBEARER is used for inter-broker communication, BROKERS must ALSO be configured with a LOGIN callback handler** to acquire tokens for client connections created by the broker."*
**② Broker `listener.name..oauthbearer.sasl.server.callback.handler.class`** — for **validating** tokens:
```java
@Override
public void handle(Callback[] callbacks) throws UnsupportedCallbackException {
for (Callback callback : callbacks) {
if (callback instanceof OAuthBearerValidatorCallback) {
OAuthBearerValidatorCallback cb = (OAuthBearerValidatorCallback) callback;
try {
cb.token(validatedToken(cb.tokenValue())); // ① validate
} catch (OAuthBearerIllegalTokenException e) {
OAuthBearerValidationResult r = e.reason();
cb.error(errorStatus(r), r.failureScope(), r.failureOpenIdConfig());
}
} else if (callback instanceof OAuthBearerExtensionsValidatorCallback) {
OAuthBearerExtensionsValidatorCallback ecb =
(OAuthBearerExtensionsValidatorCallback) callback;
ecb.inputExtensions().map().forEach((k, v) ->
ecb.valid(validateExtension(k, v))); // ② validate extensions
} else { throw new UnsupportedCallbackException(callback); }
}
}
```
##### OAUTHBEARER security considerations [#oauthbearer-security-considerations]
***
#### 3.7 Delegation tokens [#37-delegation-tokens]
> *"Delegation tokens are **shared secrets between Kafka brokers and clients** that provide a lightweight configuration mechanism **WITHOUT THE REQUIREMENT TO DISTRIBUTE SSL KEY STORES OR KERBEROS KEYTABS to client applications.** Delegation tokens can be used to **REDUCE THE LOAD ON AUTHENTICATION SERVERS**, like the Kerberos Key Distribution Center (KDC)."*
```bash
# create — "If Alice runs this command, the generated token can be used to
# IMPERSONATE ALICE. The owner of this token is User:Alice."
bin/kafka-delegation-tokens.sh --bootstrap-server localhost:9092 \
--command-config admin.props --create --max-life-time-period -1 \
--renewer-principal User:Bob
# renew — "can be run by the token OWNER (Alice) or the token RENEWER (Bob)"
bin/kafka-delegation-tokens.sh --bootstrap-server localhost:9092 \
--command-config admin.props --renew --renew-time-period -1 --hmac c2VjcmV0
```
> ⚠️ *"To create delegation tokens for the principal `User:Alice`, the client must be authenticated using Alice's credentials **for any authentication protocol OTHER THAN delegation tokens. CLIENTS AUTHENTICATED USING DELEGATION TOKENS CANNOT CREATE OTHER DELEGATION TOKENS.**"* (No privilege chaining.)
**Configuration:**
```properties
# ALL brokers must share the same master key
delegation.token.master.key=
```
> ⚠️ *"This key **can only be rotated by RESTARTING ALL BROKERS. ALL EXISTING TOKENS SHOULD BE DELETED BEFORE UPDATING THE MASTER KEY** since they can no longer be used, and **new tokens should be created AFTER the key is updated on all brokers.**"*
> *"**At least one of the SASL/SCRAM mechanisms must be enabled** on brokers to support authentication using delegation tokens."*
```properties
# CLIENT — note tokenauth="true"
sasl.mechanism=SCRAM-SHA-512
sasl.jaas.config=org.apache.kafka.common.security.scram.ScramLoginModule \
required tokenauth="true" username="MTIz" password="c2VjcmV0";
```
> *"The `KafkaPrincipal` for connections using this configuration **will be THE ORIGINAL PRINCIPAL associated with the token, e.g., `User:Alice`.**"*
**Security considerations:** *"suitable for production use **only in deployments where ZooKeeper is secure.** All the security considerations described under SCRAM also apply. **The MASTER KEY used by brokers for generating tokens must be protected using encryption or by externalizing the key in a secure password store. SHORT-LIVED delegation tokens** can be used to limit exposure. **Reauthentication** can be enabled to prevent connections operating with expired tokens."*
***
#### 3.8 ⚠️ Reauthentication — and the compromised-user problem [#38-️-reauthentication--and-the-compromised-user-problem]
**The default behavior is the problem:**
> *"Kafka brokers perform client authentication **when a connection is established.** ... Kafka uses a **background login thread** to acquire new credentials before the old ones expire, **but the new credentials are used ONLY TO AUTHENTICATE NEW CONNECTIONS by default. EXISTING CONNECTIONS THAT WERE AUTHENTICATED WITH OLD CREDENTIALS CONTINUE TO PROCESS REQUESTS until disconnection occurs due to a request timeout, an idle timeout, or network errors. LONG-LIVED CONNECTIONS MAY CONTINUE TO PROCESS REQUESTS LONG AFTER THE CREDENTIALS USED TO AUTHENTICATE THE CONNECTIONS EXPIRE.**"*
**The fix:**
```properties
connections.max.reauth.ms=
```
> \*"When set to a positive integer, Kafka brokers **determine the SESSION LIFETIME for SASL connections and inform clients of this lifetime DURING THE SASL HANDSHAKE.**
>
> **Session lifetime = the LOWER of (remaining lifetime of the credential) and (`connections.max.reauth.ms`).**
>
> **Any connection that doesn't reauthenticate within this interval IS TERMINATED BY THE BROKER.**"\*
**Four scenarios it improves:**
> ### ⚠️ COMPROMISED USERS — the incident-response procedure [#️-compromised-users--the-incident-response-procedure]
>
> \*"If a user is compromised, **action must be taken to remove the user from the system AS SOON AS POSSIBLE.** All new connections will fail to authenticate once the user is removed from the authentication server. **EXISTING connections will continue to process requests until the next reauthentication timeout. If `connections.max.reauth.ms` IS NOT CONFIGURED, NO TIMEOUT IS APPLIED AND EXISTING CONNECTIONS MAY CONTINUE TO USE THE COMPROMISED USER'S IDENTITY FOR A LONG TIME.**
>
> **Kafka does not support SSL renegotiation** due to known vulnerabilities... **Newer protocols like TLSv1.3 do not support renegotiation. SO, EXISTING SSL CONNECTIONS MAY CONTINUE TO USE REVOKED OR EXPIRED CERTIFICATES.**
>
> 💡 **`Deny` ACLs for the user principal can be used to PREVENT THESE CONNECTIONS FROM PERFORMING ANY OPERATION. SINCE ACL CHANGES ARE APPLIED WITH VERY SMALL LATENCIES ACROSS ALL BROKERS, THIS IS THE QUICKEST WAY TO DISABLE ACCESS FOR COMPROMISED USERS.**"\*
***
# 11.6 Authorization (/docs/kafka/securing-kafka/authorization)
> *"Kafka brokers manage access control using a **customizable authorizer**... When a request is processed, the broker verifies that **the principal associated with the connection is authorized to perform that request.**"*
```properties
authorizer.class.name=kafka.security.authorizer.AclAuthorizer
```
> ### `SimpleAclAuthorizer` [#simpleaclauthorizer]
>
> *"`AclAuthorizer` was introduced in Apache Kafka **2.3.** Older versions from **0.9.0.0** onward had `kafka.security.auth.SimpleAclAuthorizer`, **which has been DEPRECATED but is still supported.**"*
#### 6.1 AclAuthorizer [#61-aclauthorizer]
> *"**ACLs are stored in ZooKeeper and CACHED IN MEMORY by EVERY broker** to enable high-performance lookup. ACLs are loaded into the cache when the broker starts up, and **the cache is kept up-to-date using notifications based on a ZOOKEEPER WATCHER.**"*
*(This is why Deny ACLs are the fastest incident response — the ZK watcher propagates them near-instantly to every broker.)*
#### The six components of an ACL binding [#the-six-components-of-an-acl-binding]
#### The evaluation rule and the implicit grants [#the-evaluation-rule-and-the-implicit-grants]
#### Who needs what — the practical summary [#who-needs-what--the-practical-summary]
#### The full ACL → request mapping (Table 11-1) [#the-full-acl--request-mapping-table-11-1]
| ACL | Kafka requests | Notes |
| ----------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **`Cluster:ClusterAction`** | Inter-broker requests, including controller requests and follower fetch requests for replication | ⚠️ **"Should only be granted to brokers."** |
| `Cluster:Create` | `CreateTopics` and auto-topic creation | *"Use `Topic:Create` for **fine-grained** access control"* |
| `Cluster:Alter` | `CreateAcls`, `DeleteAcls`, `AlterReplicaLogDirs`, `ElectReplicaLeader`, `AlterPartitionReassignments` | |
| `Cluster:AlterConfigs` | `AlterConfigs`/`IncrementalAlterConfigs` for **broker and broker logger**, `AlterClientQuotas` | |
| `Cluster:Describe` | `DescribeAcls`, `DescribeLogDirs`, `ListGroups`, `ListPartitionReassignments`, describing authorized operations for cluster in `Metadata` | *"Use `Group:Describe` for fine-grained control for `ListGroups`"* |
| `Cluster:DescribeConfigs` | `DescribeConfigs` for broker and broker logger, `DescribeClientQuotas` | |
| **`Cluster:IdempotentWrite`** | Idempotent `InitProducerId` and `Produce` | ⚠️ **"Only required for NONTRANSACTIONAL idempotent producers."** |
| `Topic:Create` | `CreateTopics` and auto-topic creation | |
| `Topic:Delete` | `DeleteTopics`, **`DeleteRecords`** | |
| `Topic:Alter` | **`CreatePartitions`** | |
| `Topic:AlterConfigs` | `AlterConfigs`/`IncrementalAlterConfigs` for topics | |
| `Topic:Describe` | `Metadata` for topic, `OffsetForLeaderEpoch`, `ListOffset`, `OffsetFetch` | |
| `Topic:DescribeConfigs` | `DescribeConfigs` for topics; returning configs in `CreateTopics` response | |
| **`Topic:Read`** | Consumer `Fetch`, `OffsetCommit`, `TxnOffsetCommit`, `OffsetDelete` | *"Should be granted to consumers."* |
| **`Topic:Write`** | `Produce`, `AddPartitionToTxn` | *"Should be granted to producers."* |
| **`Group:Read`** | `JoinGroup`, `SyncGroup`, `LeaveGroup`, `Heartbeat`, `OffsetCommit`, `AddOffsetsToTxn`, `TxnOffsetCommit` | *"Required for consumers using group management or Kafka-based offset management. **Also required for TRANSACTIONAL PRODUCERS to commit offsets within a transaction.**"* |
| `Group:Describe` | `FindCoordinator`, `DescribeGroup`, `ListGroups`, `OffsetFetch` | |
| `Group:Delete` | `DeleteGroups`, `OffsetDelete` | |
| **`TransactionalId:Write`** | `Produce` and `InitProducerId` with transactions, `AddPartitionToTxn`, `AddOffsetsToTxn`, `TxnOffsetCommit`, `EndTxn` | *"Required for transactional producers."* |
| `TransactionalId:Describe` | `FindCoordinator` for transaction coordinator | |
| `DelegationToken:Describe` | `DescribeTokens` | |
#### Managing ACLs [#managing-acls]
```bash
# ① Broker ACLs created DIRECTLY IN ZOOKEEPER — "useful to create broker ACLs
# PRIOR TO STARTING BROKERS"
bin/kafka-acls.sh --add --cluster --operation ClusterAction \
--authorizer-properties zookeeper.connect=localhost:2181 \
--allow-principal User:kafka
# ② producer convenience flag; LITERAL by default
bin/kafka-acls.sh --bootstrap-server localhost:9092 \
--command-config admin.props --add --topic customerOrders \
--producer --allow-principal User:Alice
# ③ PREFIXED ACL: Bob may read all topics starting with "customer"
bin/kafka-acls.sh --bootstrap-server localhost:9092 \
--command-config admin.props --add --resource-pattern-type PREFIXED \
--topic customer --operation Read --allow-principal User:Bob
```
#### ⚠️ The two dangerous shortcuts [#️-the-two-dangerous-shortcuts]
```properties
super.users=User:Carol;User:Admin # NOTE: SEMICOLON-separated!
allow.everyone.if.no.acl.found=true
```
> ### SUPER USER SEPARATOR [#super-user-separator]
>
> *"Unlike other list configurations in Kafka that are **comma-separated**, `super.users` are separated by a **SEMICOLON** since user principals such as **distinguished names from SSL certificates OFTEN CONTAIN COMMAS.**"*
**`super.users`:**
> *"Super users are granted access for **all operations on all resources without any restrictions** and **CANNOT BE DENIED ACCESS USING `Deny` ACLs. If Carol's credentials are compromised, Carol must be REMOVED FROM `super.users`, AND BROKERS MUST BE RESTARTED** to apply the changes. **It is SAFER to grant specific access using ACLs to users in production systems to ensure access can be revoked easily.**"*
**`allow.everyone.if.no.acl.found`:**
> *"all users are granted access to resources **without any ACLs.** May be useful **when enabling authorization for the first time or during development, BUT IS NOT SUITABLE FOR PRODUCTION** since **access may be granted UNINTENTIONALLY to NEW resources.** Access may also be **UNEXPECTEDLY REMOVED when ACLs for a matching PREFIX or WILDCARD are added, if the condition for `no.acl.found` NO LONGER APPLIES.**"*
That second failure mode is genuinely surprising: **adding an ACL can revoke access** for everyone else who was relying on the "no ACL found" fallback.
#### 6.2 Customizing authorization [#62-customizing-authorization]
**Example 1 — restrict certain requests to the internal listener:**
```java
public class CustomAuthorizer extends AclAuthorizer {
private static final Set internalOps =
Utils.mkSet(CREATE_ACLS.id, DELETE_ACLS.id);
private static final String internalListener = "INTERNAL";
@Override
public List authorize(
AuthorizableRequestContext context, List actions) {
if (!context.listenerName().equals(internalListener) && // ①
internalOps.contains((short) context.requestType()))
return Collections.nCopies(actions.size(), DENIED);
else
return super.authorize(context, actions); // ②
}
}
```
① *"Authorizers are given **the request context with metadata that includes LISTENER NAMES, SECURITY PROTOCOL, REQUEST TYPES**, etc., enabling custom authorizers to add or remove restrictions **based on the context.**"*
② *"We **REUSE functionality from the built-in Kafka authorizer using the public API.**"*
**Example 2 — group/role-based access control (RBAC) from LDAP:**
```scala
class RbacAuthorizer extends AclAuthorizer {
@volatile private var groups = Map.empty[KafkaPrincipal, Set[KafkaPrincipal]]
.withDefaultValue(Set.empty) // ① from LDAP
@volatile private var roles = Map.empty[KafkaPrincipal, Set[KafkaPrincipal]]
.withDefaultValue(Set.empty) // ② from LDAP
override def authorize(context: AuthorizableRequestContext,
actions: util.List[Action]): util.List[AuthorizationResult] = {
val principals = groups(context.principal) + context.principal // ③
val allPrincipals = principals.flatMap(roles) ++ principals
val contexts = allPrincipals.map(authorizeContext(context, _)) // ⑤
actions.asScala.map { action =>
val authorized = contexts.exists( // ④
super.authorize(_, List(action).asJava).get(0) == ALLOWED)
if (authorized) ALLOWED else DENIED
}.asJava
}
// authorizeContext(...) wraps the original context, swapping the principal
}
```
③ *"We perform authorization for **the user as well as for ALL the groups and roles of the user.**"*
④ *"**If ANY of the contexts are authorized, we return ALLOWED. Note that this example DOESN'T SUPPORT `Deny` ACLs for groups or roles.**"*
```bash
# ACLs for a GROUP principal
bin/kafka-acls.sh ... --add --topic customer --producer \
--resource-pattern-type PREFIXED --allow-principal Group:Sales
# ACLs for a ROLE principal
bin/kafka-acls.sh ... --add --cluster --operation Alter \
--allow-principal=Role:Operator
```
#### 6.3 Authorization security considerations [#63-authorization-security-considerations]
***
# 11.5 Encryption (/docs/kafka/securing-kafka/encryption)
#### Three layers, three different threats [#three-layers-three-different-threats]
**Heap dumps are the underrated threat.** TLS + disk encryption still leaves plaintext in broker RAM.
#### 5.1 End-to-end encryption [#51-end-to-end-encryption]
**Where it hooks in:** *"**Serializers and deserializers can be integrated with an encryption library** to perform encryption of the message **during serialization**, and decryption **during deserialization.**"* (Ch. 3/Ch. 4.)
**Details:**
##### Key rotation — and the compacted-topic problem [#key-rotation--and-the-compacted-topic-problem]
> *"**Periodic key rotation is recommended**... since frequent rotation **limits the number of compromised messages in case of a breach and also protects against brute-force attacks. CONSUMPTION MUST BE SUPPORTED WITH BOTH OLD AND NEW KEYS DURING THE RETENTION PERIOD of messages encrypted with the old key.** Many KMS systems support graceful key rotation out of the box for symmetric encryption without requiring special handling in Kafka clients."*
> ⚠️ *"**For COMPACTED TOPICS, messages encrypted with old keys MAY BE RETAINED FOR A LONG TIME, and it may be necessary to RE-ENCRYPT OLD MESSAGES. To avoid interference with newer messages, PRODUCERS AND CONSUMERS MUST BE OFFLINE DURING THIS PROCESS.**"*
That's a real operational cost of combining end-to-end encryption with compacted topics: **key rotation becomes a maintenance window.**
> ### ⚠️ COMPRESSION OF ENCRYPTED MESSAGES [#️-compression-of-encrypted-messages]
>
> \*"**Compressing messages AFTER encryption is UNLIKELY TO PROVIDE ANY BENEFIT** in terms of space reduction compared to compressing prior to encryption. Serializers may be configured to **perform compression BEFORE encrypting**, or applications may compress prior to producing. **In either case, IT IS BETTER TO DISABLE COMPRESSION IN KAFKA since it adds overhead without providing any additional benefit.**
>
> For messages transmitted over an insecure transport layer, **KNOWN SECURITY EXPLOITS OF COMPRESSED ENCRYPTED MESSAGES must also be taken into account.**"\*
*(Encrypted data is high-entropy → incompressible. And compression-plus-encryption oracles like CRIME/BREACH are the "known exploits" being referenced.)*
##### ⚠️ Encrypting message *keys* — the hash-equivalence problem [#️-encrypting-message-keys--the-hash-equivalence-problem]
> \*"In many environments, especially when TLS is used, **message keys do not require encryption** since they typically do not contain sensitive data. But in some cases, clear-text keys may not comply with regulatory requirements.
>
> **Since message keys are used for PARTITIONING and COMPACTION, transformation of keys MUST PRESERVE THE REQUIRED HASH EQUIVALENCE to ensure that a key RETAINS THE SAME HASH VALUE even if encryption parameters are altered.**
>
> **One approach:** store **a SECURE HASH of the original key AS the message key**, and store **the ENCRYPTED message key in the message PAYLOAD or in a HEADER.**
>
> Since **Kafka serializes message key and value INDEPENDENTLY, a PRODUCER INTERCEPTOR can be used to perform this transformation.**"\*
***
# 11.1 The five security procedures, and the reference data flow (/docs/kafka/securing-kafka/five-security-procedures-reference)
#### The reference flow used throughout the chapter [#the-reference-flow-used-throughout-the-chapter]
#### The seven guarantees a secure deployment must provide [#the-seven-guarantees-a-secure-deployment-must-provide]
| Guarantee | What it means in the flow |
| ----------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Client authenticity** | *"the broker should authenticate the client to ensure that the message is really coming from Alice"* |
| **Server authenticity** | *"Alice's client should verify that the connection is to the REAL broker"* |
| **Data privacy** | *"**ALL connections where the message flows, as well as ALL DISKS where messages are stored**, should be encrypted or physically secured"* |
| **Data integrity** | *"**Message digests** should be included for data transmitted over insecure networks to detect tampering"* |
| **Access control** | Leader verifies Alice may **write** to `customerOrders`; verifies Bob may **read**; *"If Bob's consumer uses group management, the broker should ALSO verify that Bob has access to the CONSUMER GROUP"* |
| **Auditability** | *"An audit trail that shows all operations performed by **brokers, Alice, Bob, and other clients**"* |
| **Availability** | *"quotas and limits to avoid some users hogging all the available bandwidth or overwhelming the broker with DoS attacks. **ZooKeeper should be locked down** since broker availability is dependent on ZooKeeper availability and the integrity of metadata stored in ZooKeeper"* |
Note that **availability is a security concern here**, not just an ops concern — and that it explicitly includes ZooKeeper.
***
# 11. Securing Kafka (/docs/kafka/securing-kafka)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 11*
> **The framing, which mirrors Ch. 7's on reliability:** *"Like performance and reliability, **security is an aspect of the system that must be addressed for the system AS A WHOLE, rather than component by component. THE SECURITY OF A SYSTEM IS ONLY AS STRONG AS THE WEAKEST LINK**, and security processes and policies must be enforced across the system, **including the underlying platform.**"*
**And the honest trade-off statement:** *"While it is always preferable to use the strongest and latest security features available, **trade-offs are often necessary since INCREASED SECURITY IMPACTS PERFORMANCE, COST, AND USER EXPERIENCE.**"*
***
# 11.9 Securing the platform (/docs/kafka/securing-kafka/securing-platform)
> \*"Security design for a production system should use a **THREAT MODEL** that addresses security threats **not just for individual components but also for the system as a whole.** Threat models **build an abstraction of the system and identify potential threats and the associated risks. Once the threats are evaluated, DOCUMENTED, and PRIORITIZED BASED ON RISKS, mitigation strategies must be implemented for EACH potential threat.**
>
> When assessing potential threats, **it is important to consider EXTERNAL threats AS WELL AS INSIDER THREATS.**"\*
**The platform-level defenses:**
#### 9.1 Password protection — two mechanisms [#91-password-protection--two-mechanisms]
##### A. Custom `ConfigProvider` (works for brokers *and* clients) [#a-custom-configprovider-works-for-brokers-and-clients]
```java
public class GpgProvider implements ConfigProvider {
@Override public void configure(Map configs) {}
@Override
public ConfigData get(String path) {
try {
String passphrase = System.getenv("PASSPHRASE"); // ① from the ENV
String data = Shell.execCommand( // ② gpg decrypt
"gpg", "--decrypt", "--passphrase", passphrase, path);
Properties props = new Properties();
props.load(new StringReader(data)); // ③ parse
Map map = new HashMap<>();
for (String name : props.stringPropertyNames())
map.put(name, props.getProperty(name));
return new ConfigData(map);
} catch (IOException e) {
throw new RuntimeException(e); // ④ FAIL FAST
}
}
@Override
public ConfigData get(String path, Set keys) { // ⑤ subset
ConfigData configData = get(path);
Map data = configData.data().entrySet()
.stream().filter(e -> keys.contains(e.getKey()))
.collect(Collectors.toMap(Map.Entry::getKey, Map.Entry::getValue));
return new ConfigData(data, configData.ttl());
}
@Override public void close() {}
}
```
**Encrypt the credentials file:**
```bash
gpg --symmetric --output credentials.props.gpg \
--passphrase "$PASSPHRASE" credentials.props
```
**Reference it indirectly — the `${provider:path:key}` syntax:**
```properties
username=${gpg:/path/to/credentials.props.gpg:username}
password=${gpg:/path/to/credentials.props.gpg:password}
config.providers=gpg
config.providers.gpg.class=com.example.GpgProvider
```
*(Same mechanism Ch. 9 §2.6 referenced for Connect secret providers — Vault/AWS/Azure providers exist in the community.)*
##### B. Broker-only: encrypted configs in ZooKeeper, no custom code [#b-broker-only-encrypted-configs-in-zookeeper-no-custom-code]
```bash
bin/kafka-configs.sh --zookeeper localhost:2181 --alter \
--entity-type brokers --entity-name 0 --add-config \
'listener.name.external.ssl.keystore.password=server-ks-password,\
password.encoder.secret=encoder-secret'
```
> *"The following command can be executed **BEFORE STARTING BROKERS** to store **encrypted SSL key store passwords for brokers in ZooKeeper. THE PASSWORD ENCODER SECRET MUST BE CONFIGURED IN EACH BROKER'S CONFIGURATION FILE TO DECRYPT THE VALUE.**"*
***
# 11.8 Securing ZooKeeper (/docs/kafka/securing-kafka/securing-zookeeper)
> *"ZooKeeper stores Kafka metadata that is **critical for maintaining the availability of Kafka clusters**, and hence **it is VITAL to secure ZooKeeper in addition to securing Kafka.**"*
#### 8.1 SASL [#81-sasl]
> *"SASL configuration for ZooKeeper is provided using the Java system property **`java.security.auth.login.config`**"* pointing at a JAAS file.
**ZooKeeper server JAAS (`Server` section):**
```txt
Server {
com.sun.security.auth.module.Krb5LoginModule required
useKeyTab=true storeKey=true
keyTab="/path/to/zk.keytab"
principal="zookeeper/zk1.example.com@EXAMPLE.COM";
};
```
**ZooKeeper server config:**
```properties
authProvider.sasl=org.apache.zookeeper.server.auth.SASLAuthenticationProvider
kerberos.removeHostFromPrincipal=true
kerberos.removeRealmFromPrincipal=true
```
> ### ⚠️ BROKER PRINCIPAL — why those two flags matter [#️-broker-principal--why-those-two-flags-matter]
>
> *"By default, ZooKeeper uses **the FULL Kerberos principal**, e.g. `kafka/broker1.example.com@EXAMPLE.COM`, as the client identity. **When ACLs are enabled for ZooKeeper authorization, ZooKeeper servers SHOULD be configured with `kerberos.removeHostFromPrincipal=true` and `kerberos.removeRealmFromPrincipal=true` TO ENSURE THAT ALL BROKERS HAVE THE SAME PRINCIPAL.**"*
**Kafka broker's ZooKeeper client JAAS (`Client` section):**
```txt
Client {
com.sun.security.auth.module.Krb5LoginModule required
useKeyTab=true storeKey=true
keyTab="/path/to/broker1.keytab"
principal="kafka/broker1.example.com@EXAMPLE.COM";
};
```
#### 8.2 SSL [#82-ssl]
> ⚠️ **The key difference from Kafka:** *"Like Kafka, SSL may be configured to enable client authentication, **BUT UNLIKE KAFKA, connections with BOTH SASL AND SSL client authentication AUTHENTICATE USING BOTH PROTOCOLS AND ASSOCIATE MULTIPLE PRINCIPALS WITH THE CONNECTION. ZooKeeper authorizer grants access to a resource IF ANY OF THE PRINCIPALS associated with the connection have access.**"*
**ZooKeeper server:**
```properties
secureClientPort=2181
serverCnxnFactory=org.apache.zookeeper.server.NettyServerCnxnFactory
authProvider.x509=org.apache.zookeeper.server.auth.X509AuthenticationProvider
ssl.keyStore.location=/path/to/zk.ks.p12
ssl.keyStore.password=zk-ks-password
ssl.keyStore.type=PKCS12
ssl.trustStore.location=/path/to/zk.ts.p12
ssl.trustStore.password=zk-ts-password
ssl.trustStore.type=PKCS12
```
**Kafka broker → ZooKeeper:**
```properties
zookeeper.ssl.client.enable=true
zookeeper.clientCnxnSocket=org.apache.zookeeper.ClientCnxnSocketNetty
zookeeper.ssl.keystore.location=/path/to/zkclient.ks.p12
zookeeper.ssl.keystore.password=zkclient-ks-password
zookeeper.ssl.keystore.type=PKCS12
zookeeper.ssl.truststore.location=/path/to/zkclient.ts.p12
zookeeper.ssl.truststore.password=zkclient-ts-password
zookeeper.ssl.truststore.type=PKCS12
```
#### 8.3 ZooKeeper authorization [#83-zookeeper-authorization]
```properties
zookeeper.set.acl=true # on the BROKERS
```
> *"the broker **sets ACLs for ZooKeeper nodes WHEN CREATING THE NODE. By default, metadata nodes are READABLE BY EVERYONE but MODIFIABLE ONLY BY BROKERS.** Additional ACLs may be added if required for internal admin users who may need to update metadata directly. **SENSITIVE PATHS, LIKE NODES CONTAINING SCRAM CREDENTIALS, ARE NOT WORLD-READABLE BY DEFAULT.**"*
***
# 11.11 The security decision guide (/docs/kafka/securing-kafka/security-decision-guide)
#### Production hardening checklist [#production-hardening-checklist]
#### How the chapter's guarantees map to mechanisms [#how-the-chapters-guarantees-map-to-mechanisms]
| Guarantee | Mechanisms |
| ----------------------- | ---------------------------------------------------------------------------------------------------------------------- |
| **Client authenticity** | SASL, or SSL with client auth; **reauthentication** to limit compromise exposure |
| **Server authenticity** | SSL with **hostname validation**, or mutual-auth SASL (**Kerberos, SCRAM**) |
| **Data privacy** | TLS in transit; disk/volume encryption at rest; **end-to-end** for fine-grained control against admins/cloud providers |
| **Data integrity** | TLS detects tampering; **digital signatures** in messages under end-to-end encryption |
| **Access control** | Customizable authorizer; built-in **`AclAuthorizer`** with fine-grained ACLs |
| **Auditability** | Authorizer logs + request logs; message-header audit metadata |
| **Availability** | **Quotas** + connection management against DoS; **ZooKeeper** secured with SSL, SASL, and ACLs |
***
# 11.2 Security protocols — the 2×2 matrix (/docs/kafka/securing-kafka/security-protocols-2-2)
> *"Kafka brokers are configured with **listeners on one or more endpoints**... **Each listener can be configured with its own security settings.** Security requirements on a **private internal listener** that is physically protected... may be different from the requirements of an **external listener accessible over the public internet.**"*
**Two standard technologies:**
**Each Kafka security protocol = a transport layer + an optional authentication layer:**
| Protocol | Description | Suitability |
| ------------------- | -------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------- |
| **PLAINTEXT** | *"PLAINTEXT transport with **no authentication**"* | *"only for use within **private networks** for processing data that is **not sensitive**"* |
| **SSL** | *"SSL transport with **optional SSL client authentication**"* | ✅ *"suitable for use in **insecure networks**"* |
| **SASL\_PLAINTEXT** | *"PLAINTEXT transport with SASL client authentication. **Some SASL mechanisms also support server authentication.**"* | ⚠️ *"**Does NOT support encryption** and hence is suitable **only within private networks**"* |
| **SASL\_SSL** | *"SSL transport with SASL authentication"* | ✅ *"suitable for use in insecure networks"* |
> ### TLS/SSL background [#tlsssl-background]
>
> *"TLS relies on a **Public Key Infrastructure (PKI)** to create, manage, and distribute digital certificates that can be used for **asymmetric encryption, AVOIDING THE NEED FOR DISTRIBUTING SHARED SECRETS** between servers and clients. **Session keys generated during the TLS handshake enable SYMMETRIC encryption with higher performance for subsequent data transfer.**"*
#### Multi-listener configuration — the canonical production example [#multi-listener-configuration--the-canonical-production-example]
```properties
listeners=EXTERNAL://:9092,INTERNAL://10.0.0.2:9093,BROKER://10.0.0.2:9094
advertised.listeners=EXTERNAL://broker1.example.com:9092,\
INTERNAL://broker1.local:9093,\
BROKER://broker1.local:9094
listener.security.protocol.map=EXTERNAL:SASL_SSL,INTERNAL:SSL,BROKER:SSL
inter.broker.listener.name=BROKER
```
**Selecting the inter-broker listener:** `inter.broker.listener.name` **or** `security.inter.broker.protocol`.
> ⚠️ *"**BOTH server-side AND client-side configuration options must be provided in the BROKER configuration** for the security protocol used for inter-broker communication. **This is because brokers need to establish CLIENT connections for that listener.**"*
**Client side:**
```properties
security.protocol=SASL_SSL
bootstrap.servers=broker1.example.com:9092,broker2.example.com:9092
```
> 💡 *"**Metadata returned to clients contains ONLY the endpoints corresponding to THE SAME LISTENER as the bootstrap servers.**"*
That last point matters operationally: a client bootstrapping against the EXTERNAL port will only ever be told about EXTERNAL endpoints. Listener isolation is enforced by the metadata response, not just by firewalls.
***
# 11.4 Security updates without downtime (/docs/kafka/securing-kafka/security-updates-without-downtime)
> *"Kafka deployments need regular maintenance to **rotate secrets, apply security fixes, and update to the latest security protocols.** Many of these are performed using **rolling updates**... Some tasks like **updating SSL key stores and trust stores can be performed using DYNAMIC CONFIG UPDATES WITHOUT RESTARTING BROKERS.**"*
#### PLAINTEXT → SASL\_SSL (add a listener, migrate, remove) [#plaintext--sasl_ssl-add-a-listener-migrate-remove]
#### PLAIN → SCRAM-SHA-256 (same listener port, five steps) [#plain--scram-sha-256-same-listener-port-five-steps]
Note the shape: **enable both → migrate clients → migrate inter-broker → remove the old.** Same pattern as the listener migration, one level down.
***
# 11.12 Self-test (/docs/kafka/securing-kafka/self-test)
Name the five security procedures and what each establishes.
List the seven guarantees a secure Kafka deployment must provide. Which one includes ZooKeeper, and why?
Draw the 2×2 of transport × authentication and name all four security protocols. Which two are safe on an insecure network?
Why must a broker's config include *client-side* settings for the inter-broker listener?
A client bootstraps against the EXTERNAL port. Which endpoints will it learn about, and why does that matter?
What principal do unauthenticated connections get, and in which two situations?
What's the CPU cost of SSL, and what Kafka optimization does it defeat?
Which stores does a broker need, and under what two conditions does it need a trust store?
Why must a broker certificate carry the hostname in SAN or CN? What attack does hostname verification prevent?
Contrast `ssl.client.auth=required` and `requested`. What principal does an unauthenticated client get under `requested`?
What becomes the `KafkaPrincipal` by default under SSL client auth, and how do you customize it?
When can you omit trust store configuration entirely?
How do you rotate a broker's key store without a restart?
Why are TLS handshakes a DoS vector, and what two controls mitigate it?
List Kafka's four SASL mechanisms. Which two have built-in implementations unsuitable for production, and why each?
Name the three SASL callback handler types and what each is for.
Why is `sasl.jaas.config` preferred over a JAAS config file?
Why must a Kerberos broker principal include the hostname? Why may client principals omit it?
Give three Kerberos-specific security considerations beyond "use TLS."
Why is the built-in SASL/PLAIN password store both insecure *and* inflexible?
How do you support password rotation on the server side? On the client side?
What does declaring a config with the `PASSWORD` type buy you?
What does SCRAM improve over PLAIN? What are its three security prerequisites?
You delete a SCRAM user. What immediately stops working, and what doesn't?
What are delegation tokens for? What SASL mechanism carries them, and what's the resulting principal?
What's the master-key rotation procedure for delegation tokens, and why is it disruptive?
What is the default behavior of a long-lived connection when its credentials expire? What config fixes it, and how is session lifetime computed?
A user's credentials are compromised. Give the full playbook in priority order, and explain why the first step is first.
Why doesn't SSL renegotiation help with revoked certificates?
Walk the four steps to migrate PLAINTEXT → SASL\_SSL with no downtime. What makes it safe?
Walk the five steps to migrate PLAIN → SCRAM on the same port.
Name the three encryption layers and the specific threat each addresses. What does the third one cover that the first two don't?
In end-to-end encryption, what does the broker see? Why does that matter for cloud deployments?
Why should you compress before encrypting — and then disable Kafka compression?
Why does naively encrypting message keys break Kafka? Give the two things it breaks and the recommended workaround.
What makes key rotation painful on a compacted topic?
List the six components of an ACL binding. Which permission type wins?
State the ACL evaluation rule and both implicit grants.
What ACLs does each of these need: a plain producer; an idempotent non-transactional producer; a transactional producer; a consumer in a group; a broker?
Why is `super.users` dangerous, and what's the separator gotcha?
Give both failure modes of `allow.everyone.if.no.acl.found=true`.
What does the request context give a custom authorizer? Name two things you could build with it.
Which log level records granted operations vs denied ones? What's the practical implication?
What information does a single TRACE/DEBUG request-log line contain? Name three distinct uses for it.
Why must `kerberos.removeHostFromPrincipal` and `removeRealmFromPrincipal` be true for ZooKeeper?
How does ZooKeeper's combined SASL+SSL authentication differ from Kafka's, and what's the authorization consequence?
What does `zookeeper.set.acl=true` do? What's readable by whom afterward?
Describe both mechanisms for keeping passwords out of config files. Where does the chain of trust bottom out in each?
What is a threat model, and which category of threat does the chapter specifically remind you not to forget?
**Previous:** [Chapter 10 — Cross-Cluster Data Mirroring](10-cross-cluster-data-mirroring.md)
**Next:** [Chapter 12 — Administering Kafka](12-administering-kafka.md)
# 14.8 What actually breaks in production — Ch. 14 consolidated (/docs/kafka/stream-processing/actually-breaks-production-ch)
| # | Symptom | Root cause | Fix |
| -- | -------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| 1 | **Aggregates reset to zero after a restart** | State kept in **local variables** — *"when the application is stopped or crashes, the state is lost, WHICH CHANGES THE RESULTS"* | Use a proper state store (Kafka Streams `Materialized` + changelog topic) |
| 2 | **Windowed results are wrong after a producer outage** | A producer returned with hours of backlog for **already-closed windows** | Configure a **grace period**; use **event time**, not processing time |
| 3 | **The same event produces different results in different runs** | Used **processing time** — *"it can even differ for TWO THREADS IN THE SAME APPLICATION"* | Event time via `TimestampExtractor` |
| 4 | **Timestamps are wrong for CDC-sourced events** | Kafka's auto timestamp is *record-creation* time, not the original event time | *"Add the event time as a FIELD IN THE RECORD ITSELF so BOTH timestamps are available"* |
| 5 | **Window results are nonsensical across regions** | Mixed time zones | Standardize the whole pipeline on one time zone; store the zone in the record if you must mix |
| 6 | **Aggregation is slow, spilling to disk** | Local state exceeds available memory | Partition into smaller substreams; *"spilling to disk has SIGNIFICANT PERFORMANCE IMPACT"* |
| 7 | **Database overwhelmed by enrichment lookups** | Per-record external lookup: **5–15 ms each**, and *"stream systems handle 100K–500K events/s but the DB only \~10K/s"* | **Stream-table join** over a CDC-maintained local copy |
| 8 | **Enrichment uses stale data** | Hand-rolled cache with a refresh interval — *"too often → hammering the DB; too long → stale"* | CDC-driven cache updates |
| 9 | **Local aggregation can't compute a global result (top 10)** | *"all the top 10 stocks could be in partitions assigned to OTHER instances"* | **Multiphase**: local aggregate → single-partition summary topic → single instance |
| 10 | **A Kafka Streams join refuses to run / produces nothing** | *"Kafka Streams REQUIRES that all topics that participate in a join have **THE SAME NUMBER OF PARTITIONS** and be **PARTITIONED BASED ON THE JOIN KEY**"* | Repartition or recreate topics to match |
| 11 | **A stream-stream join matches unrelated events** | Window too wide, or symmetric when causality is one-directional | `JoinWindows.of(1s).before(0s)` — express causality as window asymmetry |
| 12 | **Late events silently dropped** | No grace period configured | Define the reconciliation period; remember *"the longer the windows stay available, the MORE MEMORY is required"* |
| 13 | **Late corrections never reach consumers** | Output topic not compacted | *"Those are USUALLY COMPACTED TOPICS"* — a new result for the same key replaces the old |
| 14 | **Reprocessing corrupted results / lost data** | Reset offsets **and** state in place | *"OUR RECOMMENDATION IS TO USE THE FIRST METHOD"* — run a **second app as a new consumer group** and switch clients over |
| 15 | **`groupByKey()` doesn't group anything** | It doesn't — *"despite its name, this operation DOES NOT DO ANY GROUPING. Rather, it ensures the stream is PARTITIONED based on the record key"* | Understand it as a repartition assertion |
| 16 | **Serialization failure on an intermediate result** | Provided Serdes for input/output but not for the **aggregation result object** | *"Remember to provide a Serde for EVERY object you want to store in Kafka — input, output, AND, IN SOME CASES, INTERMEDIATE RESULTS"* |
| 17 | **Windowed output can't be deserialized** | Windowed Serde missing the window size — *"deserialization REQUIRES the window size, because ONLY THE START TIME is stored"* | `WindowedSerdes.timeWindowedSerdeFrom(class, windowSize)` |
| 18 | **Two Kafka Streams apps collide on internal topics/stores** | Duplicate `APPLICATION_ID_CONFIG` — it *"names the internal local stores AND THE TOPICS related to them"* | *"THIS NAME MUST BE UNIQUE for each Kafka Streams application working with the same Kafka cluster"* |
| 19 | **Enabled topology optimization; nothing changed** | Called `build()` without the props | *"If you only call `build()` without passing the config, OPTIMIZATION IS STILL DISABLED"* |
| 20 | **Optimization changed the results** | Optimizations rewrite the physical plan | *"TEST with and without... and VALIDATE THAT THE RESULTS ARE IDENTICAL in various known scenarios"* |
| 21 | **Unit tests pass but production fails** | `TopologyTestDriver` *"does NOT simulate Kafka Streams CACHING behavior... there are ENTIRE CLASSES OF ERRORS IT WILL NOT DETECT"* | Add integration tests — **Testcontainers** preferred |
| 22 | **Adding threads doesn't increase throughput** | Task count == partition count; you've saturated it | More partitions (mind Ch. 3 §9.4); more instances only helps up to partition count |
| 23 | **After an instance fails, a subset of keys goes stale for minutes** | Task must **replay the changelog** to warm the state store | **`segment.bytes` = 100 MB** (not 1 GB) + low `min.compaction.lag.ms`; and **standby replicas** |
| 24 | **Recovery replays far more than expected** | The **active segment is never compacted** — with 1 GB segments, up to 1 GB/partition is uncompacted | Same fix as #23 |
| 25 | **Changelog topics grow without bound** | Compaction not working on internal topics (Ch. 13 §6: *silently halted cleaner threads*) | Kafka uses compaction *"to make sure they don't grow endlessly and that re-creating the state is ALWAYS FEASIBLE"* — verify it's actually running |
| 26 | **Duplicate contributions to aggregates after a failure** | At-least-once processing | `processing.guarantee=exactly_once` (or `exactly_once_beta` on 2.6+/2.5+ brokers) |
| 27 | **Rebalances pause the whole app** | Eager rebalancing | Kafka Streams inherits **cooperative rebalancing** and **static group membership** (Ch. 4) |
| 28 | **Built a stream processing app for pure data movement** | Wrong tool | *"Reconsider whether you want... a SIMPLER INGEST-FOCUSED SYSTEM LIKE KAFKA CONNECT"* |
| 29 | **Sub-millisecond SLA not met by a streaming app** | Wrong paradigm | *"REQUEST-RESPONSE PATTERNS ARE OFTEN BETTER SUITED"*; if you must stream, avoid microbatch frameworks |
***
# 14.9 Consolidated reference (/docs/kafka/stream-processing/consolidated-reference)
#### The design patterns and when each applies [#the-design-patterns-and-when-each-applies]
#### Kafka Streams operational checklist [#kafka-streams-operational-checklist]
#### How this chapter closes the loop on the whole book [#how-this-chapter-closes-the-loop-on-the-whole-book]
***
# 14.0 Further reading — the chapter's own bibliography (/docs/kafka/stream-processing/further-reading-chapter-s)
> *"This chapter is intended as **just a quick introduction to the large and fascinating world of stream processing and Kafka Streams. THERE ARE ENTIRE BOOKS WRITTEN ON THESE SUBJECTS.**"*
**Concepts — "the basic concepts of stream processing from a DATA ARCHITECTURE perspective":**
| Book | Author(s) | Publisher | What it's for |
| ------------------------------------- | ---------------------------------------- | --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Making Sense of Stream Processing** | Martin Kleppmann | O'Reilly | *"discusses the benefits of **RETHINKING APPLICATIONS AS STREAM PROCESSING APPLICATIONS** and how to **REORIENT DATA ARCHITECTURES around the idea of event streams**"* |
| **Streaming Systems** | Tyler Akidau, Slava Chernyak, Reuven Lax | O'Reilly | *"a great **GENERAL INTRODUCTION** to the topic of stream processing and some of the **basic ideas in the space**"* |
| **Flow Architectures** | James Urquhart | O'Reilly | *"**targeted at CTOs** and discusses **the IMPLICATIONS of stream processing TO THE BUSINESS**"* |
**Frameworks — "specific details of specific frameworks":**
| Book | Author(s) | Publisher |
| ------------------------------------------------- | ------------------------------- | --------- |
| **Mastering Kafka Streams and ksqlDB** | Mitch Seymour | O'Reilly |
| **Kafka Streams in Action** | William P. Bejeck Jr. | Manning |
| **Event Streaming with Kafka Streams and ksqlDB** | William P. Bejeck Jr. | Manning |
| **Stream Processing with Apache Flink** | Fabian Hueske, Vasiliki Kalavri | O'Reilly |
| **Stream Processing with Apache Spark** | Gerard Maas, Francois Garillot | O'Reilly |
**Referenced elsewhere in the chapter:**
* **"There Is No Now"** — Justin Sheehy — *"an excellent paper"* on how complex time gets in distributed systems (§2.2)
* **"Crossing the Streams"** — Kafka Summit 2020 talk, plus a more in-depth blog post — on **foreign-key joins** (§3.5)
* **"Beyond the DSL"** — presentation — *"a great introduction"* to the low-level **Processor API** (§4)
* **"Testing Kafka Streams — A Deep Dive"** — blog post — deeper explanations and detailed code examples of topologies and tests (§5.3)
* A **blog post and a Kafka Summit talk** on Kafka Streams **scalability and high availability** (§5.6)
* The Kafka Streams **developer guide** — for the low-level Processor API (§4)
> ⚠️ **Version caveat:** *"Kafka Streams is **still an evolving framework. EVERY MAJOR RELEASE DEPRECATES APIs AND MODIFIES SEMANTICS.** This chapter documents APIs and semantics **as of Apache Kafka 2.8.** We avoided using any API planned for deprecation in 3.0, **but our discussion of JOIN SEMANTICS and TIMESTAMP HANDLING does NOT include any of the changes planned for release 3.0.**"*
***
# 14.7 How to choose a stream processing framework (/docs/kafka/stream-processing/how-choose-stream-processing)
#### 7.1 Four application types → four different answers [#71-four-application-types--four-different-answers]
Note that **two of the four answers are "you may not want stream processing at all."** That's an unusually honest framing for a chapter selling a stream processing library.
#### 7.2 Four global considerations [#72-four-global-considerations]
| Criterion | The questions to ask |
| ------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| **Operability of the system** | *"Is it easy to **deploy** to production? Is it easy to **monitor and troubleshoot**? Is it easy to **scale up and down**? Does it **integrate well with your existing infrastructure**? **What if there is a mistake and you need to REPROCESS data?**"* |
| **Usability of APIs and ease of debugging** | *"I've seen **ORDERS-OF-MAGNITUDE DIFFERENCES in the time it takes to write a high-quality application AMONG DIFFERENT VERSIONS OF THE SAME FRAMEWORK. Development time and time-to-market are important**, so you need to choose a system that makes you efficient."* |
| **Makes hard things easy** | *"**ALMOST EVERY SYSTEM WILL CLAIM they can do advanced windowed aggregations and maintain local stores, BUT THE QUESTION IS: DO THEY MAKE IT EASY FOR YOU? Do they handle GRITTY DETAILS around SCALE AND RECOVERY, OR DO THEY SUPPLY LEAKY ABSTRACTIONS AND MAKE YOU HANDLE MOST OF THE MESS?**"* |
| **Community** | *"there's **no replacement for a vibrant and active community.** Good community means you get **new features regularly**, the quality is relatively good (**no one wants to work on bad software**), **bugs get fixed quickly**, and user questions get answers. **It also means that if you get a strange error and Google it, YOU WILL FIND INFORMATION ABOUT IT because other people are using this system and seeing the same issues.**"* |
***
# 14. Stream Processing (/docs/kafka/stream-processing)
> *Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 14*
> **The historical framing:** *"Kafka was traditionally seen as a powerful message bus, capable of delivering streams of events **but without processing or transformation capabilities.** ... many companies had a system containing **many streams of interesting data, stored for long amounts of time and perfectly ordered, JUST WAITING FOR SOME STREAM PROCESSING FRAMEWORK TO SHOW UP AND PROCESS THEM.** In other words, **in the same way that data processing was significantly more difficult BEFORE DATABASES WERE INVENTED, STREAM PROCESSING WAS HELD BACK BY THE LACK OF A STREAM PROCESSING PLATFORM.**"*
**Since 0.10.0, Kafka ships Kafka Streams** — *"a powerful stream processing library as part of its collection of client libraries... This allows developers to **consume, process, and produce events IN THEIR OWN APPS, WITHOUT RELYING ON AN EXTERNAL PROCESSING FRAMEWORK.**"*
***
# 14.5 Kafka Streams architecture (/docs/kafka/stream-processing/kafka-streams-architecture)
#### 5.1 Building a topology [#51-building-a-topology]
> *"Topology (also called **DAG**, or directed acyclic graph) is **a set of operations and transitions that EVERY EVENT MOVES THROUGH from input to output.** ... **Even a simple app has a NONTRIVIAL topology.**"*
#### 5.2 💡 Optimizing a topology — three steps, and why step 2 matters [#52--optimizing-a-topology--three-steps-and-why-step-2-matters]
> *"**By DEFAULT, Kafka Streams executes applications built with the DSL API by MAPPING EACH DSL METHOD INDEPENDENTLY to a lower-level equivalent. By evaluating each DSL method independently, OPPORTUNITIES TO OPTIMIZE THE OVERALL RESULTING TOPOLOGY WERE MISSED.**"*
**Enabling it:**
```java
StreamsConfig.TOPOLOGY_OPTIMIZATION = StreamsConfig.OPTIMIZE
builder.build(props) // ⚠ MUST pass the config
```
> ⚠️ *"**If you only call `build()` WITHOUT PASSING THE CONFIG, OPTIMIZATION IS STILL DISABLED.**"*
>
> *"Currently, Apache Kafka only contains **a few optimizations, mostly around REUSING TOPICS where possible.** It is recommended to **test applications with AND without optimizations and to COMPARE EXECUTION TIMES AND VOLUMES OF DATA WRITTEN TO KAFKA, and of course, VALIDATE THAT THE RESULTS ARE IDENTICAL in various known scenarios.**"*
#### 5.3 Testing a topology [#53-testing-a-topology]
**The primary tool: `TopologyTestDriver`**
> *"Since its introduction in version **1.1.0**, its API has undergone significant improvements, and **versions since 2.4 are convenient and easy to use.** These tests **look like normal unit tests. We define input data, produce it to MOCK INPUT TOPICS, run the topology with the test driver, read the results from MOCK OUTPUT TOPICS, and validate.**"*
> ### ⚠️ The gap in `TopologyTestDriver` [#️-the-gap-in-topologytestdriver]
>
> *"since it **DOES NOT SIMULATE KAFKA STREAMS CACHING BEHAVIOR** (an optimization... **entirely unrelated to the state store itself**, which IS simulated), **THERE ARE ENTIRE CLASSES OF ERRORS THAT IT WILL NOT DETECT.**"*
**Two integration test frameworks:**
| Framework | How | Verdict |
| ------------------------ | ---------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| **EmbeddedKafkaCluster** | *"runs Kafka brokers **inside the JVM that runs the tests**"* | |
| **Testcontainers** | *"runs **Docker containers** with Kafka brokers (and many other components, as needed)"* | ✅ **Recommended** — *"since by using Docker it **FULLY ISOLATES Kafka, its dependencies, and its RESOURCE USAGE from the application we are trying to test**"* |
*Further reading: the "Testing Kafka Streams — A Deep Dive" blog post.*
#### 5.4 💡 Scaling a topology — tasks are the unit of parallelism [#54--scaling-a-topology--tasks-are-the-unit-of-parallelism]
> *"Kafka Streams scales by **allowing MULTIPLE THREADS of executions within one instance** AND by **supporting LOAD BALANCING between DISTRIBUTED INSTANCES.** We can run on one machine with multiple threads or on multiple machines; **in either case, ALL ACTIVE THREADS WILL BALANCE THE WORK.**"*
**The scaling recipe:**
*(Partition count is once again the ceiling on parallelism — Ch. 2, Ch. 4, and now Ch. 14.)*
##### Task dependency case 1: joins [#task-dependency-case-1-joins]
> \*"if we join two streams... **we need data from a partition in EACH stream before we can emit a result.** Kafka Streams handles this by **ASSIGNING ALL THE PARTITIONS NEEDED FOR ONE JOIN TO THE SAME TASK** so that the task can consume from all the relevant partitions and perform the join independently.
>
> ⚠️ **THIS IS WHY KAFKA STREAMS CURRENTLY REQUIRES THAT ALL TOPICS THAT PARTICIPATE IN A JOIN OPERATION HAVE THE SAME NUMBER OF PARTITIONS AND BE PARTITIONED BASED ON THE JOIN KEY.**"\*
**That's a hard, checkable precondition** — the most common cause of "my Kafka Streams join won't start."
##### Task dependency case 2: repartitioning (the shuffle) [#task-dependency-case-2-repartitioning-the-shuffle]
> \*"in the ClickStream example, all our events are keyed by **user ID. But what if we want to generate statistics per PAGE? Or per ZIP CODE?** Kafka Streams will **REPARTITION the data by zip code** and run an aggregation with the new partitions. If task 1... reaches a processor that repartitions the data (a `groupBy` operation), **it will need to SHUFFLE, or send events to other tasks.**
>
> 💡 **UNLIKE OTHER STREAM PROCESSOR FRAMEWORKS, KAFKA STREAMS REPARTITIONS BY WRITING THE EVENTS TO A NEW TOPIC WITH NEW KEYS AND PARTITIONS. Then ANOTHER SET OF TASKS reads events from the new topic and continues processing.**"\*
**The shuffle becomes a Kafka topic.** That single decision is what makes Kafka Streams a library rather than a cluster framework: there is no shuffle service, no inter-worker network protocol, no coordinated barrier — just producers and consumers.
#### 5.5 Surviving failures [#55-surviving-failures]
#### ⚠️ 5.6 The real problem: recovery *speed* [#️-56-the-real-problem-recovery-speed]
> *"**While the high-availability methods described here work well in theory, REALITY INTRODUCES SOME COMPLEXITY. ONE IMPORTANT CONCERN IS THE SPEED OF RECOVERY.** When a thread has to start processing a task that used to run on a failed thread, **it FIRST NEEDS TO RECOVER ITS SAVED STATE — the current aggregation windows, for instance. Often this is done by REREADING INTERNAL TOPICS from Kafka in order to WARM UP Kafka Streams state stores. DURING THE TIME IT TAKES TO RECOVER THE STATE OF A FAILED TASK, THE STREAM PROCESSING JOB WILL NOT MAKE PROGRESS ON THAT SUBSET OF ITS DATA, LEADING TO REDUCED AVAILABILITY AND STALE DATA.**"*
**Two techniques, both specific and actionable:**
***
# 14.4 Kafka Streams by example (/docs/kafka/stream-processing/kafka-streams-example)
> *"Apache Kafka has **two stream APIs — a low-level PROCESSOR API and a high-level STREAMS DSL.**"* The DSL *"allows us to define the application by defining **a CHAIN OF TRANSFORMATIONS** to events... Transformations can be as simple as a filter or as complex as a stream-to-stream join. **The lower-level API allows us to create OUR OWN transformations.**"*
**The lifecycle:**
#### 4.1 Word count — map/filter + local-state aggregation [#41-word-count--mapfilter--local-state-aggregation]
**Configuration:**
```java
Properties props = new Properties();
props.put(StreamsConfig.APPLICATION_ID_CONFIG, "wordcount"); // ①
props.put(StreamsConfig.BOOTSTRAP_SERVERS_CONFIG, "localhost:9092"); // ②
props.put(StreamsConfig.DEFAULT_KEY_SERDE_CLASS_CONFIG, // ③
Serdes.String().getClass().getName());
props.put(StreamsConfig.DEFAULT_VALUE_SERDE_CLASS_CONFIG,
Serdes.String().getClass().getName());
```
① *"**Every Kafka Streams application MUST have an application ID.** It is used to **coordinate the instances** of the application and also **when NAMING THE INTERNAL LOCAL STORES AND THE TOPICS related to them. THIS NAME MUST BE UNIQUE for each Kafka Streams application working with the same Kafka cluster.**"*
② *"Kafka Streams applications **also use Kafka for COORDINATION.**"*
③ Default **Serdes** — *"If needed, we can **override these defaults later** when building the streams topology."*
> *"you can also configure **the producer and consumer EMBEDDED in Kafka Streams** by adding **any producer or consumer config** to the `Properties` object"* — i.e. everything from Ch. 3 and Ch. 4 applies.
**The topology:**
```java
StreamsBuilder builder = new StreamsBuilder();
KStream source = builder.stream("wordcount-input"); // ①
final Pattern pattern = Pattern.compile("\\W+");
KStream counts = source.flatMapValues(value-> // ②
Arrays.asList(pattern.split(value.toLowerCase())))
.map((key, value) -> new KeyValue(value, value))// ②
.filter((key, value) -> (!value.equals("the"))) // ③
.groupByKey() // ④
.count().mapValues(value-> Long.toString(value)).toStream(); // ⑤⑥
counts.to("wordcount-output"); // ⑦
```
① Point at the input topic.
② *"Each event is a **line of words**; we split it up using a regular expression into a series of individual words. Then we take each word (**currently a VALUE**) and **PUT IT IN THE EVENT RECORD KEY so it can be used in a group-by operation.**"*
③ *"We filter out the word `the`, **just to show how easy filtering is.**"*
④ *"we **group by key**, so we now have **a collection of events for each unique word**."*
⑤ *"We **count** how many events we have in each collection."*
⑥ *"The result of counting is a **`Long`**. We convert it to a **`String`** so it will be easier for humans to read."*
⑦ Write back to Kafka.
```java
KafkaStreams streams = new KafkaStreams(builder.build(), props);
streams.start();
// usually the stream application would be running forever
Thread.sleep(5000L);
streams.close();
```
#### 💡 The deployment payoff [#-the-deployment-payoff]
> \*"we can run the entire example **on our machine WITHOUT INSTALLING ANYTHING EXCEPT APACHE KAFKA.** If our input topic contains multiple partitions, we can **run MULTIPLE instances of the WordCount application (just run the app in several different terminal tabs), AND WE HAVE OUR FIRST KAFKA STREAMS PROCESSING CLUSTER.** The instances **talk to one another and coordinate the work.**
>
> **ONE OF THE BIGGEST BARRIERS TO ENTRY FOR SOME STREAM PROCESSING FRAMEWORKS IS THAT LOCAL MODE IS VERY EASY TO USE, BUT THEN TO RUN A PRODUCTION CLUSTER, WE NEED TO INSTALL YARN OR MESOS, THEN INSTALL THE PROCESSING FRAMEWORK ON ALL THOSE MACHINES, AND THEN LEARN HOW TO SUBMIT OUR APP TO THE CLUSTER. WITH KAFKA'S STREAMS API, WE JUST START MULTIPLE INSTANCES OF OUR APP — AND WE HAVE A CLUSTER. THE EXACT SAME APP IS RUNNING ON OUR DEVELOPMENT MACHINE AND IN PRODUCTION.**"\*
*(This is Ch. 1's "APIs and libraries, not a structured runtime like YARN" design stance, cashed in.)*
#### 4.2 Stock market statistics — windowed aggregation [#42-stock-market-statistics--windowed-aggregation]
**The goal:** from a stream of trades (ticker, ask price, ask size), produce **per-five-second-window**: best (minimum) ask price, number of trades, and average ask price — *"All statistics will be updated **every second**."*
**Serdes are the main config difference:**
```java
props.put(StreamsConfig.DEFAULT_VALUE_SERDE_CLASS_CONFIG, TradeSerde.class.getName());
```
```java
static public final class TradeSerde extends WrapperSerde {
public TradeSerde() {
super(new JsonSerializer(), new JsonDeserializer(Trade.class));
}
}
```
> *"we used the **Gson library from Google** to generate a JSON serializer and deserializer from our Java object. Then we created a small wrapper..."*
>
> 💡 *"**Nothing fancy, but remember to provide a Serde object for EVERY object you want to store in Kafka — INPUT, OUTPUT, AND, IN SOME CASES, INTERMEDIATE RESULTS.** To make this easier, we recommend **generating these Serdes through a library like Gson, Avro, Protobuf, or something similar.**"*
**The topology:**
```java
KStream, TradeStats> stats = source
.groupByKey() // ①
.windowedBy(TimeWindows.of(Duration.ofMillis(windowSize)) // ②
.advanceBy(Duration.ofSeconds(1)))
.aggregate( // ③
() -> new TradeStats(), // ③
(k, v, tradestats) -> tradestats.add(v), // ④
Materialized.>
as("trade-aggregates") // ⑤
.withValueSerde(new TradeStatsSerde())) // ⑥
.toStream() // ⑦
.mapValues((trade) -> trade.computeAvgPrice()); // ⑧
stats.to("stockstats-output", // ⑨
Produced.keySerde(
WindowedSerdes.timeWindowedSerdeFrom(String.class, windowSize)));
```
**① ⚠️ The `groupByKey()` misnomer:**
> *"**Despite its name, this operation DOES NOT DO ANY GROUPING. Rather, it ENSURES THAT THE STREAM OF EVENTS IS PARTITIONED BASED ON THE RECORD KEY.** Since we wrote the data into a topic with a key and didn't modify the key before calling `groupByKey()`, **the data is still partitioned by its key — SO THIS METHOD DOES NOTHING IN THIS CASE.**"*
**② The window:** 5 seconds, advancing every second → **a hopping window** (advance \< size), so windows overlap.
**③④ The aggregation:** *"The `aggregate` method will **split the stream into OVERLAPPING WINDOWS** (a five-second window every second) and then apply an aggregate method on all the events in the window."* First parameter = **an initializer** producing the result object; second = **the aggregator** (`add` updates minimum price, trade count, total prices).
**⑤⑥ The state store:** *"windowing aggregation **requires maintaining a state and a LOCAL STORE.** The last parameter is the **configuration of the state store. `Materialized` is the store configuration object**, and we configure the store name as `trade-aggregates`. **This can be ANY unique name.**"* Plus a Serde for the aggregation result.
**⑦ Table → stream:** *"The result of the aggregation is **A TABLE with THE TICKER AND THE TIME WINDOW AS THE PRIMARY KEY** and the aggregation result as the value. We are **turning the table back into a stream of events.**"*
**⑧ Deriving the average:** *"right now the aggregation results include **the SUM of prices and NUMBER of trades.** We go over these records and **use the existing statistics to calculate average price.**"* — a nice illustration: you aggregate *sum + count*, not *average*, because average isn't associative.
**⑨ The windowed Serde:** *"Since the results are part of a windowing operation, we create a **`WindowedSerde` that stores the result in a windowed data format that INCLUDES THE WINDOW TIMESTAMP. The window size is passed as part of the Serde, EVEN THOUGH IT ISN'T USED IN THE SERIALIZATION (DESERIALIZATION REQUIRES THE WINDOW SIZE, BECAUSE ONLY THE START TIME OF THE WINDOW IS STORED IN THE OUTPUT TOPIC).**"*
> 💡 *"**One thing to notice is HOW LITTLE WORK WAS NEEDED TO MAINTAIN THE LOCAL STATE of the aggregation — JUST PROVIDE A SERDE AND NAME THE STATE STORE. Yet this application will SCALE TO MULTIPLE INSTANCES and AUTOMATICALLY RECOVER FROM A FAILURE of each instance** by shifting processing of some partitions to one of the surviving instances."*
#### 4.3 ClickStream enrichment — both join types in one topology [#43-clickstream-enrichment--both-join-types-in-one-topology]
**The business goal:** *"join all three streams to get **a 360-degree view into each user activity. What did the users search for? What did they click as a result? Did they change their 'interests' in their user profile?** ... Product recommendations are often based on this kind of information — **the user searched for bikes, clicked on links for 'Trek,' and is interested in travel, so we can advertise bikes from Trek, helmets, and bike tours to exotic locations like Nebraska.**"*
```java
KStream views = // ①
builder.stream(Constants.PAGE_VIEW_TOPIC,
Consumed.with(Serdes.Integer(), new PageViewSerde()));
KStream searches = // ①
builder.stream(Constants.SEARCH_TOPIC,
Consumed.with(Serdes.Integer(), new SearchSerde()));
KTable profiles = // ②
builder.table(Constants.USER_PROFILE_TOPIC,
Consumed.with(Serdes.Integer(), new ProfileSerde()));
// ── STREAM-TABLE JOIN ────────────────────────────────────────────────────
KStream viewsWithProfile = views.leftJoin(profiles, // ③
(page, profile) -> { // ④
if (profile != null)
return new UserActivity(
profile.getUserID(), profile.getUserName(),
profile.getZipcode(), profile.getInterests(),
"", page.getPage());
else
return new UserActivity(-1, "", "", null, "", page.getPage());
});
// ── STREAM-STREAM (WINDOWED) JOIN ───────────────────────────────────────
KStream userActivityKStream =
viewsWithProfile.leftJoin(searches, // ⑤
(userActivity, search) -> { // ⑥
if (search != null)
userActivity.updateSearch(search.getSearchTerms());
else
userActivity.updateSearch("");
return userActivity;
},
JoinWindows.of(Duration.ofSeconds(1)).before(Duration.ofSeconds(0)), // ⑦
StreamJoined.with(Serdes.Integer(), // ⑧
new UserActivitySerde(),
new SearchSerde()));
```
② *"We also define **a `KTable`** for the user profiles. **A `KTable` is A MATERIALIZED STORE THAT IS UPDATED THROUGH A STREAM OF CHANGES.**"*
③ *"**In a stream-table join, EACH EVENT IN THE STREAM RECEIVES INFORMATION FROM THE CACHED COPY of the profile table.** We are doing a **left-join, so clicks without a known user WILL BE PRESERVED.**"*
④ 💡 *"**Unlike in databases, WE GET TO DECIDE HOW TO COMBINE THE TWO VALUES INTO ONE RESULT.**"*
⑦ **The interesting part:**
> *"a stream-to-stream join is **a join with A TIME WINDOW. Joining ALL clicks and searches for each user DOESN'T MAKE MUCH SENSE — we want to join each search with clicks THAT ARE RELATED TO IT**, that is, clicks that occurred **a short period of time AFTER the search.** So we define a join window of one second. **We invoke `of` to create a window of one second BEFORE AND AFTER each search, and then we call `before` with A ZERO-SECONDS INTERVAL to make sure we ONLY JOIN CLICKS THAT HAPPEN ONE SECOND AFTER EACH SEARCH AND NOT BEFORE.**"*
⑧ *"the Serde of the join result... a Serde for **the key that both sides of the join have in common** and the Serde for **both values** that will be included in the result."*
**The two patterns, summarized:**
> *"One joins **a stream with a table to ENRICH all streaming events** with information in the table. **This is similar to joining a FACT TABLE with a DIMENSION when running queries on a data warehouse.** The second example joins **two streams based on a TIME WINDOW. THIS OPERATION IS UNIQUE TO STREAM PROCESSING.**"*
***
# 14.10 Self-test (/docs/kafka/stream-processing/self-test)
Define a data stream. What single word is doing most of the work?
Give the three attributes of event streams beyond unboundedness, and contrast each with a database table.
Explain the deposit/withdrawal example. What does it prove about ordering?
A canceled transaction — what happens to the original event? What's the redo-log analogy?
Why does the book credit replayability specifically for Kafka's success in stream processing?
What does the definition of stream processing deliberately say *nothing* about?
Contrast the three programming paradigms on latency, blocking behavior, and database analogue.
What is the "gap" that stream processing fills, stated in concrete numbers?
A job runs at 2 a.m., reads 500 records, outputs a result, and exits. Is it stream processing? Why not?
Name the three notions of time. Which matters most, which should be avoided, and why?
When is log-append time an acceptable substitute for event time?
Give all four rules Kafka Streams uses to assign output timestamps.
Why must the whole pipeline share one time zone, and what do you do if it can't?
Why is storing state in a local variable unreliable? (The book admits doing this — where?)
Contrast local and external state on speed, size, sharing, and availability.
What single design consequence follows from local state's memory limit?
State the stream-table duality. What can a table answer that a stream can't, and vice versa?
How do you convert a table to a stream? A stream to a table? What's the name for the second operation?
Name the three window dimensions. What's the trade-off in window size?
Define hopping, tumbling, and session windows.
What's a grace period, and what question does it answer?
Why does key-based partitioning make local state *correct*? Name the two Kafka guarantees involved.
Name the three problems local state creates and how Kafka Streams solves the second one (two mechanisms).
Why is log compaction essential to the changelog-topic approach?
Why can't a local-state aggregate compute the daily top 10? Describe the multiphase solution and why phase 2 can be single-instance.
What does Kafka Streams do that MapReduce didn't, regarding multiple reduce phases?
Give the three problems with per-record external lookups, with numbers.
State the caching dilemma, and how CDC resolves it.
Why is a table-table join never windowed?
Why is a stream-stream join necessarily windowed? What must be true of the two topics?
How does Kafka Streams guarantee one task sees both sides of a join for a given key?
Give three real-world causes of out-of-sequence events.
List the four things an app must do to handle late events. Which one is the fundamental difference from batch?
How does Kafka Streams implement late-window correction? What Kafka feature does it rely on?
Give both reprocessing variants and the three steps of the recommended one. Why is it "much safer"?
What are interactive queries, and when are they worth it?
What does `APPLICATION_ID_CONFIG` control besides coordination?
What does `groupByKey()` actually do? When is it a no-op?
Why do you aggregate sum-and-count rather than average?
Why does the windowed Serde need the window size even though it isn't serialized?
Why is "just start multiple instances" a significant claim? What does it save you compared with other frameworks?
Name the three processor kinds in a topology. What must a topology start and end with?
Give the three execution steps of a Kafka Streams app. Which one admits optimization, and what's the easy mistake when enabling it?
What does `TopologyTestDriver` *not* simulate? Which integration framework is recommended, and why?
What determines the number of tasks? What are the two ways to scale, and what caps both?
How does Kafka Streams implement a shuffle, and why is that architecturally significant?
Why can the two sides of a repartition run fully in parallel despite one depending on the other?
Kafka Streams always recovers. So what is the actual availability problem, and what are the two specific fixes?
Why does segment size affect recovery time? (Reference the active-segment rule.)
What is beaconing, and why is it a good fit for stream processing?
For each of the four application types, say whether stream processing is the right answer and what to look for.
Give the four global framework-selection criteria. Which one is about leaky abstractions?
**Previous:** [Chapter 13 — Monitoring Kafka](13-monitoring-kafka.md)
**Next:** [Appendices A & B](15-appendices.md) · [Back to the index](README.md)
# 14.2 Stream processing concepts (/docs/kafka/stream-processing/stream-processing-concepts)
#### 2.1 Topology [#21-topology]
> *"A processing topology **starts with one or more SOURCE STREAMS** that are passed through **a graph of STREAM PROCESSORS connected through event streams**, until results are written to **one or more SINK STREAMS.** Each stream processor is **a computational step applied to the stream of events in order to transform the events.**"*
Examples used in the chapter: **filter, count, group-by, left-join.**
#### 2.2 ⚠️ Time — "probably the most important concept in stream processing and often the most confusing" [#22-️-time--probably-the-most-important-concept-in-stream-processing-and-often-the-most-confusing]
> *"having a common notion of time is critical because **most stream applications perform operations on TIME WINDOWS.** For example, our stream application might calculate a **moving five-minute average of stock prices.** In that case, **we need to know what to do when one of our producers goes offline for two hours due to network issues and RETURNS WITH TWO HOURS' WORTH OF DATA — most of the data will be relevant for five-minute time windows THAT HAVE LONG PASSED AND FOR WHICH THE RESULT WAS ALREADY CALCULATED AND STORED.**"*
*(Recommended reading: Justin Sheehy's paper "There Is No Now.")*
**The three notions of time:**
**Kafka Streams' mechanism:** *"assigns time to each event based on the **`TimestampExtractor`** interface. Developers can use different implementations, which can use **either of the three time semantics or A COMPLETELY DIFFERENT CHOICE OF TIMESTAMP, INCLUDING EXTRACTING A TIMESTAMP FROM THE CONTENTS OF THE EVENT ITSELF.**"*
#### 💡 The four output-timestamp rules [#-the-four-output-timestamp-rules]
> *"When Kafka Streams writes output to a Kafka topic, it assigns a timestamp to each event based on the following rules:"*
> *"When using the **lower-level Processor API** rather than the DSL, Kafka Streams includes **APIs for MANIPULATING THE TIMESTAMPS OF RECORDS DIRECTLY**, so developers can implement timestamp semantics that match the required business logic."*
> ### ⚠️ MIND THE TIME ZONE [#️-mind-the-time-zone]
>
> *"**THE ENTIRE DATA PIPELINE SHOULD STANDARDIZE ON A SINGLE TIME ZONE; OTHERWISE, RESULTS OF STREAM OPERATIONS WILL BE CONFUSING AND OFTEN MEANINGLESS.** If you must handle data streams with different time zones, you need to make sure you can **convert events to a single time zone BEFORE performing operations on time windows. Often this means STORING THE TIME ZONE IN THE RECORD ITSELF.**"*
#### 2.3 State [#23-state]
**When you don't need it:** *"if all we need to do is read a stream of online shopping transactions, find the transactions over $10,000, and email the relevant salesperson, we can probably write this in **just a few lines of code** using a Kafka consumer and SMTP library."*
**When you do:** *"Stream processing becomes really interesting when we have operations that **involve MULTIPLE EVENTS: counting the number of events by type, moving averages, joining two streams**... we need to keep track of more information... **We call this information a STATE.**"*
> ### ⚠️ The local-variable trap [#️-the-local-variable-trap]
>
> *"It is often tempting to **store the state in variables that are LOCAL to the stream processing app**, such as a simple hash table. **In fact, WE DID JUST THAT IN MANY EXAMPLES IN THIS BOOK. HOWEVER, THIS IS NOT A RELIABLE APPROACH because WHEN THE APPLICATION IS STOPPED OR CRASHES, THE STATE IS LOST, WHICH CHANGES THE RESULTS.**"*
*(That's the honest self-critique of Ch. 4's moving-average and word-count examples — and Ch. 7 §5.2's "consumers may need to maintain state.")*
**Two kinds of state:**
| | Local / internal | External |
| ---------- | ------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Where | *"an embedded, in-memory database running WITHIN the application"* | *"an external data store, often a NoSQL system like **Cassandra**"* |
| Visibility | *"accessible ONLY by a specific INSTANCE"* | *"can be accessed from **multiple instances** or even **different applications**"* |
| ✅ | **"extremely fast"** | *"virtually **unlimited size**"* |
| ❌ | *"we are **limited to the amount of memory available**"* | *"the extra **LATENCY and COMPLEXITY** introduced with an additional system, as well as **AVAILABILITY** — the application needs to handle the possibility that the external system is not available"* |
> 💡 *"As a result, **MANY OF THE DESIGN PATTERNS IN STREAM PROCESSING FOCUS ON WAYS TO PARTITION THE DATA INTO SUBSTREAMS THAT CAN BE PROCESSED USING A LIMITED AMOUNT OF LOCAL STATE.**"*
> *"Most stream processing apps **try to AVOID having to deal with an external store**, or at least limit the latency overhead by **caching information in the local state and communicating with the external store AS RARELY AS POSSIBLE. THIS USUALLY INTRODUCES CHALLENGES WITH MAINTAINING CONSISTENCY BETWEEN THE INTERNAL AND EXTERNAL STATE.**"*
#### 💡 2.4 Stream-table duality — the central idea [#-24-stream-table-duality--the-central-idea]
> \*"A **TABLE** is a collection of records, each identified by its primary key... **Table records are MUTABLE.** Querying a table allows **checking the state of the data AT A SPECIFIC POINT IN TIME.** ... **Unless the table was specifically designed to include history, WE WILL NOT FIND THEIR PAST CONTACTS in the table.**
>
> **Unlike tables, STREAMS CONTAIN A HISTORY OF CHANGES.** A stream is a string of events wherein **each event CAUSED a change. A table contains a CURRENT STATE of the world, which is THE RESULT of many changes.**
>
> **From this description, it is clear that STREAMS AND TABLES ARE TWO SIDES OF THE SAME COIN — the world always changes, and SOMETIMES WE ARE INTERESTED IN THE EVENTS THAT CAUSED THOSE CHANGES, whereas OTHER TIMES WE ARE INTERESTED IN THE CURRENT STATE of the world. SYSTEMS THAT ALLOW US TO TRANSITION BACK AND FORTH BETWEEN THE TWO WAYS OF LOOKING AT DATA ARE MORE POWERFUL THAN SYSTEMS THAT SUPPORT JUST ONE.**"\*
**The shoe-store example:**
#### 2.5 Time windows — the three dimensions nobody thinks about [#25-time-windows--the-three-dimensions-nobody-thinks-about]
> *"**Most operations on streams are WINDOWED operations**, operating on slices of time: moving averages, top products sold this week, 99th percentile load... **Join operations on two streams are ALSO windowed.** **VERY FEW PEOPLE STOP AND THINK ABOUT THE TYPE OF WINDOW THEY WANT.**"*
**Aligned vs unaligned windows:**
#### 2.6 Processing guarantees [#26-processing-guarantees]
> *"**A KEY REQUIREMENT for stream processing applications is the ability to process each record EXACTLY ONCE, regardless of failures. WITHOUT EXACTLY-ONCE GUARANTEES, STREAM PROCESSING CAN'T BE USED IN CASES WHERE ACCURATE RESULTS ARE NEEDED.**"*
*(Full mechanism: Ch. 8. `exactly_once_beta` is what allows one transactional producer to handle many partitions — Ch. 8 §4.1.)*
***
# 14.3 Stream processing design patterns (/docs/kafka/stream-processing/stream-processing-design-patterns)
#### 3.1 Single-event processing (map/filter) [#31-single-event-processing-mapfilter]
> *"The **most basic** pattern... also known as a **map/filter pattern** because it is commonly used to **filter unnecessary events** from the stream or **transform each event.** (The term *map* is based on the map/reduce pattern in which the map stage transforms events and the reduce stage aggregates them.)"*
#### 3.2 Processing with local state [#32-processing-with-local-state]
> *"**Most stream processing applications are concerned with AGGREGATING information, especially WINDOW aggregation.**"* Example: min/max stock prices per day, moving average.
**Why local state suffices — the partitioning argument:**
**Three problems local state creates:**
*(Log-compacted changelog topics are the Ch. 1 §3.6 / Ch. 6 §7 compaction feature doing exactly the job it was designed for: "changelog-type data, where only the last update is interesting.")*
#### 3.3 Multiphase processing / repartitioning [#33-multiphase-processing--repartitioning]
> *"Local state is great if we need a **group-by** type of aggregate. **But what if we need a result that uses ALL AVAILABLE INFORMATION?**"*
**The example:** *"suppose we want to publish the **top 10 stocks each day** — the 10 stocks that gained the most from opening to closing. **Obviously, NOTHING WE DO LOCALLY ON EACH APPLICATION INSTANCE IS ENOUGH because ALL THE TOP 10 STOCKS COULD BE IN PARTITIONS ASSIGNED TO OTHER INSTANCES.**"*
> 💡 *"This type of multiphase processing is **very familiar to those who write MapReduce code**, where you often have to resort to multiple reduce phases. **If you've ever written map-reduce code, you'll remember that YOU NEEDED A SEPARATE APP FOR EACH REDUCE STEP. UNLIKE MAPREDUCE, MOST STREAM PROCESSING FRAMEWORKS ALLOW INCLUDING ALL STEPS IN A SINGLE APP, with the framework handling the details of which application instance will run each step.**"*
#### 3.4 Stream-table join (processing with external lookup) [#34-stream-table-join-processing-with-external-lookup]
**The obvious approach, and why it fails:**
**The caching dilemma:**
> *"To get good performance and availability, we need to **CACHE** the information... **Managing this cache can be challenging though — HOW DO WE PREVENT THE INFORMATION IN THE CACHE FROM GETTING STALE? If we refresh events TOO OFTEN, WE ARE STILL HAMMERING THE DATABASE, and the cache isn't helping much. If we wait TOO LONG to get new events, WE ARE DOING STREAM PROCESSING WITH STALE INFORMATION.**"*
**The resolution — CDC:**
*(Ch. 9's Debezium recommendation is exactly the tool for the CDC leg.)*
#### 3.5 Table-table join [#35-table-table-join]
> \*"There is **no reason why we can't have those materialized tables in BOTH SIDES** of the join operation.
>
> **Joining two tables is ALWAYS NONWINDOWED and joins THE CURRENT STATE of both tables at the time the operation is performed.**"\*
#### 3.6 Streaming join (windowed join) [#36-streaming-join-windowed-join]
> *"What makes a stream 'real'? ... **When we use a stream to represent a TABLE, we can IGNORE MOST OF THE HISTORY because we only care about the current state. BUT WHEN WE JOIN TWO STREAMS, WE ARE JOINING THE ENTIRE HISTORY, trying to match events in one stream with events in the other that have THE SAME KEY and HAPPENED IN THE SAME TIME WINDOWS. THIS IS WHY A STREAMING JOIN IS ALSO CALLED A WINDOWED JOIN.**"*
**The example:** search queries ⋈ clicks on search results.
> *"We want to match search queries with the results they clicked on so that we will know **which result is most popular for which query.** Obviously, we want to match results **based on the search term BUT ONLY MATCH THEM WITHIN A CERTAIN TIME WINDOW. We assume the result is clicked SECONDS AFTER the query was entered.** So we keep **a small, few-seconds-long window on each stream** and match the results from each window."*
**How Kafka Streams makes it work:**
#### 3.7 Out-of-sequence events [#37-out-of-sequence-events]
> *"a challenge **not just in stream processing but ALSO in traditional ETL systems. Out-of-sequence events happen QUITE FREQUENTLY AND EXPECTEDLY in IoT scenarios.**"*
**Three real causes:**
**The four things an application must do:**
**How frameworks support it:**
> *"Several frameworks, including **Google's Dataflow** and **Kafka Streams**, have built-in support for the notion of **event time INDEPENDENT of the processing time**... typically done by **maintaining MULTIPLE AGGREGATION WINDOWS available for update in the local state** and giving developers the ability to configure how long to keep those window aggregates available. **⚠ Of course, THE LONGER THE AGGREGATION WINDOWS ARE KEPT AVAILABLE FOR UPDATES, THE MORE MEMORY IS REQUIRED to maintain the local state.**"*
**And the compaction trick that makes updates work:**
> *"The Kafka Streams API **always writes aggregation results to result topics. Those are usually COMPACTED TOPICS, which means that ONLY THE LATEST VALUE FOR EACH KEY IS PRESERVED.** In case the results of an aggregation window need to be updated as a result of a late event, **Kafka Streams will simply WRITE A NEW RESULT for this aggregation window, WHICH WILL EFFECTIVELY REPLACE THE PREVIOUS RESULT.**"*
#### 💡 3.8 Reprocessing — two variants, one strong recommendation [#-38-reprocessing--two-variants-one-strong-recommendation]
*(This is Ch. 5 §6.4's shoe-counting warning, resolved architecturally: don't reset offsets and state — run a second app.)*
#### 3.9 Interactive queries [#39-interactive-queries]
> *"Most of the time the users of stream processing applications get the results **by reading them from an output topic. In some cases, however, IT IS DESIRABLE TO TAKE A SHORTCUT AND READ THE RESULTS FROM THE STATE STORE ITSELF.** This is common **when the result is a TABLE (e.g., the top 10 best-selling books) and the stream of results is really a STREAM OF UPDATES to this table — it is MUCH FASTER AND EASIER to just read the table directly from the stream processing application state.**"*
***
# 14.6 Stream processing use cases (/docs/kafka/stream-processing/stream-processing-use-cases)
> *"stream processing — or continuous processing — is useful in cases where we want our events to be processed **in quick order rather than wait for hours until the next batch, but also where we are NOT EXPECTING A RESPONSE TO ARRIVE IN MILLISECONDS.**"*
#### Customer service — the hotel story [#customer-service--the-hotel-story]
> *"Suppose we just reserved a room at a large hotel chain... A few minutes later, when the confirmation still hasn't arrived, we call customer service. Suppose the desk tells us: **'I don't see the order in our system, but THE BATCH JOB that loads the data from the reservation system to the hotels and the customer service desk ONLY RUNS ONCE A DAY, so please call back tomorrow. You should see the email within 2–3 business days.'** This doesn't sound like very good service, yet **WE'VE HAD THIS CONVERSATION MORE THAN ONCE WITH A LARGE HOTEL CHAIN.**"*
**What you actually want:** *"**EVERY SYSTEM in the hotel chain to get an update about a new reservation SECONDS OR MINUTES after the reservation is made** — including the customer service center, the hotel, the email system, the website... the customer service center to be able to **immediately pull up all the details about ANY of our PAST VISITS to ANY of the hotels**, and the reception desk to know that **we are a loyal customer so they can give us an upgrade.**"*
#### Internet of Things — predictive maintenance [#internet-of-things--predictive-maintenance]
> \*"A very common use case... is to try to **PREDICT WHEN PREVENTIVE MAINTENANCE IS NEEDED.** This is similar to application monitoring **but applied to HARDWARE**, and is common in **manufacturing, telecommunications (identifying faulty cellphone towers), cable TV (identifying faulty box-top devices BEFORE USERS COMPLAIN)**...
>
> The goal: **process events arriving from devices at a large scale and IDENTIFY PATTERNS THAT SIGNAL THAT A DEVICE REQUIRES MAINTENANCE.** These patterns can be **dropped packets for a switch, MORE FORCE REQUIRED TO TIGHTEN SCREWS in manufacturing, or USERS RESTARTING THE BOX MORE FREQUENTLY for cable TV.**"\*
#### Fraud detection / anomaly detection [#fraud-detection--anomaly-detection]
> *"detecting **credit card fraud, stock trading fraud, video-game cheaters, and cybersecurity risks.** In all these fields, there are large benefits to catching fraud as early as possible, so **a near real-time system that is capable of responding quickly — PERHAPS STOPPING A BAD TRANSACTION BEFORE IT IS EVEN APPROVED — is much preferred to a batch job that detects fraud THREE DAYS AFTER THE FACT, when cleanup is much more complicated.**"*
**💡 The beaconing example — a genuinely instructive one:**
> \*"In cybersecurity, there is a method known as **BEACONING. When the hacker plants malware inside the organization, it will OCCASIONALLY REACH OUTSIDE TO RECEIVE COMMANDS. IT CAN BE DIFFICULT TO DETECT THIS ACTIVITY SINCE IT CAN HAPPEN AT ANY TIME AND ANY FREQUENCY.**
>
> **Typically, NETWORKS ARE WELL DEFENDED AGAINST EXTERNAL ATTACKS BUT MORE VULNERABLE TO SOMEONE INSIDE THE ORGANIZATION REACHING OUT.**
>
> By processing the large stream of network connection events and **recognizing a PATTERN of communication as ABNORMAL (for example, detecting that THIS HOST TYPICALLY DOESN'T ACCESS THOSE SPECIFIC IPs), the security organization can be alerted early, before more harm is done.**"\*
***
# 14.1 What is stream processing? (/docs/kafka/stream-processing/stream-processing)
> *"There is **a lot of confusion** about what stream processing means. **Many definitions MIX UP IMPLEMENTATION DETAILS, PERFORMANCE REQUIREMENTS, DATA MODELS, and many other aspects of software engineering.** A similar thing has happened in the world of relational databases — **the abstract definitions of the relational model are getting forever entangled in the implementation details and specific limitations of the popular database engines.**"*
#### 1.1 A data stream = an unbounded dataset [#11-a-data-stream--an-unbounded-dataset]
> *"First and foremost, **a data stream is AN ABSTRACTION REPRESENTING AN UNBOUNDED DATASET. Unbounded means INFINITE AND EVER GROWING.** The dataset is unbounded because **over time, new records keep arriving.** This definition is used by **Google, Amazon, and pretty much everyone else.**"*
**Its universality:** *"this simple model can be used to represent **just about EVERY business activity we care to analyze** — credit card transactions, stock trades, package deliveries, network events through a switch, sensors in manufacturing equipment, emails sent, moves in a game... **The list is endless because PRETTY MUCH EVERYTHING CAN BE SEEN AS A SEQUENCE OF EVENTS.**"*
#### 1.2 The three additional attributes — each contrasted with a database table [#12-the-three-additional-attributes--each-contrasted-with-a-database-table]
**What the definition deliberately says nothing about:**
> *"neither the definition nor the attributes say anything about **the DATA contained in the events or the NUMBER OF EVENTS PER SECOND.** ... **While it is often ASSUMED that data streams are 'BIG DATA' and involve millions of events per second, THE SAME TECHNIQUES APPLY EQUALLY WELL (AND OFTEN BETTER) TO SMALLER STREAMS OF EVENTS WITH ONLY A FEW EVENTS PER SECOND OR MINUTE.**"*
#### 1.3 The three programming paradigms [#13-the-three-programming-paradigms]
> *"Stream processing refers to **the ongoing processing of one or more event streams.** Stream processing is **a PROGRAMMING PARADIGM — just like request-response and batch processing.**"*
> **The definitional boundary:** *"the definition **doesn't mandate any specific framework, API, or feature.** As long as we are **continuously reading data from an unbounded dataset, doing something to it, and emitting output**, we are doing stream processing. **BUT THE PROCESSING HAS TO BE CONTINUOUS AND ONGOING. A process that starts every day at 2:00 a.m., reads 500 records from the stream, outputs a result, and goes away DOESN'T QUITE CUT IT.**"*
***
# Data Types (/docs/postgres/data-types/data-types)
[SQL Basics](/postgres/sql-basics) used `serial`, `text`, and
`timestamptz` without much comment. Here's the fuller type system —
picking the right one avoids subtle correctness and performance bugs
down the line.
## Numeric [#numeric]
Never use `real`/`double precision` for money or anything requiring
exact decimal arithmetic — binary floating point can't represent
most decimal fractions precisely, and rounding errors compound.
`numeric(p, s)` is exact and is the right default for currency.
## Character [#character]
Unlike some databases, Postgres has no internal performance difference
between `text` and `varchar` — `varchar(n)` just adds a length check.
Default to `text` unless you have an actual business rule that caps
length.
## Boolean, date/time, UUID [#boolean-datetime-uuid]
Prefer `timestamptz` over `timestamp` almost always. `timestamp`
stores exactly what you give it with no timezone context, so the
same column can silently mix values from different timezones.
`timestamptz` normalizes everything to UTC internally and converts
on display based on the session's timezone setting — it's the
behavior most applications actually want.
## ENUM, domains, composite types [#enum-domains-composite-types]
An **enum** restricts a column to a fixed set of labels, enforced at
the type level. A **domain** is a named constraint wrapped around an
existing base type — `positive_int` behaves like `integer` everywhere
but rejects non-positive values, so you don't have to repeat a
`CHECK (value > 0)` on every table that needs it. A **composite type**
groups fields into a single reusable structure, similar to a row
type.
Next: [JSON & Arrays](/postgres/json-arrays) — for the shapes that
don't fit neatly into fixed columns.
# Data Types (/docs/postgres/data-types)
Choosing the right shape for your data, including JSONB and arrays.
# JSON & Arrays (/docs/postgres/data-types/json-arrays)
Beyond the fixed columns in [Data Types](/postgres/data-types),
Postgres has first-class support for semi-structured data — useful,
but easy to overuse.
## json vs jsonb [#json-vs-jsonb]
Almost always reach for `jsonb`, not `json`. `json` stores the
input text verbatim (preserving key order and whitespace) and has
to re-parse it on every access. `jsonb` stores a parsed binary
representation — slightly slower to write, meaningfully faster to
query, and the only one of the two that indexes support.
## Querying jsonb [#querying-jsonb]
`->` extracts a field as jsonb (keep chaining), `->>` extracts as
text (use it at the end of a chain when you need a plain value to
compare or display). `@>` checks whether the left jsonb value
contains the right one — this is what a GIN index accelerates; see
[Indexing](/postgres/indexing).
## Arrays [#arrays]
Arrays are a native type for any base type (`integer[]`, `text[]`,
even `jsonb[]`), not an add-on — useful for small, denormalized lists
that don't warrant a separate join table.
jsonb is a relief valve, not a schema strategy. It's the right tool
for genuinely variable or sparse attributes (event payloads,
third-party API responses), but a table where every column is
really jsonb loses constraints, foreign keys, and most of the
planner's ability to reason about your data. If a field is always
present and has a known type, it belongs in a real column.
That closes out the Foundations and Data Types groups. Next up:
[Heap Storage](/postgres/heap-storage) — how a row you just inserted
actually ends up on disk.
# Architecture (/docs/postgres/foundations/architecture)
With Postgres running from [Docker & Setup](/postgres/setup), it's
worth being precise about what's actually listening on the other end
of that connection before writing any SQL.
## Client / server model [#client--server-model]
Postgres is a client/server database: a single long-running server
process (the **postmaster**) listens on port 5432 and accepts
connections. Critically, the postmaster doesn't handle queries
itself — for every new connection, it **forks a dedicated backend
process** to serve that one client for the lifetime of the
connection.
This is different from thread-per-connection designs. One process
per connection is simple and gives strong isolation (one backend
crashing doesn't take down the server), but process creation is
comparatively expensive — which is why a high connection count
becomes a real cost, not just a number. That tradeoff is exactly why
connection poolers like PgBouncer exist; see
[Scaling & Pooling](/postgres/scaling).
Alongside backend processes, the postmaster also starts a handful of
permanent background processes: the **background writer** (flushes
dirty pages from shared buffers), **checkpointer** (see
[Write-Ahead Log](/postgres/wal)), **autovacuum launcher** (see
[VACUUM & Autovacuum](/postgres/vacuum)), and **WAL writer**. You can
see all of these with `ps aux | grep postgres` inside the container.
## Database cluster, databases, schemas, tables [#database-cluster-databases-schemas-tables]
The term **database cluster** in Postgres does not mean a group of
machines — it means the entire collection of databases managed by one
running Postgres server instance, all living under one data directory
on disk (`/var/lib/postgresql/data` in this repo's container).
"Cluster" here is Postgres-specific terminology and trips up anyone
coming from a distributed-systems background. One `docker compose
up` gives you one cluster, which can contain many databases. It has
nothing to do with multiple machines — that's what
[Replication](/postgres/replication) and
[Scaling & Pooling](/postgres/scaling) are about.
Inside a cluster:
* A **database** is an isolated namespace — connections attach to
exactly one database at a time, and by default databases can't
query across each other.
* A **schema** is a namespace *inside* a database — a way to group
tables without needing separate databases. Every database starts
with a `public` schema, and objects created without specifying a
schema go there by default.
* A **table** lives inside a schema, addressed as
`schema_name.table_name` (or just `table_name` if it's in your
`search_path`, which defaults to including `public`).
## OLTP, again [#oltp-again]
This process-per-connection, cluster/database/schema/table structure
exists to serve the [OLTP workload](/postgres/why-postgres) Postgres
is built for: many concurrent clients, each doing small, isolated
units of work against a shared, structured dataset.
Next: [SQL Basics](/postgres/sql-basics) — creating that structure
and putting rows into it.
# Foundations (/docs/postgres/foundations)
Get it running, understand the process model, and write your first schema.
# Docker & Setup (/docs/postgres/foundations/setup)
Same as [ClickHouse](/clickhouse/setup), everything here runs in
Docker first. The image/container/volume model is identical across
both modules — worth being precise about once.
## From image to running server [#from-image-to-running-server]
A Docker **image** is a read-only template. A **container** is a
running process created from that image, with its own writable layer
on top. A **volume** is storage that lives outside the container's
lifecycle, so data survives when the container is removed and
recreated.
The image is pulled once and cached locally. Every
`docker compose up` after that reuses it — only the container is
created and destroyed. `docker compose down` removes the container
but leaves `postgres_data` untouched; only `down -v` deletes the
volume, and with it, every database in the cluster.
## The compose file [#the-compose-file]
This repo's `postgres/compose.yaml` is intentionally minimal — one
service, one port, one named volume:
Unlike ClickHouse's dual HTTP/native-protocol ports, Postgres exposes
a single port: `5432`, the Postgres wire protocol. Every client —
`psql`, drivers, connection poolers — speaks it.
## Environment & credentials [#environment--credentials]
`.env` supplies the database name, user, and password Postgres
bootstraps on first start — first start only. Changing these after
the volume already has data does nothing until you wipe the volume,
because the role and database were already created on disk.
The official Postgres image only runs its first-start initialization
scripts when `/var/lib/postgresql/data` is empty. If you change
`.env` after the volume already has data, nothing happens — the
`admin` role and `learning` database already exist from the first
run. Wipe the volume (`docker compose down -v`) to re-init.
## Everyday commands [#everyday-commands]
## Verifying it works [#verifying-it-works]
Once the container is up, connect and run a few sanity commands. The
first two are plain SQL; the last two are `psql` **meta-commands**
(client-side shortcuts, not SQL — they start with a backslash and
aren't sent to the server as queries):
With the server running and reachable, the next lesson,
[Architecture](/postgres/architecture), covers what's actually
happening on the other end of that connection.
# SQL Basics (/docs/postgres/foundations/sql-basics)
With a database and schema in place from
[Architecture](/postgres/architecture), this covers the SQL you'll use
constantly: defining tables and moving rows in and out of them.
## Creating structure [#creating-structure]
A few constraints worth knowing by default:
* `serial` allocates an auto-incrementing integer backed by a
sequence — the modern equivalent, preferred in new schemas, is
`GENERATED ALWAYS AS IDENTITY`.
* `NOT NULL` and `UNIQUE` are enforced on every write, not advisory.
* `DEFAULT now()` fills the column automatically if the INSERT
doesn't specify it.
`ALTER TABLE` and `DROP TABLE` work as you'd expect — Postgres will
enforce constraints on `ALTER TABLE ... ADD CONSTRAINT` against
existing data, and `DROP TABLE` takes an `ACCESS EXCLUSIVE` lock (see
[Locks & Deadlocks](/postgres/locks)).
## Moving rows [#moving-rows]
UPDATE and DELETE here are ordinary, fast, row-targeted operations —
not batch rewrites.
This is where Postgres's SQL diverges most visibly from
[ClickHouse's dialect](/clickhouse/sql-basics). ClickHouse treats
UPDATE/DELETE as heavyweight asynchronous mutations and doesn't
enforce `NOT NULL` or foreign keys by default, because it's
optimized for append-heavy analytical ingestion, not row-level
editing. Postgres enforces every constraint immediately and expects
UPDATE/DELETE to be routine — that's the OLTP contract from
[Why PostgreSQL](/postgres/why-postgres).
Next: [Data Types](/postgres/data-types) — picking the right column
types for what you're storing.
# EXPLAIN & EXPLAIN ANALYZE (/docs/postgres/indexing-query-planning/explain)
The [Query Planner](/postgres/query-planner) picks a strategy silently —
`EXPLAIN` is how you see which one it actually chose, and whether its
estimates matched reality.
## EXPLAIN vs EXPLAIN ANALYZE [#explain-vs-explain-analyze]
`EXPLAIN` alone shows the *planned* execution without running the query:
`EXPLAIN ANALYZE` actually **executes** the query and adds real timing and
row counts alongside the estimates:
`EXPLAIN ANALYZE` really runs the query — including a real `DELETE`,
`UPDATE`, or `INSERT`. Wrap it in a transaction you roll back if you're
analyzing a write: `BEGIN; EXPLAIN ANALYZE UPDATE ...; ROLLBACK;`.
Add `BUFFERS` to see actual page I/O — how much came from shared buffers
(cache) versus disk, which is often more useful than timing alone for
diagnosing a slow query:
`hit` is a page already in shared memory; `read` came from disk (or the OS
page cache). A query that's all `hit` on a warm cache but still slow points
away from I/O and toward CPU-bound work — sorting, hashing, or function
evaluation.
## Reading the plan tree [#reading-the-plan-tree]
Plans nest: the innermost, most-indented lines execute first, feeding rows
up to the operations above them. For a join:
Read this bottom-up: scan `accounts` filtering `balance > 1000`, build a
hash table from the result, scan `orders` in full, then probe the hash
table for each `orders` row to find matches.
## The one signal that matters most [#the-one-signal-that-matters-most]
Every line has both an **estimated** row count (`rows=310` in the cost
line) and an **actual** row count (`rows=298` in the actual-time line).
When these diverge wildly — the planner expected 10 rows and got 100,000 —
every decision built on top of that estimate (which join strategy, which
scan type) is likely wrong too, even though each individual choice looked
reasonable given the (bad) numbers it had.
A large estimate/actual mismatch almost always traces back to stale or
insufficient statistics — see [Query Planner](/postgres/query-planner).
Run `ANALYZE` on the table first before assuming the query itself needs
rewriting.
With planning and diagnosis covered, the next group of lessons moves back
to SQL itself — starting with
[Joins, CTEs & Recursion](/postgres/joins-ctes).
# Indexing & Query Planning (/docs/postgres/indexing-query-planning)
Making the planner choose the fast path.
# Indexing (/docs/postgres/indexing-query-planning/indexing)
Without an index, every query is a sequential scan — Postgres reads every
page of a table, checking every row against your `WHERE` clause. An index is
a separate, ordered structure that lets Postgres jump straight to matching
rows instead. Postgres ships with six index types, each suited to a
different kind of query.
## B-Tree — the default [#b-tree--the-default]
`CREATE INDEX` with no type specified builds a B-Tree, and it's the right
choice for the large majority of columns: equality, ranges, sorting, and
`IS NULL` all use it.
## Hash — equality only [#hash--equality-only]
A Hash index supports only `=` comparisons, never ranges or sorting. It's
rarely worth choosing over a B-Tree in practice — B-Tree equality lookups
are already fast, and Hash indexes were unlogged (not crash-safe) before
Postgres 10. Reach for it only when you're certain the column is queried
exclusively with `=` and the index would be large enough for the smaller
per-entry size to matter.
## GIN — composite values [#gin--composite-values]
**Generalized Inverted Index.** Built for columns that hold *multiple*
values per row — arrays, `jsonb`, full-text search vectors — where you need
to ask "does this row contain X?" A GIN index maps each individual element
to the rows containing it:
See [JSON & Arrays](/postgres/json-arrays) for the operators GIN accelerates
on those types.
## GiST — geometric and "nearest" queries [#gist--geometric-and-nearest-queries]
**Generalized Search Tree.** Supports queries with no strict ordering —
geometric containment/overlap, nearest-neighbor (`ORDER BY point <->
target`), and range-type exclusion constraints (e.g. "no two bookings for
the same room can overlap").
## BRIN — huge, naturally-ordered tables [#brin--huge-naturally-ordered-tables]
**Block Range Index.** Instead of indexing every row, BRIN stores the
min/max value per *range of pages* (128 pages by default). It's tiny — often
a few hundred KB even for a billion-row table — and works well when a
column correlates with physical insertion order, like a timestamp on an
append-only log table:
BRIN only helps if the column is roughly sorted on disk. On a table where
rows are updated out of order (moving a "recent" row to a new physical
page), BRIN's per-range min/max becomes wide and useless. B-Tree is still
the safer default unless you've confirmed the correlation.
## SP-GiST — space-partitioned data [#sp-gist--space-partitioned-data]
**Space-Partitioned GiST.** For data with a natural, non-balanced tree
structure — IP address ranges, phone number prefixes, quadtree-style
geometric data. Niche, but the right tool when your data actually has that
shape.
## Making an index more useful [#making-an-index-more-useful]
* **Partial index** — index only the rows you actually query:
`CREATE INDEX idx_active ON users (email) WHERE active = true;` — smaller
and faster than indexing every row when most queries filter on `active`.
* **Covering index** — add extra columns with `INCLUDE` so an index-only
scan can satisfy a query without touching the heap at all:
`CREATE INDEX idx_covering ON accounts (id) INCLUDE (balance);`
* **Expression index** — index the *result* of an expression, not the raw
column: `CREATE INDEX idx_lower_email ON users (lower(email));` lets a
case-insensitive lookup (`WHERE lower(email) = 'x@example.com'`) use an
index at all.
`CREATE INDEX` takes a `SHARE` lock that blocks writes for the duration of
the build. On a live table, use `CREATE INDEX CONCURRENTLY` instead — it
takes roughly twice as long and can't run inside a transaction block, but
never blocks reads or writes.
Having the right index doesn't guarantee Postgres will use it — that
decision belongs to the planner, covered next in
[Query Planner](/postgres/query-planner).
# Query Planner (/docs/postgres/indexing-query-planning/query-planner)
Having an index from the [Indexing](/postgres/indexing) lesson doesn't
guarantee Postgres uses it. Every query is handed to the **planner**, which
generates several candidate execution strategies, estimates the cost of
each, and picks the cheapest one it can find — not necessarily the fastest
one that exists.
The planner never looks at your actual rows to make its choice — only at
statistics gathered in advance and the cost constants below. That's worth
keeping in mind whenever a plan looks "wrong": it may be a perfectly
rational decision given stale or incomplete information.
## The cost model [#the-cost-model]
Costs are abstract units, not milliseconds, built from a handful of tunable
constants:
`random_page_cost` defaulting to 4x `seq_page_cost` reflects spinning-disk
seek latency. On SSD-backed storage — which is most production deployments
today — random reads are much closer in cost to sequential ones, and it's
common to tune `random_page_cost` down to `1.1`, which makes the planner
choose index scans more readily.
## Statistics [#statistics]
The planner's cost estimates are only as good as its statistics about the
actual data — row counts, most-common values, and value distribution per
column — collected by `ANALYZE` (which autovacuum also runs automatically):
Stale statistics are one of the most common causes of a "why did the
planner pick a terrible plan" incident — a bulk load or a big `DELETE`
can shift the data enough that old estimates are simply wrong until the
next `ANALYZE` runs. Running `ANALYZE` manually after a large batch
operation is often the fix.
## Scan strategies [#scan-strategies]
For a single table, the planner chooses between:
* **Sequential scan** — read every page in order. Cheapest when a large
fraction of the table matches, since it avoids random I/O entirely.
* **Index scan** — walk the index, then fetch each matching row from the
heap. Good when a small fraction of rows match.
* **Bitmap heap scan** — walk the index to build an in-memory bitmap of
matching *pages*, then read those pages in physical order. A middle
ground: it beats a plain index scan when enough rows match that jumping
around the heap row-by-row would thrash, but still beats a full
sequential scan.
## Join strategies [#join-strategies]
For joining two tables, the planner picks from:
* **Nested loop** — for each row in the outer table, scan (or index-probe)
the inner table. Best when one side is small or well-indexed on the join
column.
* **Hash join** — build an in-memory hash table from the smaller side, then
stream the larger side through it. Good for large, unsorted inputs with
no useful index.
* **Merge join** — if both inputs are already sorted on the join key (or
the planner sorts them first), walk both in order simultaneously. Good
for large, pre-sorted inputs.
The planner estimates the cost of every viable combination of scan and join
strategy for a query and picks the cheapest total. You rarely need to force
a particular plan — but you do need to be able to read the one it chose,
which is the whole subject of the next lesson,
[EXPLAIN & EXPLAIN ANALYZE](/postgres/explain).
# Production Failure Scenarios (/docs/postgres/monitoring-performance/failure-scenarios)
Every earlier lesson explains a mechanism. This one is about what happens
when those mechanisms are neglected — the handful of failure modes that
account for most Postgres production incidents.
## Table bloat from a stalled autovacuum [#table-bloat-from-a-stalled-autovacuum]
MVCC means every `UPDATE`/`DELETE` leaves the old row version in place
until [`VACUUM`](/postgres/vacuum) reclaims it. If autovacuum can't keep
up — a long-running transaction holding back the cleanup horizon, or
autovacuum simply undertuned for the write rate — dead tuples accumulate.
Tables and indexes bloat, sequential scans get slower (more dead rows to
skip over per live row), and eventually disk fills.
## Transaction ID wraparound [#transaction-id-wraparound]
Postgres transaction IDs are 32-bit and MVCC visibility is defined in terms
of comparing them, which only works if "old" and "new" stay well-ordered —
so IDs are compared cyclically, and rows need to be periodically **frozen**
(marked as "always visible") so old IDs can safely be reused. If autovacuum
can't freeze fast enough — same root cause as bloat, usually a stuck
transaction or disabled autovacuum — Postgres eventually forces the issue:
past a warning threshold it logs aggressively, and past the hard limit it
**shuts down and refuses new transactions** rather than risk data
corruption from ID wraparound. This is one of the few Postgres failure
modes that is a full outage, not just degraded performance.
Wraparound is almost always preceded by weeks of warnings in the logs
("database is not accepting commands to avoid wraparound data loss") —
it's a slow-motion incident that becomes a sudden one only because
nobody was watching `age(datfrozenxid)`.
## Lock storms [#lock-storms]
A single query holding a lock longer than expected — an `ALTER TABLE`
that takes an `ACCESS EXCLUSIVE` lock, or an `idle in transaction`
connection sitting on a row lock — can cascade: every subsequent query
touching that table queues up behind it. `pg_stat_activity` fills with
backends in `state = active` but making no progress, all waiting on the
same [lock](/postgres/locks). From the outside this looks identical to the
database being "down," even though nothing has crashed.
## Replication lag growing unbounded [#replication-lag-growing-unbounded]
A [replica](/postgres/replication) that can't keep up — undersized
hardware, a network blip, or a heavy query holding a snapshot open on the
replica — falls further behind with every write on the primary. Left
unaddressed, reads against that replica become increasingly stale, and if
the primary's WAL retention runs out before the replica catches up, the
replica can no longer resume and needs to be rebuilt from a fresh base
backup.
## Connection exhaustion [#connection-exhaustion]
Every backend is a process holding real memory. Without
[pooling](/postgres/scaling), each new app instance opening its own
connection pool adds up fast, and once `max_connections` is hit, new
clients get `FATAL: too many connections` — including the operational
tooling you'd want to use to diagnose the problem.
Reserve a few connections for exactly this situation:
`superuser_reserved_connections` keeps a small headroom so an admin can
still connect and run diagnostics when the pool is otherwise exhausted.
## OOM from a runaway query [#oom-from-a-runaway-query]
As covered in [Performance Tuning](/postgres/performance-tuning),
[`work_mem`](/postgres/performance-tuning) is per-operation, not
per-query. A handful of concurrent queries each doing several large sorts
or hash joins can multiply past available RAM, and the Linux OOM killer
doesn't negotiate — it picks a process (often the postmaster or a busy
backend) and kills it, which can crash the whole instance and trigger
[crash recovery](/postgres/wal) on restart.
## The pattern underneath all of these [#the-pattern-underneath-all-of-these]
Every scenario above is a mechanism from an earlier lesson, left
unmonitored past the point where it self-corrects: vacuum falling behind,
a lock held too long, a replica falling behind, connections piling up,
memory multiplying. None of them are exotic — they're the ordinary
mechanisms this module covers, given enough time and enough neglect. The
[Monitoring](/postgres/monitoring) lesson's catalog views are exactly what
catches each of these while they're still a warning in a dashboard, not
yet a page at 3am. That's the real point of this whole module: understand
what Postgres is doing internally, and production failures stop being
mysterious — they're just the mechanism you already know, showing up
somewhere you weren't looking.
That closes out the module. Back to the [PostgreSQL overview](/postgres)
for the full lesson list, from [why Postgres exists](/postgres/why-postgres)
through everything covered here.
# Monitoring & Performance (/docs/postgres/monitoring-performance)
Watching it run, tuning it, and what breaks in production.
# Monitoring (/docs/postgres/monitoring-performance/monitoring)
Postgres exposes almost everything it's doing through **catalog views** —
plain tables you `SELECT` from, no separate tooling required to get
started.
## pg\_stat\_activity — what's running right now [#pg_stat_activity--whats-running-right-now]
This is the first table to check when something feels stuck: `state`
(`active`, `idle`, `idle in transaction`), how long a query has been
running, and — if it's waiting — what it's waiting on (see
[Locks](/postgres/locks) for `wait_event_type = 'Lock'`). Found a runaway
query?
`idle in transaction` is a special kind of dangerous: a connection sitting
in an open transaction, doing nothing, still holds whatever locks and
MVCC snapshot it acquired. It's a common cause of both blocked queries
and — because Postgres can't clean up rows newer than the oldest open
snapshot — [table bloat](/postgres/vacuum).
## pg\_stat\_database and pg\_locks [#pg_stat_database-and-pg_locks]
`pg_stat_database` gives per-database counters — cache hit ratio,
transactions committed/rolled back, deadlocks — good for a single "is this
database healthy" glance:
`pg_locks` is the raw lock table underneath `pg_stat_activity`'s
`wait_event` column — join it to itself to see exactly which backend is
blocking which (see [Locks & Deadlocks](/postgres/locks) for the full
query).
## pg\_stat\_user\_tables — the missing-index signal [#pg_stat_user_tables--the-missing-index-signal]
Per-table access patterns are one of the most actionable views in the
whole catalog:
A table with a high `seq_scan` count and a large `n_live_tup` (row count)
but a low `idx_scan` count is a table the planner keeps deciding to scan
sequentially — either it's genuinely small enough that a scan is cheaper,
or it's missing an index that would make the [planner](/postgres/query-planner)
choose differently. `n_dead_tup` climbing steadily is the same bloat signal
covered in [VACUUM & Autovacuum](/postgres/vacuum).
## pg\_stat\_io [#pg_stat_io]
Postgres 17 unified what used to be scattered across several views into
one: I/O broken down by backend type, I/O object (relation, temp file, WAL),
and operation (read/write/extend), including hits vs actual reads.
## Logs [#logs]
Catalog views reset on restart and only show recent activity; logs are the
durable record. The setting that matters most day-to-day:
Set this too low (or to 0, logging everything) on a busy database and the
log itself becomes a performance and disk problem. Start high (a few
hundred ms) and lower it only while actively hunting a specific slow
query.
## Prometheus & Grafana [#prometheus--grafana]
For dashboards and alerting rather than ad-hoc queries, `postgres_exporter`
runs many of the queries above on a schedule and exposes them as
Prometheus metrics.
The same three views above — `pg_stat_activity`, `pg_stat_user_tables`, and
`pg_stat_database` — cover the majority of exporter dashboards you'll find;
learning to read them directly in `psql` first makes the dashboards make
sense rather than just being colored numbers. Next: turning what you can
now see into faster queries — [Performance Tuning](/postgres/performance-tuning).
# Performance Tuning (/docs/postgres/monitoring-performance/performance-tuning)
[Monitoring](/postgres/monitoring) tells you something is slow. This lesson
is about the levers for making it fast — finding the specific slow query,
skipping unnecessary planning work, and sizing the memory Postgres has to
work with.
## pg\_stat\_statements [#pg_stat_statements]
`pg_stat_activity` only shows what's running *right now*. `pg_stat_statements`
aggregates every query the server has run, grouped by query shape (literal
values normalized away), ranked by total time, calls, or average latency:
This is almost always where real tuning starts: not by guessing, but by
asking Postgres which query shape has actually cost the most cumulative
time. A query that runs in 2ms but fires 50,000 times an hour can outrank
one slow 2-second report that runs twice a day.
## Prepared statements [#prepared-statements]
A normal query gets parsed and planned fresh every single time it's sent.
A **prepared statement** does that work once and reuses the plan on
subsequent executions with different parameter values:
Most drivers do this transparently for parameterized queries. It matters
most for simple, extremely frequent queries, where planning time is a real
fraction of total execution time — for a query that's mostly waiting on
I/O or scanning millions of rows, re-planning cost is noise.
Postgres switches from a plan tailored to each call's specific parameter
values to one "generic" plan (safe for any value) after enough
executions of the same prepared statement. If a query's ideal plan
genuinely depends on the parameter value (e.g. a very selective vs. very
common value for an indexed column), this generic-plan switch can make a
previously-fast prepared query suddenly slow. This is also why
[connection pooling](/postgres/scaling) in transaction mode complicates
prepared statements — the backend a client lands on next may not be the
one that prepared the statement.
## Key memory settings [#key-memory-settings]
* **`shared_buffers`** — Postgres's own cache of table/index pages,
separate from the OS page cache. A common starting point is \~25% of
system RAM; higher isn't always better, since the OS cache backs up
whatever doesn't fit here anyway.
* **`work_mem`** — memory available for *each* sort, hash join, or hash
aggregate operation before it spills to disk.
* **`maintenance_work_mem`** — a separate, usually much larger budget for
[`VACUUM`](/postgres/vacuum) and index builds, which benefit from more
memory but run far less often than everyday queries.
* **`effective_cache_size`** — doesn't allocate anything; it just tells the
[planner](/postgres/query-planner) roughly how much memory is available
for caching across shared\_buffers *and* the OS, so it can judge whether an
index scan's random I/O is likely to hit cache or hit disk.
`work_mem` is per sort/hash operation, not per query and not per
connection — a single complex query with three sorts and two hash joins
can use `5 × work_mem`, and that multiplies again by however many such
queries run concurrently. A `work_mem` that looks conservative at 100
connections can OOM a server at 500. Size it against worst-case
concurrency, not the happy path — this is exactly the mechanism behind
the OOM failure mode in [Production Failure
Scenarios](/postgres/failure-scenarios).
Tuning, in order: find the actual slow query with `pg_stat_statements`,
check its plan with [`EXPLAIN ANALYZE`](/postgres/explain), fix the
`work_mem`/index gap it reveals, then size the memory settings above for
the concurrency you actually run — not before, and not the other way
around. Last: what happens when tuning doesn't happen in time —
[Production Failure Scenarios](/postgres/failure-scenarios).
# Replication & Scaling (/docs/postgres/replication-scaling)
Running it across machines, safely, over time.
# Replication (/docs/postgres/replication-scaling/replication)
Every write to Postgres already produces a byte-for-byte record of what
changed: the [write-ahead log](/postgres/wal). Replication is what happens
when you ship that log to another server and have it replay the same
changes. Nothing about replication requires a second "replication engine" —
it's WAL, sent somewhere else, applied in order.
## Physical (streaming) replication [#physical-streaming-replication]
A **replica** connects to the **primary** and streams WAL as it's
generated, applying each record to its own copy of the data files. Because
it's replaying the exact same WAL the primary wrote, the replica ends up
byte-for-byte identical — same tables, same indexes, same bloat, same
everything.
A replica is read-only by default — it will reject writes with `ERROR:
cannot execute INSERT in a read-only transaction`. This is what makes read
replicas useful: point read-heavy traffic at them without any risk of them
diverging from the primary.
### Synchronous vs asynchronous [#synchronous-vs-asynchronous]
By default, replication is **asynchronous** — the primary commits a
transaction and returns success to the client without waiting for any
replica to receive it. This means a crashed primary can lose the last few
transactions that never made it to a replica.
**Synchronous replication** (`synchronous_standby_names`) makes the primary
wait for at least one replica to confirm it has *received* the WAL before
the client's `COMMIT` returns. This trades latency (every commit now waits
on a network round-trip) for a durability guarantee: an acknowledged commit
is not lost even if the primary dies immediately after.
Synchronous replication protects against data loss, not availability. If
the only synchronous replica goes down, the primary blocks on every
commit until it comes back (or you reconfigure). Most production setups
use `ANY 1 (replica_a, replica_b)` so one healthy replica out of several
is enough.
## Replication slots [#replication-slots]
Without a slot, a lagging replica is the primary's problem to solve on its
own — it will happily recycle old WAL segments once `wal_keep_size` is
exceeded, and a replica that falls behind that point can no longer catch
up; it has to be rebuilt from scratch.
A **replication slot** flips that: the primary tracks exactly how far each
slot's consumer has confirmed receiving WAL, and refuses to remove any
segment a slot still needs — no matter how far behind it falls.
A slot is a promise the primary keeps forever unless you break it. If a
replica using a slot dies and nobody drops the slot, the primary keeps
every WAL segment since that replica's last checkpoint — indefinitely.
On a busy database this fills the disk in hours, not days. Monitor
`pg_replication_slots` for slots where `active = false`.
## Logical replication [#logical-replication]
Physical replication ships raw bytes and requires an identical replica —
same major version, same full copy of every database in the cluster.
**Logical replication** instead decodes WAL back into row-level changes
(`INSERT`/`UPDATE`/`DELETE` on specific tables) and replays those as SQL,
which unlocks things physical replication can't do:
* Replicate a subset of tables, not the whole cluster.
* Replicate into a database with extra tables, columns, or a different
major Postgres version — useful for near-zero-downtime upgrades.
* Feed changes into non-Postgres consumers (this is what tools like
Debezium build on).
Logical replication needs `wal_level = logical` (a superset of `replica`)
because it has to decode enough information from WAL to reconstruct each
row's before/after values, not just apply raw page changes.
Logical replication doesn't ship DDL — if you `ALTER TABLE` on the
publisher, you generally have to apply the same change on the subscriber
yourself. This is a common source of "why did replication just stop"
incidents.
This repo's local `postgres/compose.yaml` runs a single instance, so
there's no replica to point at — but everything above works the same
whether the replica is a second container on your laptop or a separate
machine in another region. The concepts, not the container count, are what
matter here. Next: what to do when a single primary and its replicas
aren't enough — [Scaling & Pooling](/postgres/scaling).
# Scaling & Pooling (/docs/postgres/replication-scaling/scaling)
[Replication](/postgres/replication) gives you copies of your data. Scaling
is about deciding what to do with them — and what to do once copies alone
aren't enough.
## Vertical scaling, and its ceiling [#vertical-scaling-and-its-ceiling]
The simplest lever is a bigger machine: more RAM (more room for
[`shared_buffers`](/postgres/performance-tuning) and OS page cache), faster
disks, more CPU cores for parallel query execution. It requires no
architecture changes and it's usually the right first move. It also has a
hard ceiling — at some point there isn't a bigger single machine to buy,
and a single primary can only accept writes as fast as one WAL stream can
be fsynced to disk.
## Read replicas [#read-replicas]
If your workload is read-heavy (most web apps are), [streaming
replicas](/postgres/replication) let you scale reads roughly linearly by
adding more of them. Writes still go through the primary — replicas don't
help write throughput at all, and application code has to be aware of
**replication lag**: a read immediately after a write, on a replica, can
return stale data if it hasn't caught up yet.
"Read your own writes" bugs are the most common replica footgun: a user
submits a form, the app writes to the primary, redirects to a page that
reads from a replica, and the just-created row isn't there yet. Route
read-after-write paths back to the primary, or read from the same
connection you wrote with.
## Connection pooling [#connection-pooling]
Postgres backends are not free — each connection is a full OS process with
its own memory (including [`work_mem`](/postgres/performance-tuning)
headroom), and the default `max_connections` (100) is reached faster than
most people expect once an app has a few dozen instances each holding their
own pool.
**PgBouncer** sits between clients and Postgres and multiplexes many client
connections onto far fewer real backend connections:
Pooling modes trade compatibility for reuse:
* **Session** — one client connection maps to one backend for the whole
session. Fully compatible (session-level state like `SET` and advisory
locks works normally), least reuse.
* **Transaction** — a backend is only held for the duration of one
transaction, then returned to the pool. Much better reuse, but session
state (prepared statements outside the pooler's support, `SET` outside a
transaction) doesn't survive across transactions.
* **Statement** — a backend is returned after every single statement. Best
reuse, breaks multi-statement transactions entirely; rarely used.
Transaction pooling is the common default, but it silently breaks
session-scoped features. Advisory locks (see [Locks &
Deadlocks](/postgres/locks)) and `SET` statements meant to persist across
a session are the two things most likely to bite you the first time you
put PgBouncer in front of an app that wasn't written with it in mind.
## Sharding — the last resort [#sharding--the-last-resort]
Sometimes writes themselves need to scale past what one primary can do.
**Sharding** splits data across multiple independent Postgres instances by
some key (e.g. `customer_id`), so each shard only holds and serves a slice
of the data. Postgres has no built-in sharding; it's done either at the
application layer (your code decides which shard to query) or with an
extension like **Citus**, which distributes tables and rewrites queries
across a cluster of Postgres nodes transparently.
Sharding is powerful and also the most operationally expensive option on
this list — cross-shard joins, transactions, and rebalancing all become
hard problems. Reach for it only after vertical scaling, read replicas, and
pooling are genuinely exhausted, not as a first move.
## High availability [#high-availability]
None of the above automatically handles the primary failing. Tools like
**Patroni** watch a primary's health, hold leader election via a
distributed store (etcd/Consul/ZooKeeper), and promote a replica to
primary automatically on failure — the same "one primary, several standbys,
automatic failover" shape as most other database HA systems.
This repo's `postgres/compose.yaml` runs a single node for learning
purposes; none of these tools are wired up locally, but the vertical →
replicas → pooling → sharding progression is the order production systems
actually climb it in. Next: [Users, Roles &
Security](/postgres/security).
# Configuration & Extensions (/docs/postgres/security-ops/configuration)
Two files, sitting next to each other in the data directory, control almost
everything about how a Postgres server behaves and who it will talk to.
## postgresql.conf [#postgresqlconf]
Server-wide settings live here: memory ([`shared_buffers`,
`work_mem`](/postgres/performance-tuning)), WAL behavior
([`wal_level`](/postgres/replication)), logging, connection limits, and
hundreds more.
Not every setting takes effect the same way once changed:
* Some (`work_mem`, `log_min_duration_statement`) apply on the next
`SIGHUP` — `SELECT pg_reload_conf();` or `docker exec postgres pg_ctl
reload`, no downtime.
* Some (`shared_buffers`, `max_connections`) are fixed at process startup
and need a full restart, because they size memory that's allocated once
when the postmaster starts.
`ALTER SYSTEM` is the SQL-native way to change settings without editing the
file by hand — it writes to `postgresql.auto.conf`, which is read after
(and overrides) `postgresql.conf`.
`ALTER SYSTEM SET` doesn't apply the change — it only writes it. Forgetting
the follow-up `pg_reload_conf()` (or restart, for startup-only settings)
is one of the most common "why didn't my setting take" moments.
## pg\_hba.conf [#pg_hbaconf]
This is the file that actually enforces the authentication methods
discussed in [Users, Roles & Security](/postgres/security) — "hba" stands
for host-based authentication. Each line is a rule matched top-to-bottom
against connection type, database, user, and source address; the first
match wins.
* `trust` — no password at all. Fine for `local` connections in a
throwaway container, never for anything reachable from a network.
* `password` — plaintext over the wire (only safe combined with
[`sslmode=require` or stronger](/postgres/security)).
* `scram-sha-256` — the modern default: a salted-challenge password
exchange that never sends the password itself, even over an unencrypted
connection.
* `cert` — the client presents a TLS client certificate instead of a
password.
A trailing `reject` rule for `0.0.0.0/0` is worth keeping explicit even
though "no matching rule" also rejects a connection by default — it
documents the intent, and it fails loudly (a clear "no pg\_hba.conf entry"
error) instead of silently depending on rule ordering someone might
reorder later.
## Extensions [#extensions]
Postgres ships a small core and leans on extensions for almost everything
beyond it. `CREATE EXTENSION` loads one into the current database:
Some extensions (like `pg_stat_statements`) also need to be added to
`shared_preload_libraries` in `postgresql.conf` and the server restarted,
because they hook into query execution from the moment the server starts —
`CREATE EXTENSION` alone isn't enough for those.
## Major-version upgrades [#major-version-upgrades]
Changing a Docker image tag from `postgres:16` to `postgres:17` does not
upgrade a running cluster — the on-disk format isn't guaranteed compatible
across major versions, so the new binary will refuse to start against old
data files. A real upgrade uses `pg_upgrade` (which can do it in place,
fast, via hard links) or a `pg_dump`/`pg_restore` round trip, or — the
zero-downtime option — [logical replication](/postgres/replication) into a
new cluster running the target version, followed by a cutover.
Configuration decides how the server runs and who's allowed to touch it.
Next: [Monitoring](/postgres/monitoring) covers watching what it's actually
doing while it runs.
# Security & Ops (/docs/postgres/security-ops)
Users, permissions, and day-to-day configuration.
# Users, Roles & Security (/docs/postgres/security-ops/security)
Postgres has no separate concept of a "user" — a **role** that has been
granted `LOGIN` privilege *is* a user. Everything else (permissions,
row-level security, transport encryption) builds on that one primitive.
## Roles [#roles]
Roles can own objects, hold privileges directly, and be members of other
roles (which is how Postgres does "groups" — a group is just a role other
roles inherit from).
## Privileges [#privileges]
Privileges are granted on specific objects — databases, schemas, tables,
even individual columns — not globally by default.
New tables don't automatically inherit privileges you granted on
existing ones — `GRANT SELECT ON orders TO app_user` only covers
`orders`. Use `ALTER DEFAULT PRIVILEGES` if you want privileges to apply
to tables created later, or every new table becomes a silent access gap.
## Row-Level Security [#row-level-security]
Table-level `GRANT`s control which tables a role can touch. **Row-Level
Security (RLS)** goes further and controls *which rows* — the same query,
run by two different roles, can see two different sets of rows from the
same table.
Once RLS is enabled, every query against `orders` — including ones the
application didn't anticipate, like an ad-hoc `SELECT *` from a debugging
session — is filtered by the policy. This is the standard way to build
multi-tenant Postgres safely: application bugs that forget a `WHERE
tenant_id = ...` clause fail closed instead of leaking another tenant's
rows.
RLS policies don't apply to the table owner or superusers by default
(`FORCE ROW LEVEL SECURITY` changes that). This is deliberate — admin and
migration tooling usually needs unfiltered access — but it means testing
RLS as a superuser will look like it isn't working at all.
Compare this to ClickHouse's [row
policies](/clickhouse/row-policies-quotas): the mechanism (a `USING`
predicate scoped to a table and role) is nearly identical — RLS is one of
the places the two databases' SQL genuinely converges, since both borrow
the same standard SQL feature.
## SSL/TLS [#ssltls]
By default, a local connection like this module's `docker exec ... psql`
never leaves the container, so encryption isn't in play. Any connection
over a network should set `sslmode` on the client:
`sslmode` has a range from `disable` up through `verify-full` (which also
validates the server's certificate matches the hostname, closing off
man-in-the-middle attacks). `require` alone encrypts the connection but
doesn't verify who's on the other end of it — enough to stop passive
eavesdropping, not enough to stop impersonation.
## Authentication methods [#authentication-methods]
Whether a connection is accepted at all — and which method it must
authenticate with — isn't decided here; it's decided per-connection by
matching rules in `pg_hba.conf`, covered next in [Configuration &
Extensions](/postgres/configuration).
# SQL & Querying (/docs/postgres/sql-querying)
Joins, windows, views, and partitioning.
# Joins, CTEs & Recursion (/docs/postgres/sql-querying/joins-ctes)
[SQL Basics](/postgres/sql-basics) covered `SELECT` against a single
table. Almost nothing interesting stays in one table for long —
joins combine rows across tables, `GROUP BY`/`HAVING` collapse them
into summaries, and CTEs (`WITH ...`) let you name intermediate
results instead of nesting subqueries five levels deep.
## Join types [#join-types]
Sample schema for this lesson: `employees(id, name, manager_id,
department_id)` and `departments(id, name)`.
`INNER JOIN` is the default `JOIN` — rows drop out entirely if either
side has no match. `LEFT JOIN` (and its mirror, `RIGHT JOIN`) keeps
every row from the "kept" side and fills unmatched columns with
`NULL`. `FULL JOIN` keeps everything from both sides. In practice
`RIGHT JOIN` is rare — it's almost always written as a `LEFT JOIN`
with the tables swapped instead, since that reads more naturally
left-to-right.
## GROUP BY and HAVING [#group-by-and-having]
`GROUP BY` collapses matching rows into one row per group; `HAVING`
filters groups the same way `WHERE` filters rows, but it runs
*after* aggregation, so it can reference aggregate functions where
`WHERE` can't:
`WHERE` filters rows before grouping happens; `HAVING` filters the
groups produced by that grouping. `WHERE count(*) > 3` is a syntax
error for exactly this reason — at the point `WHERE` runs, `count(*)`
doesn't exist yet.
## CTEs: naming intermediate results [#ctes-naming-intermediate-results]
A CTE (`WITH ... AS (...)`) gives a subquery a name you can reference
later in the statement, one or more times:
This is purely a readability tool for non-recursive CTEs — every one
of them can be rewritten as a nested subquery. The value is that a
chain of `WITH` clauses reads top-to-bottom like a sequence of steps,
instead of nesting inside-out.
Before Postgres 12, every CTE was an **optimization fence**: the
planner executed it in isolation and materialized the full result
before the outer query ran, so a filter in the outer query couldn't
push down into it. Postgres 12+ inlines non-recursive CTEs by
default, treating `WITH` more like a named subquery the planner is
free to rewrite — filters and joins can push down into it just like
a regular subquery. If you specifically need the old fencing
behavior (say, to guarantee a side-effecting or expensive
computation runs exactly once), force it explicitly with `WITH x AS
MATERIALIZED (...)`. Recursive CTEs can't be inlined and remain
fenced either way.
## Recursive CTEs: walking a tree [#recursive-ctes-walking-a-tree]
`WITH RECURSIVE` is the standard way to walk hierarchical data — org
charts, category trees, bill-of-materials — in plain SQL, without a
procedural loop. It has two parts joined by `UNION ALL`: an **anchor**
that seeds the starting rows, and a **recursive term** that references
the CTE's own name and is re-executed against only the rows produced
by the previous iteration, until it returns nothing.
Each pass only sees the rows the *previous* pass added — not the
whole accumulated result — which is what makes this a proper
breadth-first traversal and not an infinite self-join. `UNION ALL`
(not `UNION`) is deliberate: deduplicating would require comparing
every new row against everything accumulated so far, which is both
slower and wrong if you genuinely expect duplicate rows (e.g. the
same employee ID reachable via two manager paths in a non-tree
graph).
A recursive CTE over a graph with cycles (not a strict tree) can
loop forever unless you break it yourself — track visited IDs in an
array column and add `WHERE NOT e.id = ANY(oc.path)` to the
recursive term, or cap it with a `depth < n` guard.
Next, a different way to keep every row instead of collapsing them
into groups: [Window Functions](/postgres/window-functions).
# Partitioning (/docs/postgres/sql-querying/partitioning)
An [index](/postgres/indexing) speeds up finding rows inside a table.
Partitioning changes what "a table" even means physically: instead of
one heap file, a partitioned table is a thin logical wrapper over
several real child tables, each holding a distinct slice of the rows.
Postgres decides which slice a row belongs in at insert time and which
slices a query even needs to look at, at plan time.
## Declaring a partitioned table [#declaring-a-partitioned-table]
Postgres supports three partitioning strategies, chosen with
`PARTITION BY`:
The parent table (`events`, `orders`, `sessions`) holds no rows of its
own — every row physically lives in exactly one child partition, and
Postgres routes `INSERT`s there automatically based on the partition
key.
## Partition pruning [#partition-pruning]
The whole payoff shows up in `EXPLAIN`: when a query's `WHERE` clause
can be matched against the partition key, the planner eliminates
non-matching partitions before scanning anything in them.
`events_2026_01` never appears in the plan at all — not "scanned and
found empty," genuinely never opened. For a query that only ever
touches recent data, this turns a scan over years of history into a
scan over one month, without an index in sight.
## Bulk operations become metadata operations [#bulk-operations-become-metadata-operations]
Because each partition is a real, independent table, whole ranges of
data can be dropped or moved as a single fast catalog change instead
of a row-by-row `DELETE`:
This is the other half of why time-series workloads reach for
partitioning: "keep 13 months, drop anything older" becomes a monthly
`DROP TABLE` against a single small partition instead of a `DELETE`
that has to find and remove millions of scattered rows — and unlike
`DELETE`, it doesn't leave dead tuples for [VACUUM](/postgres/vacuum)
to clean up afterward.
Partitioning is not free, and more partitions is not automatically
better. Every query against a partitioned table has to consider
every partition during planning, even ones that get pruned — with
hundreds or thousands of tiny partitions, planning time itself
becomes the bottleneck, sometimes exceeding the actual execution
time. A good partition key produces a modest number of meaningfully
large partitions (monthly, not hourly; by region, not by user ID)
— not the finest granularity you can think of. If a table is only
tens of thousands of rows, it almost certainly doesn't need
partitioning at all.
A global unique constraint (a `PRIMARY KEY` that's unique across
*all* partitions, not just within one) requires the partition key to
be part of that key. Postgres can enforce uniqueness per-partition
trivially, but enforcing it cluster-wide would mean checking every
other partition on every insert — so it simply requires the
constraint to include the partition column instead.
Next: the mechanism that makes every one of these writes durable in
the first place, in [Write-Ahead Log](/postgres/wal).
# Views, Functions & Triggers (/docs/postgres/sql-querying/views-functions)
The queries from [Window Functions](/postgres/window-functions) are
useful exactly once, typed out in full, every time. Views name a
query so you can reuse it like a table. Functions and procedures go
further, wrapping logic — not just a single `SELECT` — behind a
callable name. Triggers run that logic automatically, in response to
writes, without the application ever calling anything explicitly.
## Views: a saved query, not saved data [#views-a-saved-query-not-saved-data]
A plain view stores no data of its own — every query against it
re-runs the underlying `SELECT` against current data, live. It's
purely a naming and permissions convenience: you can `GRANT SELECT`
on the view without granting access to the underlying tables, and
hide a complex join behind a simple name.
## Materialized views: a saved result [#materialized-views-a-saved-result]
A materialized view runs the query once and stores the result like a
real table, on disk. Reads against it are fast and don't touch the
underlying tables at all — but the data goes stale the moment
anything underneath changes, until the next `REFRESH`. Plain
`REFRESH MATERIALIZED VIEW` takes an `ACCESS EXCLUSIVE` lock and
blocks reads against it for the duration; `CONCURRENTLY` avoids that
by building the new result alongside the old one and swapping, at the
cost of requiring a unique index on the view and roughly doubling the
disk space used during the refresh.
Postgres materialized views don't refresh themselves — there's no
built-in equivalent of ClickHouse's insert-triggered materialized
views that stay continuously up to date. You own the refresh
schedule, typically via `pg_cron` or an external job.
## Functions: SQL and PL/pgSQL [#functions-sql-and-plpgsql]
The simplest function is a named, parameterized SQL query:
For anything with branching, loops, or multiple statements, PL/pgSQL
(Postgres's procedural extension of SQL) is the usual choice:
## Procedures: functions that can control transactions [#procedures-functions-that-can-control-transactions]
Functions always run inside the transaction of the caller and can
never issue `COMMIT` or `ROLLBACK` themselves. Procedures
(`CREATE PROCEDURE`, invoked with `CALL`, added in Postgres 11) can
— useful for batch jobs that need to commit progress in chunks
instead of holding one enormous transaction open, or that need to
keep going even if one chunk fails.
## Triggers: functions that fire automatically [#triggers-functions-that-fire-automatically]
A trigger function looks like a normal PL/pgSQL function but returns
`trigger` and reads the special `NEW`/`OLD` row variables that hold
the row being inserted/updated/deleted:
`OLD` holds the row's values before the change, `NEW` holds them
after — `OLD` is `NULL` for an `INSERT` trigger, `NEW` is `NULL` for a
`DELETE` trigger. `BEFORE` triggers can inspect and modify `NEW`
before it's written (returning `NULL` from a `BEFORE` trigger cancels
the operation entirely); `AFTER` triggers, like this one, run once the
change has already happened and are the natural fit for logging,
since they can't affect whether the write succeeds.
Triggers are invisible at the call site — an `UPDATE payments SET
amount = ...` gives no syntactic hint that it also writes to
`payments_audit`. That's convenient until it's a debugging trap:
unexplained rows, slower-than-expected writes, or a trigger that
throws and silently rolls back an otherwise-fine transaction. Keep
trigger logic small, and check `information_schema.triggers` (or
`\d payments` in `psql`) before assuming a table's writes have no
side effects.
Next: what happens when a single table like `payments` grows too
large to manage as one physical object, in
[Partitioning](/postgres/partitioning).
# Window Functions (/docs/postgres/sql-querying/window-functions)
`GROUP BY`, from [Joins, CTEs & Recursion](/postgres/joins-ctes),
collapses matching rows into one row per group — you lose the
individual rows in exchange for the aggregate. A window function
computes the same kinds of aggregates but keeps every row: a running
total next to each transaction, a rank next to each score, the
previous row's value next to the current one.
## The anatomy of OVER [#the-anatomy-of-over]
* `PARTITION BY` groups rows the same way `GROUP BY` would — but every
row in the group survives in the output, not just the aggregate.
* `ORDER BY` inside `OVER (...)` defines the order the window function
walks rows in *within each partition*. It has nothing to do with the
final result set's order — you still need an outer `ORDER BY` for
that.
* The frame clause, `ROWS BETWEEN ... AND ...`, controls exactly which
rows around the current one are included in the calculation.
## Ranking functions [#ranking-functions]
The three differ only in how they treat ties. Given spend values
`100, 100, 90`:
* `row_number()` gives `1, 2, 3` — always unique, breaks ties
arbitrarily (by whatever order is stable for equal keys).
* `rank()` gives `1, 1, 3` — ties share a rank, and the next rank
skips ahead by the number of tied rows.
* `dense_rank()` gives `1, 1, 2` — ties share a rank, but the next
rank is always just one higher, with no gap.
Other common window functions: `lag(col, n)` / `lead(col, n)` return
the value of a column `n` rows before/after the current one in the
ordered partition — the standard tool for period-over-period
comparisons (this month vs. last month) without a self-join.
## Frames and running aggregates [#frames-and-running-aggregates]
The frame clause is what makes running totals and moving averages
possible. Omitting it defaults to `RANGE BETWEEN UNBOUNDED PRECEDING
AND CURRENT ROW` when an `ORDER BY` is present — which is subtly
different from `ROWS`, so it's worth being explicit:
`ROWS BETWEEN 2 PRECEDING AND CURRENT ROW` is a physical window: the
current row plus the two immediately before it, always three rows
wide regardless of ties in `ORDER BY`. `RANGE` instead groups by
*value* — every row with an equal `ORDER BY` value is treated as part
of the same peer group and included together, which matters when the
ordering column has duplicates.
Window functions run as their own step in query processing — after
`WHERE`, `GROUP BY`, and `HAVING`, but before the final `SELECT
DISTINCT`, `ORDER BY`, and `LIMIT`. That ordering is why you can't
reference a window function's output directly in the same query's
`WHERE` clause (it doesn't exist yet at that stage) — you have to
wrap it in a subquery or CTE and filter the outer query instead.
## Window function vs. GROUP BY: same math, different shape [#window-function-vs-group-by-same-math-different-shape]
Both compute the same per-department average. `GROUP BY` throws away
the individual employee rows to get there; the window function keeps
every employee row and attaches the department's average to each one
— which is exactly what you need for something like "how far above or
below their department's average does each employee sit," a
calculation that's impossible to express with `GROUP BY` alone since
it needs both the individual row and the aggregate at the same time.
Next: turning a query like this into something reusable, in
[Views, Functions & Triggers](/postgres/views-functions).
# Heap Storage (/docs/postgres/storage-internals/heap-storage)
Every table you create with [SQL Basics](/postgres/sql-basics) has to live
somewhere on disk. Postgres uses a **heap** — an unordered collection of
fixed-size pages, each holding a handful of rows — and understanding its
shape explains a lot of behavior that otherwise looks like magic, from why
`SELECT *` on a huge table is slow to why deleted rows don't shrink a table
immediately.
## Pages [#pages]
A table's data file is divided into 8KB **pages** (also called blocks).
Every read and write happens a whole page at a time — Postgres never touches
less than 8KB of disk, even to read a single row. A page holds:
* a header (checksums, free space pointers)
* an array of **item pointers** (line pointers), each pointing at a tuple
later in the same page
* the tuples themselves, packed in from the end of the page backwards
## Tuples [#tuples]
A **tuple** is one physical row version. Every tuple carries a fixed header
in front of your actual column data, including:
* `xmin` — the transaction ID that created this tuple version
* `xmax` — the transaction ID that deleted/replaced it (0 if still live)
* `ctid` — this tuple's own physical address, `(page_number, item_index)`
You can see `ctid` directly:
`(0,1)` means "page 0, item pointer 1." `ctid` changes every time a row is
updated — that's the first hint that `UPDATE` isn't an in-place edit, which
is the whole subject of the next lesson, [MVCC](/postgres/mvcc).
Postgres does have a Heap-Only Tuple (HOT) optimization: if a new tuple
version fits on the *same page* and no indexed column changed, it chains
the old item pointer straight to the new tuple instead of touching every
index. It reduces index bloat, but the row still gets a new physical
location — HOT doesn't mean "updated in place."
## Relation files on disk [#relation-files-on-disk]
Each table (and each index) is backed by one or more physical files under
the data directory, named after the table's `relfilenode`, not its name:
Files are capped at 1GB each — a table larger than that spills into
`16391.1`, `16391.2`, and so on, all still logically one relation. This is
also why `ALTER TABLE ... RENAME` is instant: the file name never changes,
only the catalog entry does.
To see how many pages (and bytes) a table actually occupies:
`pg_relation_size` reports only the heap itself. A table's total footprint
— indexes, [TOAST](/postgres/toast) tables for large values, the free
space map — is usually several times larger. Use
`pg_total_relation_size('accounts')` for the real number.
Rows shrink out of a page only when [VACUUM](/postgres/vacuum) reclaims the
space old tuple versions leave behind — a delete or update doesn't free
anything immediately. Next: what happens when a single column's value is too
big to fit in a page at all, in [TOAST, FSM & Visibility Map](/postgres/toast).
# Storage Internals (/docs/postgres/storage-internals)
How Postgres actually stores rows on disk, and why old ones stick around.
# MVCC (/docs/postgres/storage-internals/mvcc)
Postgres never blocks a reader to let a writer finish, and never blocks a
writer to let a reader finish. That guarantee — high concurrency without
constant locking — comes from **MVCC** (Multi-Version Concurrency Control),
and it's the single idea that explains the most about how Postgres actually
behaves under load.
## UPDATE doesn't update [#update-doesnt-update]
Recall from [Heap Storage](/postgres/heap-storage) that every tuple carries
an `xmin` (creating transaction) and `xmax` (deleting transaction) in its
header. `UPDATE` uses both: it marks the old tuple's `xmax` with the current
transaction ID and inserts a brand-new tuple with a fresh `xmin` — the old
row isn't touched or overwritten, it's just marked dead.
`DELETE` is the same idea with no new tuple: it just sets `xmax`. The old
tuple version keeps occupying its page until
[VACUUM](/postgres/vacuum) comes along and reclaims it — which is why a
table with heavy update traffic grows on disk even though its row count
stays flat.
This is exactly why `UPDATE`-heavy tables need routine vacuuming: every
update leaves a dead tuple behind. Skip vacuum for long enough and the
table can bloat to many times its logical size.
## Snapshots and visibility [#snapshots-and-visibility]
Every transaction gets a **snapshot** the moment it starts (or, under Read
Committed — see [Transactions & Isolation](/postgres/transactions) — the
moment each statement starts): a record of which transaction IDs were
already committed, in-progress, or not yet started at that instant.
A tuple is visible to your transaction only if:
* its `xmin` committed before your snapshot was taken, **and**
* its `xmax` is either unset, or belongs to a transaction that hadn't
committed by your snapshot
Session A keeps seeing the pre-update value for its whole transaction
(under Repeatable Read or Serializable) because its snapshot was taken
before B committed — not because anything was locked. B's write never had
to wait for A's read, and A's read never had to wait for B's write.
This is the actual mechanism behind "readers don't block writers, writers
don't block readers." There's no read lock being avoided — there's simply
more than one physical version of the row on disk at once, and each
transaction is handed the version consistent with its own snapshot.
## Multiple versions, one table [#multiple-versions-one-table]
At any moment a busy table can have several live versions of the "same"
logical row, each visible to a different set of in-flight transactions.
Postgres — and specifically `VACUUM` — is responsible for eventually
deciding a version is visible to *nobody* and reclaiming it.
A transaction ID is a 32-bit counter that wraps around. If a table went
unvacuumed forever, old `xmin` values would eventually look like they're
"in the future" again as the counter wraps — Postgres prevents this by
forcing a freeze (rewriting old `xmin`s to a special frozen marker) well
before that can happen.
MVCC explains what happens when two transactions touch the same row at the
same time. What happens when they *lock* the same row on purpose is next, in
[Transactions & Isolation](/postgres/transactions).
# TOAST, FSM & Visibility Map (/docs/postgres/storage-internals/toast)
[Heap Storage](/postgres/heap-storage) fixes every page at 8KB. That's fine
for a `numeric` or a short `text` value, but what happens when you insert a
50KB JSON blob into a single column? Postgres has to move it somewhere else
— and it needs to track, per page, how much room is left and which pages
are safe to skip during a scan. Three mechanisms handle this: TOAST, the
Free Space Map, and the Visibility Map.
## TOAST [#toast]
**TOAST** (The Oversized-Attribute Storage Technique) kicks in automatically
for variable-length columns — `text`, `jsonb`, `bytea`, arrays — once a row
would exceed roughly 2KB (Postgres targets fitting at least 4 tuples per
page). The oversized value is compressed and, if still too big, sliced into
chunks stored in a hidden companion table:
Each column has a **storage strategy** controlling this behavior:
* `PLAIN` — never compressed or moved out-of-line (used for fixed-size types
like `integer`)
* `EXTENDED` — compress first, then move out-of-line if still too large
(the default for `text`/`jsonb`)
* `EXTERNAL` — move out-of-line without compressing (faster substring
access, larger storage)
* `MAIN` — compress, but avoid moving out-of-line unless there's no other
choice
A wide `jsonb` column means every `SELECT *` — even one that never touches
that column in its `WHERE` clause — can pay for a TOAST fetch if the
column is included in the result. Select only the columns you need on hot
paths.
## Free Space Map (FSM) [#free-space-map-fsm]
Every relation keeps a **Free Space Map** — a compact side-file recording
roughly how much free space each page has. When you `INSERT`, Postgres
consults the FSM to find a page with room instead of scanning the whole
table or always appending to the end. It's what makes space freed by
[VACUUM](/postgres/vacuum) reusable rather than wasted.
## Visibility Map (VM) [#visibility-map-vm]
The **Visibility Map** is a bitmap, one bit (actually two) per page,
tracking whether a page is:
* **all-visible** — every tuple on the page is visible to all current and
future transactions, so no [MVCC](/postgres/mvcc) visibility check is
needed
* **all-frozen** — every tuple's `xmin` is old enough that it no longer
needs to be considered during transaction-ID wraparound freezing
The VM matters for two very different reasons:
**Index-only scans.** If a query only needs indexed columns, Postgres can
answer entirely from the index — skipping the heap — but only for pages
the VM marks all-visible. A table with a lot of recent write activity has
a "colder" VM and falls back to the heap more often, even for
index-covered queries.
**Vacuum skipping.** Autovacuum uses the VM to skip pages it already knows
are all-visible/all-frozen, which is why a mostly-read, rarely-written
table vacuums fast even when it's huge — most of its pages are never
revisited.
Both maps live in small companion files next to the heap (`_fsm`, `_vm`) and
are updated incrementally as you write and vacuum. Up next: the mechanism
that actually decides which tuples are visible in the first place —
[MVCC](/postgres/mvcc).
# Transactions & Concurrency (/docs/postgres/transactions-concurrency)
Isolation levels, locks, and what blocks what.
# Locks & Deadlocks (/docs/postgres/transactions-concurrency/locks)
[Transactions & Isolation](/postgres/transactions) introduced `SELECT ...
FOR UPDATE` as a way to lock a row on purpose. Postgres actually has two
separate locking systems — row-level and table-level — plus an
application-level escape hatch, and knowing which one is holding things up
is most of debugging a "why is this query just hanging" incident.
## Row-level locks [#row-level-locks]
Row locks come in a few flavors, from weakest to strongest:
* `FOR UPDATE` — locks rows as if for `UPDATE`; blocks other `FOR UPDATE`,
`FOR SHARE`, `UPDATE`, and `DELETE` on the same rows
* `FOR NO KEY UPDATE` — like `FOR UPDATE` but doesn't conflict with
`FOR KEY SHARE` (used internally for updates that don't touch a foreign
key's referenced columns)
* `FOR SHARE` — locks rows for reading; multiple transactions can hold it
concurrently, but it blocks concurrent `UPDATE`/`DELETE`
* `FOR KEY SHARE` — the weakest; only conflicts with locks that would
change a row's key values
## Table-level lock modes [#table-level-lock-modes]
Every statement takes a table-level lock too, even a plain `SELECT`
(`ACCESS SHARE`, the weakest mode — it only conflicts with `ACCESS
EXCLUSIVE`). The eight modes form a conflict matrix from weakest to
strongest:
`ALTER TABLE ... ADD COLUMN` with a non-`NULL` default used to rewrite the
whole table under an `ACCESS EXCLUSIVE` lock on older Postgres versions.
Since Postgres 11 it's instant for a constant default, but many `ALTER
TABLE` forms — adding a `CHECK` constraint, changing a column type — still
take `ACCESS EXCLUSIVE` and block every reader and writer until they
finish. Run those during low traffic, or check for a `CONCURRENTLY`
variant first.
## Deadlocks [#deadlocks]
A deadlock happens when two transactions each hold a lock the other is
waiting for:
Postgres runs a deadlock detector (checking every second by default) that
finds the cycle and aborts one of the transactions with `ERROR: deadlock
detected`, letting the other proceed. The fix is almost always at the
application level: always lock rows (or tables) in a consistent order
across every code path.
## Advisory locks [#advisory-locks]
Sometimes you want to coordinate application logic that has nothing to do
with a specific row — a cron job that shouldn't run twice concurrently, for
example. **Advisory locks** are arbitrary integer locks Postgres tracks for
you without attaching them to any table:
They're cheap, application-defined, and never conflict with real table or
row locks — a useful tool for distributed coordination when you already
have a Postgres connection open.
Locking controls what happens when queries collide. [Indexing](/postgres/indexing),
next, is about avoiding unnecessary collisions in the first place by making
queries touch far fewer rows.
# Transactions & Isolation (/docs/postgres/transactions-concurrency/transactions)
[MVCC](/postgres/mvcc) explained how snapshots let transactions avoid
blocking each other. This lesson covers the transaction itself — how to
control where it starts and ends, and how much of that concurrent activity
it's allowed to see.
## BEGIN, COMMIT, ROLLBACK, SAVEPOINT [#begin-commit-rollback-savepoint]
Every statement outside an explicit transaction runs in its own
implicit one. Wrapping several statements in `BEGIN`/`COMMIT` makes them
atomic — either all of them take effect or none do:
`SAVEPOINT` lets you roll back part of a transaction without abandoning the
whole thing — useful for "try this, and if it fails, fall back" logic
inside a single transaction:
Once any statement in a transaction errors, the *whole transaction* is
aborted and every subsequent statement is rejected with "current
transaction is aborted" — until you either `ROLLBACK` entirely or roll
back to a savepoint taken before the error.
## Isolation levels [#isolation-levels]
The SQL standard defines four isolation levels, each permitting fewer
anomalies than the last. Postgres implements three of them distinctly — Read
Uncommitted behaves exactly like Read Committed, because Postgres's MVCC
never allows a dirty read in the first place.
| Level | Dirty read | Non-repeatable read | Phantom read | Serialization anomaly |
| ---------------------------- | ------------------ | ------------------- | ------------ | --------------------- |
| Read Uncommitted | never (same as RC) | possible | possible | possible |
| **Read Committed** (default) | never | possible | possible | possible |
| Repeatable Read | never | never | never\* | possible |
| Serializable | never | never | never | never |
Postgres defaults to **Read Committed**, not Repeatable Read like some
other engines. Under Read Committed, each *statement* gets a fresh
snapshot — a long transaction can see a different value for the same row
each time it queries, if another transaction committed a change in
between.
**Repeatable Read** takes one snapshot for the whole transaction — every
query sees the data exactly as it stood when the transaction began.
Postgres's implementation also blocks phantom reads (Postgres's Repeatable
Read is actually closer to the standard's Snapshot Isolation than the bare
minimum the spec requires).
**Serializable** goes further: it detects when concurrent transactions
*would* have produced a result impossible under any serial (one-at-a-time)
execution order, and aborts one of them with a serialization error your
application must be ready to retry.
## Locking a row on purpose [#locking-a-row-on-purpose]
`SELECT ... FOR UPDATE` takes an explicit row lock as part of a read,
blocking other transactions from locking or updating the same rows until
you commit — the classic "read a balance, then update it" pattern:
This is a bridge into explicit locking — the full set of lock modes,
what conflicts with what, and how deadlocks get resolved, is next in
[Locks & Deadlocks](/postgres/locks).
# Backup & Restore (/docs/postgres/wal-vacuum-recovery/backup-restore)
[VACUUM](/postgres/vacuum) keeps a running database healthy, but it
does nothing for the failure modes that matter most: a dropped table,
a bad migration, a corrupted disk, a whole machine gone. Backups are
the only defense against those — and Postgres gives you two genuinely
different kinds, suited to different recovery goals.
## Logical backups: pg\_dump and pg\_dumpall [#logical-backups-pg_dump-and-pg_dumpall]
`pg_dump` exports a database's schema and data as SQL (or a portable
binary format) — a **logical** snapshot, independent of the exact
files Postgres happens to store it in on disk.
`pg_dump` operates on one database at a time and never touches roles
or tablespaces — that's what `pg_dumpall` is for, usually paired with
per-database `pg_dump -Fc` dumps in the same backup routine: `pg_dumpall --globals-only` for roles, individual `pg_dump`s for data.
## Restoring: pg\_restore [#restoring-pg_restore]
The custom format (`-Fc`) isn't plain SQL — it's restored with
`pg_restore`, which (unlike replaying a giant `.sql` file) can run in
parallel and lets you select individual tables:
A plain-SQL dump (`pg_dump > file.sql`) restores with `psql -f
file.sql`, statement by statement, in file order — simple, but
single-threaded and all-or-nothing. `-Fc` restores with `pg_restore`,
which can parallelize table loads and index builds across multiple
jobs and lets you cherry-pick objects to restore, at the cost of the
file no longer being human-readable.
## Physical (base) backups: pg\_basebackup [#physical-base-backups-pg_basebackup]
A logical dump reconstructs data by re-running `INSERT`s — it's slow
to restore on a large database and, on its own, only ever gets you
back to the exact moment the dump was taken. A **physical** backup
copies the actual data directory — heap files, indexes, everything —
byte for byte:
`-Xs` streams the WAL generated *during* the backup alongside it,
which matters: copying files while the database is live and being
written to produces an internally inconsistent snapshot on its own —
it's only made consistent by replaying that WAL on startup, the same
crash-recovery machinery from [Write-Ahead Log](/postgres/wal).
## Point-in-time recovery (PITR) [#point-in-time-recovery-pitr]
A base backup plus a continuous WAL archive (`archive_command`, also
covered in [Write-Ahead Log](/postgres/wal)) together let you restore
to *any* moment since the base backup was taken — not just the moment
it happened to be taken. That's PITR, and it's the only tool that
answers "someone ran a bad `DELETE` at 2:14pm, get us back to 2:13pm"
without losing every other commit up to that second.
The recipe, at a high level:
On startup, Postgres notices `recovery.signal`, pulls WAL segments
from the archive via `restore_command`, and replays them forward from
the base backup — the exact same replay mechanism as ordinary crash
recovery, just fed a longer history and told where to stop instead of
running to the end.
This is why WAL archiving genuinely is the backbone of PITR, not
just a nice-to-have: without it, a base backup alone only restores to
its own backup time, and a logical dump only restores to its dump
time. The archived WAL stream is what fills every gap in between,
letting you recover to a target measured in seconds, not "whenever
the last backup happened to run."
## Choosing between them [#choosing-between-them]
* **`pg_dump`/`pg_dumpall`** — portable across Postgres versions and
even hardware architectures, easy to inspect, good for migrating a
single database or seeding a dev environment. Slow to restore on a
large database; only restores to the exact dump moment.
* **`pg_basebackup` + WAL archiving** — restores an entire cluster
fast (it's a file copy, not replayed `INSERT`s) and, combined with
PITR, recovers to any second in the archived window. Tied to the
same Postgres major version and, generally, the same architecture.
Most production setups run both: periodic `pg_dump`s for
portability/dev-seeding, and continuous base backups + WAL archiving
as the real disaster-recovery path.
A backup that has never been restored is a hypothesis, not a safety
net — this is doubly true for PITR, since the failure mode you're
protecting against (a corrupted archive, a `restore_command` typo, a
gap in WAL continuity) is invisible until the moment you actually
need it. Practice a full restore, including PITR to an arbitrary
timestamp, on a schedule, into a throwaway environment — not for the
first time during an actual incident.
Next: instead of restoring from a backup after the fact, keeping a
second server continuously up to date in the first place, in
[Replication](/postgres/replication).
# WAL, Vacuum & Recovery (/docs/postgres/wal-vacuum-recovery)
The log that makes crash recovery and backups possible, and the cleanup that keeps tables lean.
# VACUUM & Autovacuum (/docs/postgres/wal-vacuum-recovery/vacuum)
[MVCC](/postgres/mvcc) makes concurrent reads and writes safe by never
overwriting a row in place: an `UPDATE` writes a brand-new tuple and
marks the old one dead (`xmax` set) rather than mutating it, and a
`DELETE` just marks a tuple dead without removing it. Nobody reclaims
that space automatically as part of the write — that's `VACUUM`'s
entire job, and if it falls behind, the table just keeps growing.
## Dead tuples and why they pile up [#dead-tuples-and-why-they-pile-up]
That single statement doesn't touch the old row at all — it inserts a
new tuple with the new balance and sets `xmax` on the old tuple to the
current transaction ID. The old tuple stays on its page, fully intact,
invisible to any transaction that starts after this one commits, but
still physically present and still costing space and I/O until
something removes it.
## What VACUUM actually does [#what-vacuum-actually-does]
Plain `VACUUM` scans the table, identifies tuples that are dead to
*every* current and future transaction (no open snapshot can still see
them), and marks that space reusable by recording it in the table's
**free space map** (from
[TOAST, FSM & Visibility Map](/postgres/toast)). Crucially, it does
**not** shrink the file on disk — freed space is only made available
for *future* inserts and updates to reuse. A table that had a huge
`DELETE` and then a plain `VACUUM` will occupy exactly the same number
of bytes on disk as before, just with more of those bytes marked free
internally.
`VACUUM` runs concurrently with normal reads and writes — it only
needs a lightweight lock, never blocking queries against the table.
## VACUUM FULL: the exclusive-lock alternative [#vacuum-full-the-exclusive-lock-alternative]
`VACUUM FULL` actually shrinks the file: it rewrites the entire table
into a new file containing only live tuples, then swaps it in and
drops the old one — the same technique as `CLUSTER`. That's the only
way to hand disk space back to the operating system after a huge
delete. The cost is an `ACCESS EXCLUSIVE` lock for the whole
operation, blocking every read and write against the table until it
finishes.
`VACUUM FULL` on a large table can lock it for minutes or hours.
It's a maintenance-window operation, not something to run reflexively
because `pg_relation_size` looks bigger than expected — plain
`VACUUM` (or just waiting for autovacuum) is almost always the right
first move, since disk space that's marked reusable gets consumed by
future writes anyway.
## Autovacuum [#autovacuum]
In practice you rarely run `VACUUM` by hand — the **autovacuum**
daemon does it automatically, per table, based on how many rows have
changed since the last run:
For a 10-million-row table, the default scale factor means autovacuum
waits until roughly 2 million rows are dead before it bothers — fine
for a slowly-changing table, potentially a real problem for a small,
extremely hot one.
The single most common autovacuum production incident: a small,
high-churn table (a queue, a session table, a counters table) gets
updated thousands of times a minute, and the *default* scale-factor
threshold means autovacuum only fires occasionally relative to how
fast dead tuples accumulate. Meanwhile every query against the table
has to scan past a growing number of dead tuples, index bloat grows
alongside it, and the table's disk footprint climbs even though its
logical row count never changes. The fix is a per-table override —
lower thresholds for hot tables specifically, e.g. `ALTER TABLE
queue_jobs SET (autovacuum_vacuum_scale_factor = 0.01,
autovacuum_vacuum_cost_delay = 2)`. Watch
`pg_stat_user_tables.n_dead_tup` and `last_autovacuum` for any table
under heavy write load — don't wait for query latency to surface the
problem first.
## Transaction ID wraparound and VACUUM FREEZE [#transaction-id-wraparound-and-vacuum-freeze]
Postgres transaction IDs (`xid`) are a 32-bit counter, and MVCC
visibility depends on comparing `xid`s to decide what's older or
newer. A 32-bit counter wraps around eventually, and if an old tuple's
`xmin` were left as a raw comparable number forever, wraparound would
make old committed rows suddenly look like they came from the future
— catastrophic silent data-visibility corruption.
`VACUUM` prevents this by **freezing** old tuples: once a tuple is old
enough that every current and future transaction is guaranteed to see
it as committed, its `xmin` is replaced with a special
frozen marker that's always considered "in the past," permanently,
regardless of counter wraparound.
If autovacuum is disabled, misconfigured, or simply can't keep up on
a large enough table, `age(datfrozenxid)` climbs toward the
wraparound limit. Postgres protects itself by refusing new writes
cluster-wide once it gets dangerously close — a full outage that can
only be fixed by running `VACUUM FREEZE` (often taking the affected
tables offline for a while to catch up). This is one of the few
Postgres failure modes that's entirely preventable and entirely
self-inflicted: monitor transaction age, don't disable autovacuum,
and don't let it starve on a busy table.
Next: even with vacuum keeping the live table lean, you still need a
plan for disasters vacuum can't help with, in
[Backup & Restore](/postgres/backup-restore).
# Write-Ahead Log (/docs/postgres/wal-vacuum-recovery/wal)
Every [transaction](/postgres/transactions) that commits has to
survive a crash the instant after it commits — a power failure one
millisecond after `COMMIT` returns must not be able to lose that data.
Postgres doesn't guarantee this by forcing every changed data page to
disk before acknowledging a commit; that would make every write pay
for a random-I/O flush of scattered 8KB pages. Instead it writes
everything to one append-only log first, and that log is what actually
has to hit disk before a commit is considered done.
## Write-ahead, literally [#write-ahead-literally]
The rule the WAL (Write-Ahead Log) is named after: **a change is
described in the log before the data page it modifies is allowed to
reach disk.** The log entry itself is small and sequential — cheap to
fsync — while the actual data page can be flushed later, lazily, in
whatever order is convenient.
{/* Not a chain: the WAL buffer's contents fork into two independent
destinations after the fsync gate — the client gets acked off the
fast, sequential WAL write, while the actual heap page is flushed
to its data file later, on its own schedule (checkpoint or buffer
eviction), completely decoupled from when the client saw success. */}
The client only waits on the top path — buffer, fsync, ack. The heap
page carrying the actual row data can sit dirty in
[shared buffers](/postgres/architecture) for a while after the commit
already returned; it gets written out later by the background writer
or the next checkpoint, whichever comes first. This is exactly what
makes commits fast: one small sequential write instead of N scattered
ones, per transaction.
## Why this makes crash recovery possible [#why-this-makes-crash-recovery-possible]
If Postgres crashes before a dirty heap page reaches disk, the change
still exists — durably — in the WAL. On restart, Postgres replays every
WAL record since the last checkpoint against the on-disk pages,
reconstructing exactly the state that existed at the moment of the
crash. This is why a `postgres` container that gets `kill -9`'d
mid-write comes back with all committed data intact and nothing else:
recovery isn't guesswork, it's deterministic replay of a durable log.
This same mechanism is why a `ROLLBACK`ed transaction's changes
never corrupt anything: its WAL records get written too (Postgres
doesn't know in advance a transaction will abort), but each record
carries the transaction ID that wrote it, and the visibility rules
from [MVCC](/postgres/mvcc) simply treat that transaction's tuples as
never having happened. Replay restores physical state; MVCC decides
what's visible.
## Checkpoints [#checkpoints]
Replaying WAL back to the very beginning of time on every restart
would make recovery take longer the older the database gets. A
**checkpoint** periodically flushes all currently-dirty buffers to
their data files and records the WAL position at that moment. Recovery
after a crash only ever needs to replay WAL from the most recent
checkpoint forward, not from the dawn of the cluster.
Checkpoints are a tradeoff: too infrequent and recovery after a crash
takes longer (more WAL to replay); too frequent and you pay repeated
write I/O flushing pages that would otherwise have been overwritten
again before ever being flushed.
## WAL segments and pg\_wal/ [#wal-segments-and-pg_wal]
WAL isn't one infinite file — it's split into fixed-size **segments**,
16MB each by default, stored in the `pg_wal/` directory inside the
data directory. Postgres writes into the current segment sequentially
and rolls to a new one when it fills up.
Segments no longer needed for crash recovery (older than the latest
checkpoint) are normally recycled — renamed and reused for future WAL,
rather than deleted and recreated — unless something is configured to
keep them around longer, which is exactly what archiving does.
## Archiving: the backbone of point-in-time recovery [#archiving-the-backbone-of-point-in-time-recovery]
By default, an old WAL segment is just recycled once it's no longer
needed for crash recovery — its contents are gone. Turning on
`archive_mode` runs `archive_command` against every completed segment
before it's allowed to be recycled, copying it somewhere durable (S3,
another disk, a dedicated archive server) instead of letting it be
overwritten.
A continuous archive of every WAL segment, combined with one physical
base backup taken at some point in the past, is what lets Postgres
replay forward to *any* moment since that backup — not just the
moments a full backup happened to be taken. That combination is the
entire mechanism behind
[Point-in-Time Recovery](/postgres/backup-restore), and it's also the
same stream of WAL records that
[replication](/postgres/replication) ships to standby servers in
real time instead of archiving to a file.
If `archive_mode` is on but `archive_command` starts failing
silently (disk full, permissions, network), Postgres keeps every WAL
segment since the last successful archive — it can't recycle them,
because losing an unarchived segment would break the archive's
continuity. `pg_wal/` grows without bound until the disk fills up,
which is itself an outage. Monitor archiving success, not just that
the server is up.
Next: what happens to the *old* row versions that WAL-logged updates
leave behind, in [VACUUM & Autovacuum](/postgres/vacuum).
# Why PostgreSQL (/docs/postgres/why-postgresql)
What problem it solves, and when it's the wrong choice.
# Why PostgreSQL (/docs/postgres/why-postgresql/why-postgres)
PostgreSQL is a row-oriented, general-purpose database built for OLTP:
lots of small, concurrent transactions that read and write individual
rows safely. That's a different job than
[ClickHouse](/clickhouse/why-clickhouse), which is built for OLAP —
scanning and aggregating billions of rows at once. Knowing which
category a workload falls into is the first design decision you make,
before any schema or index.
## OLTP vs OLAP [#oltp-vs-olap]
**OLTP** (Online Transaction Processing) means many short transactions
touching a handful of rows each — placing an order, updating a user's
profile, checking an inventory count. The access pattern is narrow and
deep: `WHERE id = 42`, update three columns, commit. **OLAP** (Online
Analytical Processing) means the opposite — few queries, but each one
scans millions or billions of rows to compute an aggregate. Postgres
is tuned for the first pattern: row-at-a-time storage, indexes for
fast point lookups, and a transaction model that lets thousands of
clients write concurrently without corrupting each other's data.
## A brief history [#a-brief-history]
Postgres traces back to the **POSTGRES** project at UC Berkeley in the
1980s, led by Michael Stonebraker as a successor to his earlier
Ingres database. In the mid-90s it gained a SQL interface and was
renamed **PostgreSQL**. Decades of continuous open-source development
later, it's one of the most feature-complete relational databases
that exists — full SQL, extensibility (custom types, functions,
extensions), and a reputation for taking correctness seriously over
cutting corners for speed.
## ACID, briefly [#acid-briefly]
Postgres is an ACID database:
* **Atomicity** — a transaction's changes all happen, or none do.
* **Consistency** — a transaction moves the database from one valid
state to another, respecting constraints.
* **Isolation** — concurrent transactions don't see each other's
uncommitted changes (with tunable strength — see
[Transactions & Isolation](/postgres/transactions)).
* **Durability** — once committed, a transaction's changes survive a
crash (this is what the [Write-Ahead Log](/postgres/wal) exists
for).
ACID is the whole point of choosing a relational database like
Postgres over something eventually-consistent. If your application
can tolerate stale or partial writes, you have more options; if it
can't — payments, inventory, anything where "close enough" causes
real damage — ACID guarantees are what you're paying for in
engineering complexity.
## When Postgres is the wrong choice [#when-postgres-is-the-wrong-choice]
* **Huge analytical scans.** Aggregating billions of rows across a
handful of columns is ClickHouse's job, not Postgres's — Postgres
stores whole rows together on disk, so a query touching two columns
out of fifty still reads all fifty.
* **Pure key-value caching.** If you need sub-millisecond lookups of
ephemeral data with no durability requirement, Redis is a better
fit — Postgres's durability and MVCC machinery is overhead you don't
need.
* **Unstructured document dumping ground.** Postgres can store JSON
(see [JSON & Arrays](/postgres/json-arrays)), but if literally
everything is schema-less documents with no relational structure,
a document database may fit the access patterns better.
None of this means Postgres can't do some of these things — it can,
often well enough — but knowing the shape it was designed for tells
you when you're fighting the tool instead of using it.
This module starts with [Docker & Setup](/postgres/setup) — getting a
real Postgres instance running locally before any of the internals
below mean anything concrete.