1.1 What problem does Kafka solve?
The problem is not "we need a queue." The problem is N×M point-to-point coupling in a data platform.
Topics & partitions
- 4
- max consumers
- per key
- ordering
Order preserved per key. All records for a key hash to the same partition, so a given account's events stay in sequence while 4 consumers work in parallel. Note that adding partitions later changes the hash mapping — the ordering guarantee spans the change only if you never repartition.
The problem is not "we need a queue." The problem is N×M point-to-point coupling in a data platform.
The decay sequence (this is the actual origin story)
Stage 1 — one producer, one consumer. You have an app emitting metrics and a dashboard. You open a direct TCP connection. This works. It is the correct engineering decision at this size.
A direct TCP connection. “This works. It is the correct engineering decision at this size.”
Stage 2 — requirements multiply. You want long-term metric analysis, so you add a storage/analysis service. Now the app writes to two places. Then 3 more apps start emitting metrics. Then a coworker wants pull-based alerting, so every app also runs an HTTP metrics server. Then more consumers appear pulling from those servers.
O(N×M) connections. Every new consumer requires touching every producer, and because every producer knows every consumer, adding a consumer is a deployment of the producer fleet.
The cost here is real and specific, not aesthetic:
- O(N×M) connections. Every new consumer requires touching every producer.
- Every producer knows every consumer. Adding a consumer is a deployment of the producer fleet.
- Mixed push and pull. Some paths push, some poll. Different latency, different failure modes, different code.
- Format drift. Each pair negotiated its own format. No two are the same.
Stage 3 — you build a broker. You put a single service in the middle that accepts metrics from everyone and serves them to anyone.
Congratulations — you have built a publish/subscribe messaging system. This is the point the book makes bluntly: pub/sub is not a technology you choose, it's the shape every sufficiently-grown data platform converges on. The only question is whether you build it deliberately or accidentally.
Stage 4 — the real problem. Meanwhile, a coworker independently built the same thing for log messages. Another built it for user activity tracking. You now have three separate pub/sub systems, each with its own bugs, its own operational runbook, its own scaling limits, and its own on-call rotation.
Three separate pub/sub systems, each with its own bugs, its own operational runbook, its own scaling limits, and its own on-call rotation:
- metrics pub/sub
- logging pub/sub
- user-activity pub/sub
3 systems · 3 sets of bugs · 3 scaling stories · 0 shared tooling — and you know a 4th use case is coming next quarter.
This is the problem Kafka solves: not "move messages," but one general-purpose, horizontally scalable, durable data backbone that any type of data can flow through, so the organization stops re-solving pub/sub per data type.
The mental model that follows
Kafka is described two ways in the book, and both matter:
- "A distributed commit log." A DB or filesystem commit log is a durable, ordered record of every transaction, so state can be rebuilt deterministically by replay. Kafka is that, as a service.
- "A distributed streaming platform." Same substrate, framed as continuous data-in-motion rather than storage.
The commit-log framing is the load-bearing one. It explains every property Kafka has:
| Because it's a log… | You get… |
|---|---|
| Appends are sequential | Extremely high write throughput on cheap disks |
| Reads are positional (offset) | Many independent readers, no per-consumer state on the broker |
| The log is retained, not consumed | Replay, reprocessing, late consumers, backfill |
| Log is deterministic and ordered | State can be rebuilt; stream processing is well-defined |
| Log can be sharded and copied | Horizontal scale + redundancy |