Learn Labs
1. Meet Kafka

1.1 What problem does Kafka solve?

The problem is not "we need a queue." The problem is N×M point-to-point coupling in a data platform.

Topics & partitions

message key
4
max consumers
per key
ordering
hash(key) % nproducerpartition 0hash(key)partition 1hash(key)partition 2hash(key)partition 3hash(key)
Safe

Order preserved per key. All records for a key hash to the same partition, so a given account's events stay in sequence while 4 consumers work in parallel. Note that adding partitions later changes the hash mapping — the ordering guarantee spans the change only if you never repartition.

A partition is the unit of both parallelism and ordering, and you cannot have more of one without less of the other. Kafka guarantees order within a partition and says nothing across them.

The problem is not "we need a queue." The problem is N×M point-to-point coupling in a data platform.

The decay sequence (this is the actual origin story)

Stage 1 — one producer, one consumer. You have an app emitting metrics and a dashboard. You open a direct TCP connection. This works. It is the correct engineering decision at this size.

metricsAppDashboard

A direct TCP connection. “This works. It is the correct engineering decision at this size.”

Figure 1.1.1The decay sequence (this is the actual origin story)

Stage 2 — requirements multiply. You want long-term metric analysis, so you add a storage/analysis service. Now the app writes to two places. Then 3 more apps start emitting metrics. Then a coworker wants pull-based alerting, so every app also runs an HTTP metrics server. Then more consumers appear pulling from those servers.

App1App2App3App4DashboardStorageAlertingReportsEach app also runs a metrics HTTP endpoint for pollers.Connections nobody can draw on a whiteboard anymore.

O(N×M) connections. Every new consumer requires touching every producer, and because every producer knows every consumer, adding a consumer is a deployment of the producer fleet.

Figure 1.1.2

The cost here is real and specific, not aesthetic:

  • O(N×M) connections. Every new consumer requires touching every producer.
  • Every producer knows every consumer. Adding a consumer is a deployment of the producer fleet.
  • Mixed push and pull. Some paths push, some poll. Different latency, different failure modes, different code.
  • Format drift. Each pair negotiated its own format. No two are the same.

Stage 3 — you build a broker. You put a single service in the middle that accepts metrics from everyone and serves them to anyone.

App1App2App3App4metrics pub/subDashboardStorageAlertingReportsN + M connections, not N × M
Figure 1.1.3

Congratulations — you have built a publish/subscribe messaging system. This is the point the book makes bluntly: pub/sub is not a technology you choose, it's the shape every sufficiently-grown data platform converges on. The only question is whether you build it deliberately or accidentally.

Stage 4 — the real problem. Meanwhile, a coworker independently built the same thing for log messages. Another built it for user activity tracking. You now have three separate pub/sub systems, each with its own bugs, its own operational runbook, its own scaling limits, and its own on-call rotation.

Three separate pub/sub systems, each with its own bugs, its own operational runbook, its own scaling limits, and its own on-call rotation:

  • metrics pub/sub
  • logging pub/sub
  • user-activity pub/sub

3 systems · 3 sets of bugs · 3 scaling stories · 0 shared tooling — and you know a 4th use case is coming next quarter.

This is the problem Kafka solves: not "move messages," but one general-purpose, horizontally scalable, durable data backbone that any type of data can flow through, so the organization stops re-solving pub/sub per data type.

The mental model that follows

Kafka is described two ways in the book, and both matter:

  1. "A distributed commit log." A DB or filesystem commit log is a durable, ordered record of every transaction, so state can be rebuilt deterministically by replay. Kafka is that, as a service.
  2. "A distributed streaming platform." Same substrate, framed as continuous data-in-motion rather than storage.

The commit-log framing is the load-bearing one. It explains every property Kafka has:

Because it's a log…You get…
Appends are sequentialExtremely high write throughput on cheap disks
Reads are positional (offset)Many independent readers, no per-consumer state on the broker
The log is retained, not consumedReplay, reprocessing, late consumers, backfill
Log is deterministic and orderedState can be rebuilt; stream processing is well-defined
Log can be sharded and copiedHorizontal scale + redundancy

On this page