1.2 Why wasn't something else enough?
This is answered concretely by LinkedIn's history, which is the honest version of the story.
This is answered concretely by LinkedIn's history, which is the honest version of the story.
What LinkedIn had before Kafka
System A — monitoring/metrics. Custom collectors + open-source storage/presentation. Its faults:
- Poll-based collection → you get data at the poller's convenience, not the event's.
- Large intervals between metrics → low resolution, blind spots.
- No self-service — application owners could not manage their own metrics.
- High-touch — human intervention for routine tasks.
- Inconsistent — the same measurement had different metric names in different systems.
System B — user activity tracking. Frontends periodically POSTed batches of XML to an HTTP service; batches were moved to offline platforms and parsed there. Its faults:
- XML parsing was computationally expensive, and the formatting was inconsistent.
- Schema changes broke it constantly.
- Changing what you tracked required coordinated frontend + offline work — a cross-team project for a new field.
- Hourly batching → structurally incapable of real time.
Why they couldn't just merge the two
This is the crux, and it's the best "why not something else" argument in the book:
| Monitoring system | Tracking system | |
|---|---|---|
| Model | Pull (polling) | Push (frontends post) |
| Cadence | Interval-based | Hourly batches |
| Data shape | Metrics-oriented | Activity/event-oriented |
| Robustness | Clunky but tolerable for metrics | Too fragile for metrics |
| Real-time? | Sort of | No — batch by construction |
The monitoring backend couldn't take tracking data (wrong data model, pull model incompatible with push). The tracking backend couldn't take metrics (too fragile, batch-oriented — useless for alerting).
And they desperately needed them joined. The high-value question was "how does this specific type of user activity affect application performance?" — a correlation across both datasets. With hours of batch delay, a drop in a user activity type (a strong signal that the serving app is broken) surfaced hours late. The business cost of the architecture was slow incident response.
Why not off-the-shelf?
They did evaluate existing open source properly. ActiveMQ was prototyped and rejected, for two specific reasons:
- It could not handle the scale at the time.
- The brokers would pause. This is the fatal one. When an ActiveMQ broker paused, it backed up the client connections, which interfered with the applications' ability to serve user requests.
Read that failure mode carefully, because it's the single most important operational lesson in the chapter: a telemetry system that can stall its producers has coupled your observability pipeline to your revenue path. Metrics/tracking must be a system that cannot apply backpressure into the serving tier. That requirement — "the broker must absorb, never stall the producer" — drives Kafka's whole design: sequential appends, batching, page-cache reads, disk retention as a buffer instead of memory-bound queues.
The design goals they wrote down
- Decouple producers and consumers using a push-pull model. Producers push; consumers pull at their own rate. This is the direct fix for the ActiveMQ pause problem — a slow consumer cannot stall a producer, because consumption is pull-based off a durable log.
- Persist message data inside the messaging system, to allow multiple consumers. Persistence is not for durability alone; it's what makes many independent consumers possible.
- Optimize for high throughput.
- Allow horizontal scaling as data streams grow.
The result: "an interface typical of messaging systems, but a storage layer more like a log-aggregation system." That sentence is the whole product.
Did it work?
- Combined with Apache Avro for serialization, it handled metrics and activity tracking at billions of messages/day.
- By February 2020, LinkedIn: >7 trillion messages produced and >5 PB consumed daily.
- Open sourced on GitHub late 2010 → Apache incubator July 2011 → graduated October 2012.