1.4 Why Kafka? (the property-by-property case)
Note the framing on platform features: these are not full platforms.
| Property | What it means | Why it matters |
|---|---|---|
| Multiple producers | Seamlessly handles many producers, many topics or the same topic | Many microservices write page views to one topic in a common format; consumers get a single unified stream instead of N topics to correlate |
| Multiple consumers | Many consumers read the same stream without interfering with each other | Explicitly contrasted with queues where once consumed, a message is gone. Consumers can also form a group to process each message once |
| Disk-based retention | Messages on disk with configurable, per-topic retention | Consumers need not work in real time. A slow consumer or a traffic burst → no data loss. Take a consumer offline for maintenance → producers don't back up, nothing is lost, restart resumes where it left off |
| Scalable | 1 broker (PoC) → 3 (dev) → tens/hundreds (prod). Expansions performed online with no availability impact | Also: multi-broker clusters survive individual broker failure; raise replication factor to tolerate more simultaneous failures |
| High performance | Producers, consumers, and brokers all scale out | Subsecond latency from produce to consumer availability, under very large message streams |
| Platform features | Kafka Connect (source→Kafka, Kafka→sink) and Kafka Streams (scalable, fault-tolerant stream processing) | Deliberately APIs and libraries, not a structured runtime like YARN — "a solid foundation to build on and flexibility as to where they can be run" |
Note the framing on platform features: these are not full platforms. That's a design stance — Kafka gives you libraries you can run under whatever scheduler you already have (k8s, ECS, bare metal) instead of imposing a cluster manager.
The ecosystem framing
Coupled with a message-schema system, producers and consumers need no tight coupling and no direct connections of any sort. Components can be added and removed as business cases come and go, and producers need not know who consumes the data or how many consumers exist.
Coupled with a message-schema system, producers and consumers need no tight coupling and no direct connections of any sort. Components can be added and removed as business cases come and go, and producers need not know who consumes the data or how many consumers exist.
1.3 How does it work internally? (the core object model)
A message is the unit of data — think a DB row or record.
1.5 Use cases (with the "why Kafka specifically" for each)
The win: avoids duplicating this logic in every app, and enables aggregation that would not otherwise be possible (you can't batch a user's notifications if each app sends its own…