7. Reliable Data Delivery
Chapter 7 of Kafka: The Definitive Guide — 9 sections.
Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 7 The thesis, stated in the first sentence: "Reliability is a property of a system — not of a single component... the systems that integrate with Kafka are as important as Kafka itself. And because reliability is a system concern, it cannot be the responsibility of just one person. Everyone — Kafka administrators, Linux administrators, network and storage administrators, and the application developers — must work together to build a reliable system."
And the warning that makes this chapter necessary:
"Kafka was written to be configurable enough, and its client API flexible enough, to allow all kinds of reliability trade-offs. Because of its flexibility, it is also easy to accidentally shoot ourselves in the foot when using Kafka — believing that our system is reliable when in fact it is not."
Sections
- 7.12Kafka's reliability guarantees — the exact contractThe book frames this by analogy to ACID: "Those guarantees are the reason people trust relational databases with their most critical applications — they know exactly what the syst…
- 7.22Replication recap — and what "in sync" precisely meansA replica is in sync if it is the leader, or if it is a follower that:
- 7.311Broker configuration — three knobsreplication.factor (topic) / default.replication.factor (broker, for auto-created topics).
- 7.44Using producers reliablyThat parenthetical is worth quoting to anyone waving a Kafka benchmark at you.
- 7.53Using consumers reliablyThe reliability-first default is earliest.
- 7.66Validating system reliabilityNote the shape of that expectation: a quantified duplicate budget.
- 7.7Failure catalogWhat actually breaks in production — Ch. 7 consolidated
- 7.82The reliability configuration matrix
- 7.9Self-testSelf-test