7. Reliable Data Delivery
7.8 The reliability configuration matrix
Maximum reliability — financial transactions, “never lose a message”
| Layer | Setting | Value |
|---|---|---|
| Broker / topic | replication.factor | 3 (or 4 = RF++, Ch. 2) |
min.insync.replicas | 2 ← makes acks=all meaningful | |
unclean.leader.election.enable | false | |
broker.rack | set, replicas across racks/AZs | |
zookeeper.session.timeout.ms | 18000 (2.5.0 default) | |
replica.lag.time.max.ms | 30000 (2.5.0 default) | |
| Producer | acks | all |
retries | default (~MAX_INT) | |
delivery.timeout.ms | > measured cluster recovery time | |
enable.idempotence | true ← at-least-once → exactly-once | |
| error handling | real handlers for non-retriable errors, serialization errors, exhausted retries, buffer exhaustion, timeouts | |
| Consumer | group.id | unique per application |
auto.offset.reset | earliest (or none to fail loudly) | |
enable.auto.commit | false | |
| offset commits | after processing; commit in onPartitionsRevoked() | |
| retry strategy | pause() + buffer, or a retry topic | |
| state | transactional (Ch. 8) or Kafka Streams / Flink |
Relaxed — click tracking, the “customer complaints” topic
| Setting | Value |
|---|---|
replication.factor | 2 — cheaper; durability from the storage layer, but lower availability than 3 |
min.insync.replicas | 1 (default — knowingly loss-tolerant) |
acks | 1 |
enable.auto.commit | true, all processing inside the poll loop |
auto.offset.reset | latest |
Same cluster. Per-topic overrides. That’s the point of topic-level configuration.
The dependency graph — why these settings only work together
Break any link and the chain provides no guarantee. This is what “reliability is a property of a system” means concretely.
Monitoring summary
| Layer | Signal | Why |
|---|---|---|
| Broker | UnderReplicatedPartitions | Effective RF dropped — the #6 trap |
| Broker | OfflinePartitions | Unavailable data |
| Broker | FailedProduceRequestsPerSec / FailedFetchRequestsPerSec (tagged by error) | Distinguishes benign NOT_LEADER_FOR_PARTITION from NOT_ENOUGH_REPLICAS |
| Broker | ISR shrink/expand rate | Flapping = GC/network trouble |
| Producer | error-rate, retry-rate per record | "the two metrics most important for reliability" |
| Producer | WARN logs with "0 attempts left" | Out of retries |
| Producer | ERROR logs | Complete failure: non-retriable, exhausted, or timeout |
| Consumer | Consumer lag (trend, not threshold) | Use Burrow |
| End-to-end | produced/sec vs consumed/sec reconciliation | The only way to prove nothing was lost |
| End-to-end | produce-timestamp → consume-timestamp latency | Verifies "timely" against business requirements |