Learn Labs
7. Reliable Data Delivery

7.8 The reliability configuration matrix

Maximum reliability — financial transactions, “never lose a message”

LayerSettingValue
Broker / topicreplication.factor3 (or 4 = RF++, Ch. 2)
min.insync.replicas2 ← makes acks=all meaningful
unclean.leader.election.enablefalse
broker.rackset, replicas across racks/AZs
zookeeper.session.timeout.ms18000 (2.5.0 default)
replica.lag.time.max.ms30000 (2.5.0 default)
Produceracksall
retriesdefault (~MAX_INT)
delivery.timeout.ms> measured cluster recovery time
enable.idempotencetrue ← at-least-once → exactly-once
error handlingreal handlers for non-retriable errors, serialization errors, exhausted retries, buffer exhaustion, timeouts
Consumergroup.idunique per application
auto.offset.resetearliest (or none to fail loudly)
enable.auto.commitfalse
offset commitsafter processing; commit in onPartitionsRevoked()
retry strategypause() + buffer, or a retry topic
statetransactional (Ch. 8) or Kafka Streams / Flink

Relaxed — click tracking, the “customer complaints” topic

SettingValue
replication.factor2 — cheaper; durability from the storage layer, but lower availability than 3
min.insync.replicas1 (default — knowingly loss-tolerant)
acks1
enable.auto.committrue, all processing inside the poll loop
auto.offset.resetlatest

Same cluster. Per-topic overrides. That’s the point of topic-level configuration.

The dependency graph — why these settings only work together

gives you replicas to be in-syncmakes “all in-sync replicas” mean MORE THAN ONEa rejected/failed write must be retriedretries create duplicatesdelivered ≠ processedprocessing state must match offsetsreplication.factor = 3min.insync.replicas = 2producer acks = allretries ~∞ + delivery.timeout.msgreater than the measured recovery timeenable.idempotence = truecommit offsets AFTER processingmanualtransactions (Ch. 8) /Kafka Streams (Ch. 14)broker.rackindependent failure domains

Break any link and the chain provides no guarantee. This is what “reliability is a property of a system” means concretely.

Figure 7.8.2The dependency graph — why these settings only work together

Monitoring summary

LayerSignalWhy
BrokerUnderReplicatedPartitionsEffective RF dropped — the #6 trap
BrokerOfflinePartitionsUnavailable data
BrokerFailedProduceRequestsPerSec / FailedFetchRequestsPerSec (tagged by error)Distinguishes benign NOT_LEADER_FOR_PARTITION from NOT_ENOUGH_REPLICAS
BrokerISR shrink/expand rateFlapping = GC/network trouble
Producererror-rate, retry-rate per record"the two metrics most important for reliability"
ProducerWARN logs with "0 attempts left"Out of retries
ProducerERROR logsComplete failure: non-retriable, exhausted, or timeout
ConsumerConsumer lag (trend, not threshold)Use Burrow
End-to-endproduced/sec vs consumed/sec reconciliationThe only way to prove nothing was lost
End-to-endproduce-timestamp → consume-timestamp latencyVerifies "timely" against business requirements

On this page