Learn Labs
7. Reliable Data Delivery

7.7 What actually breaks in production — Ch. 7 consolidated

Production failure catalog
0 rows
#SymptomRoot causeFix
1Producer got success; message gone. Brokers configured "perfectly."acks=1 + leader crashed before replicating; followers were still nominally in-sync because replica.lag.time.max.ms hadn't elapsedacks=all and min.insync.replicas=2
2acks=all set, message still lostmin.insync.replicas=1 and both followers were down → "all in-sync replicas" meant onemin.insync.replicas=2 — this is the only thing that makes acks=all mean what you think
3acks=all set, still lost the message on "Leader not Available"Producer didn't retry; the broker never received itDefault (infinite) retries + delivery.timeout.ms > measured recovery time
4Total silent loss during a cluster outageacks=0 — "we won't get any error if the partition is offline, a leader election is in progress, or even if the entire Kafka cluster is unavailable"Never acks=0 for data you care about; don't trust acks=0 benchmarks
5Under-replicated partitions climbing while producers are happyacks=1 lets you "write to the leader faster than it can replicate"acks=all; fix replication throughput (NIC — Ch. 2)
6Latency mysteriously improved mid-incidentA limping replica fell out of the ISR — the latency drag vanished, and so did a third of your durabilityAlert on UnderReplicatedPartitions, not just latency. Latency recovery ≠ resolution
7A replica flaps in and out of the ISRLong GC pauses disconnecting the broker from ZooKeeper; historically worsened by large max request size + large heapKafka 2.5.0+ defaults; JVM 8+ with G1; tune for large messages
8A follower fetches continuously yet is declared out of syncCondition ③: it must have had zero lag at least once in the window — steady lag is still out-of-syncFix throughput; understand the definition before "fixing" the config
9RF=3 and a partition still went offline when one switch failedAll three replicas in one rackbroker.rack everywhere; AZs as racks in cloud; audit after reassignments (Ch. 2)
10Partition offline for hours after multiple broker failuresNo in-sync replica; unclean.leader.election.enable=false (correct default)Accept the outage, or deliberately enable unclean election — and turn it back off after recovery
11Same offset returns different messages to different consumers; downstream reports disagreeUnclean leader election — offsets 100–200 were rewritten with new data; the old leader later deleted its divergent messagesThis is the known cost of unclean election. Prefer false + min.insync.replicas=2 so it can't be needed
12Producers suddenly get NotEnoughReplicasExceptionWorking as designed: in-sync count fell below min.insync.replicas → the partition went read-onlyRestore a replica and let it catch up. This is durability protecting you
13Acknowledged data lost in a correlated power eventNo fsync by default — Kafka bets on independent failure domainsReplicas in separate racks/AZs; only consider flush.messages/flush.ms after reading the throughput implications
14Cluster destabilizes in a cloud environment with variable latencyzookeeper.session.timeout.ms too low for the environment2.5.0+ default of 18 s; tune high enough to avoid flapping, low enough to catch frozen brokers
15Consume latency worse after upgrading to 2.5.0replica.lag.time.max.ms 10 s → 30 s raises the ceiling on "until a message arrives to all replicas and consumers are allowed to consume it"Understand the trade; lower it only if you accept more ISR flapping
16A single consumer only sees a fraction of messagesIt shares group.id with another application"it will need a unique group.id"
17Consumer restarts and silently skips everything that arrived while it was downauto.offset.reset=latest with an invalid/expired committed offsetearliest (accept duplicates) or none (fail loudly)
18Records read but never processed after handing them to a thread poolAutocommit committed offsets for records that had only been read"there is no choice but to use manual offset commit" once processing leaves the poll loop
19Offsets committed for the batch even though processing threw mid-batchAutocommit doesn't know what you processedManual commits after processing (Ch. 4)
20One broker is overloaded purely by offset commits"all offset commits of a single consumer group are produced to the same broker", and each commit ≈ produce with acks=allCommit less often; "committing after every message should only ever be done on very low-throughput topics"
21A failed record is silently skippedCommitted offset 31 after record 30 failed — the watermark marks 30 processed tooPattern A (pause() + buffer + retry) or Pattern B (retry topic / DLQ)
22Retry topic reorders eventsPattern B trades ordering for livenessChoose deliberately; Pattern A preserves order but head-of-line blocks
23Stateful consumer resumes at the right offset with the wrong aggregateOffsets and application state were committed separatelyWrite results + offsets atomically (Ch. 8 transactions), or use Kafka Streams / Flink
24Application "reliable" in test, loses data in productionNever tested under hanging disk / high latency / disk full / rolling restartsTrogdor or equivalent; write the expected behavior first, then measure
25delivery.timeout.ms too short; producer gives up during failoverNever measured how long leader election actually takes in your clusterRun the leader-election test with VerifiableProducer/Consumer and use the measured number
26Lag alerts flap constantly; team ignores themLag oscillates by construction — static thresholds don't workTrend/status-based checking (Burrow)
27Messages "lost somewhere" and nobody can prove whereNo end-to-end reconciliation of produced vs consumed countsBuild produce/consume counters + timestamp-based latency; note there's no open-source implementation
28Produce-to-consume latency looks impossibly lowBroker configured for append-time timestamps, overriding create-timeKnow which timestamp type your topics use
29Failed request metrics risingCould be benign (NOT_LEADER_FOR_PARTITION during maintenance) or serious"Unexplained increases should always be investigated" — the metrics are tagged with the specific error
30Hand-rolled retry loop caused duplicates and reorderingApplication-level retries layered on producer retries"if all the error handler is doing is retrying, we'll be better off relying on the producer's retry functionality"