7. Reliable Data Delivery
7.7 What actually breaks in production — Ch. 7 consolidated
| # | Symptom | Root cause | Fix |
|---|---|---|---|
| 1 | Producer got success; message gone. Brokers configured "perfectly." | acks=1 + leader crashed before replicating; followers were still nominally in-sync because replica.lag.time.max.ms hadn't elapsed | acks=all and min.insync.replicas=2 |
| 2 | acks=all set, message still lost | min.insync.replicas=1 and both followers were down → "all in-sync replicas" meant one | min.insync.replicas=2 — this is the only thing that makes acks=all mean what you think |
| 3 | acks=all set, still lost the message on "Leader not Available" | Producer didn't retry; the broker never received it | Default (infinite) retries + delivery.timeout.ms > measured recovery time |
| 4 | Total silent loss during a cluster outage | acks=0 — "we won't get any error if the partition is offline, a leader election is in progress, or even if the entire Kafka cluster is unavailable" | Never acks=0 for data you care about; don't trust acks=0 benchmarks |
| 5 | Under-replicated partitions climbing while producers are happy | acks=1 lets you "write to the leader faster than it can replicate" | acks=all; fix replication throughput (NIC — Ch. 2) |
| 6 | Latency mysteriously improved mid-incident | A limping replica fell out of the ISR — the latency drag vanished, and so did a third of your durability | Alert on UnderReplicatedPartitions, not just latency. Latency recovery ≠ resolution |
| 7 | A replica flaps in and out of the ISR | Long GC pauses disconnecting the broker from ZooKeeper; historically worsened by large max request size + large heap | Kafka 2.5.0+ defaults; JVM 8+ with G1; tune for large messages |
| 8 | A follower fetches continuously yet is declared out of sync | Condition ③: it must have had zero lag at least once in the window — steady lag is still out-of-sync | Fix throughput; understand the definition before "fixing" the config |
| 9 | RF=3 and a partition still went offline when one switch failed | All three replicas in one rack | broker.rack everywhere; AZs as racks in cloud; audit after reassignments (Ch. 2) |
| 10 | Partition offline for hours after multiple broker failures | No in-sync replica; unclean.leader.election.enable=false (correct default) | Accept the outage, or deliberately enable unclean election — and turn it back off after recovery |
| 11 | Same offset returns different messages to different consumers; downstream reports disagree | Unclean leader election — offsets 100–200 were rewritten with new data; the old leader later deleted its divergent messages | This is the known cost of unclean election. Prefer false + min.insync.replicas=2 so it can't be needed |
| 12 | Producers suddenly get NotEnoughReplicasException | Working as designed: in-sync count fell below min.insync.replicas → the partition went read-only | Restore a replica and let it catch up. This is durability protecting you |
| 13 | Acknowledged data lost in a correlated power event | No fsync by default — Kafka bets on independent failure domains | Replicas in separate racks/AZs; only consider flush.messages/flush.ms after reading the throughput implications |
| 14 | Cluster destabilizes in a cloud environment with variable latency | zookeeper.session.timeout.ms too low for the environment | 2.5.0+ default of 18 s; tune high enough to avoid flapping, low enough to catch frozen brokers |
| 15 | Consume latency worse after upgrading to 2.5.0 | replica.lag.time.max.ms 10 s → 30 s raises the ceiling on "until a message arrives to all replicas and consumers are allowed to consume it" | Understand the trade; lower it only if you accept more ISR flapping |
| 16 | A single consumer only sees a fraction of messages | It shares group.id with another application | "it will need a unique group.id" |
| 17 | Consumer restarts and silently skips everything that arrived while it was down | auto.offset.reset=latest with an invalid/expired committed offset | earliest (accept duplicates) or none (fail loudly) |
| 18 | Records read but never processed after handing them to a thread pool | Autocommit committed offsets for records that had only been read | "there is no choice but to use manual offset commit" once processing leaves the poll loop |
| 19 | Offsets committed for the batch even though processing threw mid-batch | Autocommit doesn't know what you processed | Manual commits after processing (Ch. 4) |
| 20 | One broker is overloaded purely by offset commits | "all offset commits of a single consumer group are produced to the same broker", and each commit ≈ produce with acks=all | Commit less often; "committing after every message should only ever be done on very low-throughput topics" |
| 21 | A failed record is silently skipped | Committed offset 31 after record 30 failed — the watermark marks 30 processed too | Pattern A (pause() + buffer + retry) or Pattern B (retry topic / DLQ) |
| 22 | Retry topic reorders events | Pattern B trades ordering for liveness | Choose deliberately; Pattern A preserves order but head-of-line blocks |
| 23 | Stateful consumer resumes at the right offset with the wrong aggregate | Offsets and application state were committed separately | Write results + offsets atomically (Ch. 8 transactions), or use Kafka Streams / Flink |
| 24 | Application "reliable" in test, loses data in production | Never tested under hanging disk / high latency / disk full / rolling restarts | Trogdor or equivalent; write the expected behavior first, then measure |
| 25 | delivery.timeout.ms too short; producer gives up during failover | Never measured how long leader election actually takes in your cluster | Run the leader-election test with VerifiableProducer/Consumer and use the measured number |
| 26 | Lag alerts flap constantly; team ignores them | Lag oscillates by construction — static thresholds don't work | Trend/status-based checking (Burrow) |
| 27 | Messages "lost somewhere" and nobody can prove where | No end-to-end reconciliation of produced vs consumed counts | Build produce/consume counters + timestamp-based latency; note there's no open-source implementation |
| 28 | Produce-to-consume latency looks impossibly low | Broker configured for append-time timestamps, overriding create-time | Know which timestamp type your topics use |
| 29 | Failed request metrics rising | Could be benign (NOT_LEADER_FOR_PARTITION during maintenance) or serious | "Unexplained increases should always be investigated" — the metrics are tagged with the specific error |
| 30 | Hand-rolled retry loop caused duplicates and reordering | Application-level retries layered on producer retries | "if all the error handler is doing is retrying, we'll be better off relying on the producer's retry functionality" |