8. Exactly-Once Semantics
8.7 What actually breaks in production — Ch. 8 consolidated
| # | Symptom | Root cause | Fix |
|---|---|---|---|
| 1 | Duplicates even though nothing failed | Leader wrote and replicated, then crashed before acking; the producer retried to the new leader | enable.idempotence=true |
| 2 | Aggregates (averages, counts, balances) are subtly wrong and nobody can prove why | A duplicate was folded into an aggregate — "impossible to correct the result without reprocessing the input" | Transactions / Kafka Streams processing.guarantee |
| 3 | "Out of order sequence number" in the logs; team ignores it | The producer can continue, but the gap means messages 3–26 were LOST | Review reliability config (Ch. 7); check whether unclean leader election occurred |
| 4 | Duplicates after a producer restart, with idempotence on | Each init gets a brand-new PID — the broker cannot correlate old and new producers | Use transactional.id for a stable identity across restarts |
| 5 | A revived frozen producer wrote duplicates; not detected as a zombie | "we have two totally different producers with different IDs" | Transactions + epoch fencing |
| 6 | Duplicates from two app instances reading the same source file | Idempotence is per producer instance, not a global dedup service | Partition the work; dedupe by business key upstream |
| 7 | Duplicates from your own retry loop | producer.send() called twice — "the producer has no way of knowing that the two records are in fact the same" | Rely on the built-in retry mechanism only |
| 8 | Fatal UNKNOWN_PRODUCER_ID | Pre-2.5 producer state not kept long enough; known partition-reassignment edge case where a new leader had no state | Upgrade to 2.5+ (KIP-360) |
| 9 | Exactly-once configured, consumers still see aborted records | isolation.level defaults to read_uncommitted — aborted records are physically in the log | isolation.level=read_committed on every downstream consumer |
| 10 | read_committed consumers stall for up to 15 minutes | A long-running open transaction holds back the LSO; everything after it is withheld until commit/abort or transaction.timeout.ms | Commit transactions frequently; tune transaction.timeout.ms |
| 11 | End-to-end latency worse after enabling transactions | Same LSO mechanism — read_committed consumers always lag | Treat transaction duration as a latency budget |
| 12 | Duplicate emails / double API charges despite exactly-once | "The guarantee only applies to records written to Kafka... it will not un-send an email" | Make external effects idempotent (idempotency keys), or move them outside the transaction |
| 13 | Can't atomically write to a DB and commit Kafka offsets | "There is no mechanism" — no producer is involved | Manage offsets in the database; use the DB's transaction |
| 14 | Microservice updated the DB but the Kafka message was lost (or vice versa) | Expecting Kafka transactions to span systems | Outbox pattern (Kafka-as-outbox with an idempotent DB update, or table-as-outbox when you need RDBMS constraints) |
| 15 | Source database transactions not preserved through Kafka | Consumers have no transaction boundary information and may be lagging on some topics | Not solvable with Kafka transactions; redesign |
| 16 | MirrorMaker copied records exactly once but transactions lost atomicity | Cross-cluster copy can't guarantee it sees all events in a transaction — "it can replicate part of a transaction if it is only subscribed to a subset of the topics" | Accept per-record exactly-once only |
| 17 | Pub/sub consumers still process messages twice | Transactions don't govern consumer offset commit logic | You still need consumer-side discipline (Ch. 7 §5) |
| 18 | Deadlock: producer waits for a reply that can never arrive | Published inside a transaction, then waited for a read_committed consumer to respond before committing | Never block on a response before committing the transaction |
| 19 | Zombie not fenced; duplicates in the output | Transactional ID changed between the failed instance and its replacement (A → B) | Pre-2.5: statically map transactional ID → partitions. 2.5+: pass consumer.groupMetadata() to sendOffsetsToTransaction() |
| 20 | ProducerFencedException / InvalidProducerEpochException | You are the zombie. A newer instance holds your transactional ID | "Nothing to do but die gracefully" — don't retry |
| 21 | Offsets committed outside the transaction; exactly-once silently broken | Called consumer.commitSync(), or left autocommit on | enable.auto.commit=false and never call consumer commit APIs; offsets only via sendOffsetsToTransaction() |
| 22 | Broker OOM / severe GC after weeks of normal operation | Producer-state accumulation — new transactional/producer IDs created at a high rate, retained transactional.id.expiration.ms (7 days). 3/sec ⇒ 1.8M entries ⇒ ~5 GB | Few long-lived producers; if FaaS makes that impossible, lower transactional.id.expiration.ms |
| 23 | Transaction throughput poor | Very small transactions — overhead is per transaction, and init/commit are synchronous stops | Batch more messages per transaction |
| 24 | In-flight transactions left hanging after a crash | Normal — resolved by design | initTransactions() aborts older in-flight transactions; the coordinator auto-aborts after transaction.timeout.ms; a new coordinator picks up logged intent |
| 25 | Rebalance broke transactional processing (pre-2.5) | Transactional producers needed static partition assignment; subscribe() reassigns freely | Kafka 2.5+ with group-metadata fencing; "commit transactions whenever the related partitions are revoked" |
| 26 | Kafka Streams app doesn't scale with many partitions | One transactional producer per partition-identity was required | processing.guarantee=exactly_once_beta (brokers 2.5+, Streams 2.6+) |