Learn Labs
8. Exactly-Once Semantics

8.7 What actually breaks in production — Ch. 8 consolidated

Production failure catalog
0 rows
#SymptomRoot causeFix
1Duplicates even though nothing failedLeader wrote and replicated, then crashed before acking; the producer retried to the new leaderenable.idempotence=true
2Aggregates (averages, counts, balances) are subtly wrong and nobody can prove whyA duplicate was folded into an aggregate — "impossible to correct the result without reprocessing the input"Transactions / Kafka Streams processing.guarantee
3"Out of order sequence number" in the logs; team ignores itThe producer can continue, but the gap means messages 3–26 were LOSTReview reliability config (Ch. 7); check whether unclean leader election occurred
4Duplicates after a producer restart, with idempotence onEach init gets a brand-new PID — the broker cannot correlate old and new producersUse transactional.id for a stable identity across restarts
5A revived frozen producer wrote duplicates; not detected as a zombie"we have two totally different producers with different IDs"Transactions + epoch fencing
6Duplicates from two app instances reading the same source fileIdempotence is per producer instance, not a global dedup servicePartition the work; dedupe by business key upstream
7Duplicates from your own retry loopproducer.send() called twice — "the producer has no way of knowing that the two records are in fact the same"Rely on the built-in retry mechanism only
8Fatal UNKNOWN_PRODUCER_IDPre-2.5 producer state not kept long enough; known partition-reassignment edge case where a new leader had no stateUpgrade to 2.5+ (KIP-360)
9Exactly-once configured, consumers still see aborted recordsisolation.level defaults to read_uncommitted — aborted records are physically in the logisolation.level=read_committed on every downstream consumer
10read_committed consumers stall for up to 15 minutesA long-running open transaction holds back the LSO; everything after it is withheld until commit/abort or transaction.timeout.msCommit transactions frequently; tune transaction.timeout.ms
11End-to-end latency worse after enabling transactionsSame LSO mechanism — read_committed consumers always lagTreat transaction duration as a latency budget
12Duplicate emails / double API charges despite exactly-once"The guarantee only applies to records written to Kafka... it will not un-send an email"Make external effects idempotent (idempotency keys), or move them outside the transaction
13Can't atomically write to a DB and commit Kafka offsets"There is no mechanism" — no producer is involvedManage offsets in the database; use the DB's transaction
14Microservice updated the DB but the Kafka message was lost (or vice versa)Expecting Kafka transactions to span systemsOutbox pattern (Kafka-as-outbox with an idempotent DB update, or table-as-outbox when you need RDBMS constraints)
15Source database transactions not preserved through KafkaConsumers have no transaction boundary information and may be lagging on some topicsNot solvable with Kafka transactions; redesign
16MirrorMaker copied records exactly once but transactions lost atomicityCross-cluster copy can't guarantee it sees all events in a transaction — "it can replicate part of a transaction if it is only subscribed to a subset of the topics"Accept per-record exactly-once only
17Pub/sub consumers still process messages twiceTransactions don't govern consumer offset commit logicYou still need consumer-side discipline (Ch. 7 §5)
18Deadlock: producer waits for a reply that can never arrivePublished inside a transaction, then waited for a read_committed consumer to respond before committingNever block on a response before committing the transaction
19Zombie not fenced; duplicates in the outputTransactional ID changed between the failed instance and its replacement (A → B)Pre-2.5: statically map transactional ID → partitions. 2.5+: pass consumer.groupMetadata() to sendOffsetsToTransaction()
20ProducerFencedException / InvalidProducerEpochExceptionYou are the zombie. A newer instance holds your transactional ID"Nothing to do but die gracefully" — don't retry
21Offsets committed outside the transaction; exactly-once silently brokenCalled consumer.commitSync(), or left autocommit onenable.auto.commit=false and never call consumer commit APIs; offsets only via sendOffsetsToTransaction()
22Broker OOM / severe GC after weeks of normal operationProducer-state accumulation — new transactional/producer IDs created at a high rate, retained transactional.id.expiration.ms (7 days). 3/sec ⇒ 1.8M entries ⇒ ~5 GBFew long-lived producers; if FaaS makes that impossible, lower transactional.id.expiration.ms
23Transaction throughput poorVery small transactions — overhead is per transaction, and init/commit are synchronous stopsBatch more messages per transaction
24In-flight transactions left hanging after a crashNormal — resolved by designinitTransactions() aborts older in-flight transactions; the coordinator auto-aborts after transaction.timeout.ms; a new coordinator picks up logged intent
25Rebalance broke transactional processing (pre-2.5)Transactional producers needed static partition assignment; subscribe() reassigns freelyKafka 2.5+ with group-metadata fencing; "commit transactions whenever the related partitions are revoked"
26Kafka Streams app doesn't scale with many partitionsOne transactional producer per partition-identity was requiredprocessing.guarantee=exactly_once_beta (brokers 2.5+, Streams 2.6+)