8. Exactly-Once Semantics
8.8 Deploy / monitor / configure
The configuration recipes
Tier 1 — idempotent producer (do this everywhere; essentially free):
enable.idempotence = true
acks = all (required)
retries > 0 (required)
max.in.flight.requests.per.connection ≤ 5 (required)- Gets you: no retry duplicates + guaranteed ordering with 5 in-flight.
- Costs: 1 extra startup API call, 96 bits per record batch. “If already configured with
acks=all, no performance difference.”
Tier 2 — exactly-once via Kafka Streams (the recommended path):
processing.guarantee = exactly_once
| exactly_once_beta (brokers 2.5+, Streams 2.6+,
better partition scaling)- ► “That's it.” Streams manages transactions for you.
- ► Downstream consumers still need
isolation.level=read_committed.
Tier 3 — raw transactional API (only if Streams doesn't fit):
| Where | What |
|---|---|
| Producer | transactional.id = unique + long-lived (it defines an instance) · initTransactions() on startup · beginTransaction() / send() / sendOffsetsToTransaction() / commitTransaction() · handle ProducerFenced/InvalidProducerEpoch by dying · handle other KafkaException by abort + reset consumer position |
| Consumer (in the loop) | enable.auto.commit = false · never call consumer.commitSync/commitAsync · isolation.level = read_committed · pass consumer.groupMetadata() to sendOffsetsToTransaction() (2.5+) · commit transactions when partitions are revoked |
| Broker / ops | transaction.timeout.ms (default 15 min) — the LSO stall ceiling · transactional.id.expiration.ms (default 7 days) — lower it if you cannot use long-lived producers (FaaS) |
Monitoring
| Signal | Meaning |
|---|---|
record-error-rate (producer) | Includes benign duplicate rejections — "should not cause any alarm", but a rising trend is worth understanding |
ErrorsPerSec of RequestMetrics (broker) | "includes a separate count for each type of error" — this is how you separate duplicate rejections from real errors |
| "out of order sequence number" in producer logs | ⚠️ Message loss between producer and broker — investigate config + unclean elections |
UNKNOWN_PRODUCER_ID | Producer state expired or missing (pre-2.5 bug, or expiration too low) |
ProducerFencedException rate | Zombie instances being fenced — expected during failover, alarming if constant |
| read_committed consumer lag vs read_uncommitted | Measures your LSO stall — i.e. how long your transactions stay open |
| Broker heap / GC | The producer-state accumulation leak (#22) shows up here first |
| Transaction commit latency | The synchronous stop in the produce path |
The relationship to Ch. 7
Ch. 7’s recipe: acks=all + min.insync.replicas=2 + RF≥3 + retries + delivery.timeout.ms, and commit offsets after processing.
- ·
+ enable.idempotence=true→ no retry duplicates, ordering kept - ·
+ transactions (or Streams EOS)→ atomic consume-process-produce - ·
+ isolation.level=read_committed→ downstream sees only committed
EXACTLY-ONCE — but ONLY for records written to Kafka, ONLY within the consume-process-produce pattern, and ONLY if all three parts are configured.