Learn Labs
8. Exactly-Once Semantics

8.8 Deploy / monitor / configure

The configuration recipes

Tier 1 — idempotent producer (do this everywhere; essentially free):

enable.idempotence                      = true
acks                                    = all      (required)
retries                                 > 0        (required)
max.in.flight.requests.per.connection   ≤ 5        (required)
  • Gets you: no retry duplicates + guaranteed ordering with 5 in-flight.
  • Costs: 1 extra startup API call, 96 bits per record batch. “If already configured with acks=all, no performance difference.”

Tier 2 — exactly-once via Kafka Streams (the recommended path):

processing.guarantee = exactly_once
                     | exactly_once_beta   (brokers 2.5+, Streams 2.6+,
                                            better partition scaling)
  • ► “That's it.” Streams manages transactions for you.
  • ► Downstream consumers still need isolation.level=read_committed.

Tier 3 — raw transactional API (only if Streams doesn't fit):

WhereWhat
Producertransactional.id = unique + long-lived (it defines an instance) · initTransactions() on startup · beginTransaction() / send() / sendOffsetsToTransaction() / commitTransaction() · handle ProducerFenced/InvalidProducerEpoch by dying · handle other KafkaException by abort + reset consumer position
Consumer (in the loop)enable.auto.commit = false · never call consumer.commitSync/commitAsync · isolation.level = read_committed · pass consumer.groupMetadata() to sendOffsetsToTransaction() (2.5+) · commit transactions when partitions are revoked
Broker / opstransaction.timeout.ms (default 15 min) — the LSO stall ceiling · transactional.id.expiration.ms (default 7 days) — lower it if you cannot use long-lived producers (FaaS)

Monitoring

SignalMeaning
record-error-rate (producer)Includes benign duplicate rejections — "should not cause any alarm", but a rising trend is worth understanding
ErrorsPerSec of RequestMetrics (broker)"includes a separate count for each type of error" — this is how you separate duplicate rejections from real errors
"out of order sequence number" in producer logs⚠️ Message loss between producer and broker — investigate config + unclean elections
UNKNOWN_PRODUCER_IDProducer state expired or missing (pre-2.5 bug, or expiration too low)
ProducerFencedException rateZombie instances being fenced — expected during failover, alarming if constant
read_committed consumer lag vs read_uncommittedMeasures your LSO stall — i.e. how long your transactions stay open
Broker heap / GCThe producer-state accumulation leak (#22) shows up here first
Transaction commit latencyThe synchronous stop in the produce path

The relationship to Ch. 7

Ch. 7’s recipe: acks=all + min.insync.replicas=2 + RF≥3 + retries + delivery.timeout.ms, and commit offsets after processing.

Ch. 7 gave you AT-LEAST-ONCEretries create DUPLICATESCh. 8 upgrades itidempotence + transactions + read_committedEXACTLY-ONCE
  • · + enable.idempotence=true → no retry duplicates, ordering kept
  • · + transactions (or Streams EOS) → atomic consume-process-produce
  • · + isolation.level=read_committed → downstream sees only committed

EXACTLY-ONCE — but ONLY for records written to Kafka, ONLY within the consume-process-produce pattern, and ONLY if all three parts are configured.

Figure 8.8.2The relationship to Ch. 7

On this page