8.2 The idempotent producer
The setup: producer writes to topic A partition 0; leader on broker 5, follower on broker 3.
Idempotent producer
- 4
- written
- 3
- duplicates
3 duplicates in the log. Every retry of an already-committed batch appends again. Note that acks=all does not help here: the write did succeed — it is the acknowledgement that was lost.
2.1 What "idempotent" means
"A service is called idempotent if performing the same operation multiple times has the same result as performing it a single time."
The database illustration:
NOT idempotent: UPDATE t SET x = x+1 WHERE y=5
IDEMPOTENT: UPDATE t SET x = 18 WHERE y=5x = x+1— call it 3 times, get a very different result.x = 18— “no matter how many times we run this statement, x will be equal to 18.”
2.2 The classic duplicate — nothing actually failed
Steps ① and ② mean the write SUCCEEDED and is DURABLE. ⑤ The resent message arrives at the new leader — “who ALREADY HAS a copy of the message from the previous attempt.” ► DUPLICATE.
The consequences, in the book's own concrete terms: "in others they can lead to inventory miscounts, bad financial statements, or sending someone two umbrellas instead of the one they ordered."
Why retry logic can never solve this alone: the producer cannot distinguish "the write failed" from "the write succeeded and the ack was lost." Only the broker can, because only the broker knows what it already has. Hence the fix lives at the protocol level.
2.3 How it works — PID + sequence number
“To limit the number of previous sequence numbers that have to be tracked for each partition, we also require that producers use max.inflight.requests = 5 OR LOWER (the default is 5).”
That's why the 5-in-flight limit exists. It isn't arbitrary — it's the size of the broker's dedup window.
Duplicate detection:
"When a broker receives a message that it already accepted before, it will reject the duplicate with an appropriate error. This error is logged by the producer and is reflected in its metrics but DOES NOT CAUSE ANY EXCEPTION and SHOULD NOT CAUSE ANY ALARM."
| Where | Metric |
|---|---|
| Producer client | added to record-error-rate |
| Broker | part of ErrorsPerSec of the RequestMetrics type — "which includes a separate count for each type of error" |
2.4 ⚠️ "Out of order sequence number" — an error you should investigate, not ignore
"The broker expects message number 2 to be followed by message number 3; what happens if the broker receives message number 27 instead? In such cases the broker will respond with an 'out of order sequence' error, but if we use an idempotent producer without using transactions, this error CAN BE IGNORED."
⚠️ WARNING — but it's telling you something
*"While the producer will continue normally after encountering an 'out of order sequence number' exception, this error typically indicates that MESSAGES WERE LOST between the producer and the broker — if the broker received message number 2 followed by message number 27, something must have happened to messages 3 to 26.
When encountering such an error in the logs, it is worth:
- revisiting the producer and topic configuration and making sure the producer is configured with recommended values for high reliability, and
- checking whether UNCLEAN LEADER ELECTION has occurred."*
- · “can be ignored” (the producer keeps working) ≠ “is benign” (you just silently lost 24 messages)
- · Two likely causes: weak reliability config (Ch. 7 §4), or unclean leader election truncating the log (Ch. 7 §3.2).
2.5 Behavior under failure — the two cases
Case A: Producer restart — ⚠️ idempotence does NOT survive it
"when the producer starts, if the idempotent producer is enabled, the producer will initialize and reach out to a Kafka broker to generate a producer ID. EACH INITIALIZATION OF A PRODUCER WILL RESULT IN A COMPLETELY NEW ID (assuming that we did not enable transactions)."
- · The new producer resends a message the old one already sent — different PID, different sequence number — so “the broker will NOT DETECT the duplicates; the two messages will be CONSIDERED AS TWO DIFFERENT MESSAGES.”
- · Same for a frozen producer that revives after its replacement started: “the original producer is NOT RECOGNIZED AS A ZOMBIE, because we have two totally different producers with different IDs.”
This is precisely the gap transactional.id fills (§3.4): a stable identity across restarts, which a randomly-assigned PID can never provide.
Case B: Broker failure — idempotence does survive it
The setup: producer writes to topic A partition 0; leader on broker 5, follower on broker 3. Broker 5 fails, broker 3 becomes leader. "But how will broker 3 know which sequences were already produced in order to reject duplicates?"
Four layers of producer-state durability:
- In-memory, leader
“The leader keeps updating its in-memory producer state with the Five last sequence IDs every time a new message is produced.”
- In-memory, followers
“Follower replicas update their Own in-memory buffers every time they replicate new messages from the leader.”
► “when a follower becomes a leader, It already has the latest sequence numbers in memory, and validation of newly produced messages can continue Without any issues or delays.”
- Snapshot file (for a restarted broker)
“brokers take a Snapshot of the producer state to a file when they Shut down or Every time a segment is created. When the broker starts, it reads the latest state from a file.”
Then it “keeps updating the producer state as it catches up by replicating from the current leader, and it has the most current sequence IDs in memory when it is ready to become a leader again.”
- The log itself (if the snapshot is stale after a crash)
“Producer ID and sequence ID are Also part of the message format that is written to Kafka’s logs. During crash recovery, the producer state will be recovered by reading The older snapshot and Also messages from the latest segment of each partition. A new snapshot will be stored as soon as the recovery process completes.”
(④ is why Ch. 6's batch header carries producer ID, producer epoch, and first sequence — the log is the ultimate source of truth for dedup state.)
The edge case with a satisfying answer:
"what happens if there are no messages? Imagine that a topic has two hours of retention, but no new messages arrived in the last two hours — there will be no messages to use to recover the state if a broker crashed. Luckily, NO MESSAGES ALSO MEANS NO DUPLICATES. We will start accepting messages immediately (while logging a warning about the lack of state), and create the producer state from the new messages that arrive."
2.6 ⚠️ Limitations — exactly what idempotence does NOT cover
“The idempotent producer will only prevent duplicates caused by THE RETRY MECHANISM OF THE PRODUCER ITSELF, whether the retry is caused by producer, network, or broker errors. BUT NOTHING ELSE.”
Limitation 1: Calling send() twice yourself
"Calling
producer.send()twice with the same message will create a duplicate, and the idempotent producer won't prevent it. This is because the producer has no way of knowing that the two records that were sent are in fact the same record."
"It is always a good idea to use the built-in retry mechanism of the producer rather than catching producer exceptions and retrying from the application itself; the idempotent producer makes this pattern EVEN MORE APPEALING — it is the easiest way to avoid duplicates when retrying."
(Same conclusion as Ch. 7 §4.4, now with a second reason: your retry loop is invisible to dedup; the producer's is not.)
Limitation 2: Multiple producer instances
"It is also rather common to have applications that have multiple instances or even one instance with multiple producers. If two of these producers attempt to send identical messages, the idempotent producer will not detect the duplication."
The concrete example:
► Idempotence is PER-PRODUCER-INSTANCE. It is not a global dedup service.
2.7 How to use it
enable.idempotence=true"This is the easy part... If the producer is already configured with
acks=all, there will be NO DIFFERENCE IN PERFORMANCE."
Four things change:
| Change | Detail |
|---|---|
| One extra API call at startup | to retrieve a producer ID |
| 96 bits added per record batch | producer ID (a long) + sequence ID for the first message in the batch ("sequence IDs for each message in the batch are derived from the sequence ID of the first message plus a delta") — "barely any overhead for most workloads" |
| Brokers validate sequence numbers | "from any single producer instance and guarantee the lack of duplicate messages" |
| Ordering guaranteed through all failure scenarios | "even if max.in.flight.requests.per.connection is set to more than 1 (5 is the default and also the highest value supported by the idempotent producer)" |
That last row is the one people undervalue. Without idempotence, retries>0 + max.in.flight>1 silently reorders (Ch. 3 §6). Idempotence gives you ordering AND dedup AND 5 in-flight requests simultaneously — there is no reason not to enable it.
NOTE — KIP-360 improvements in 2.5
"Idempotent producer logic and error handling improved significantly in version 2.5 (both producer and broker side) as a result of KIP-360."
Before 2.5:
- "the producer state was not always maintained for long enough, which resulted in fatal
UNKNOWN_PRODUCER_IDerrors in various scenarios."- A known edge case with partition reassignment: "the new replica became the leader before any writes happened from a specific producer, meaning that the new leader had no state for that partition."
- "previous versions attempted to rewrite the sequence IDs in some error scenarios, which could lead to duplicates."
In newer versions: "if we encounter a fatal error for a record batch, this batch and all the batches that are in flight will be rejected. The user who writes the application can handle the exception and decide whether to skip those records or retry and risk duplicates and reordering."