Learn Labs
8. Exactly-Once Semantics

8.5 3.7 Transactional IDs and fencing

"Choosing the transactional ID for producers is important and a bit more challenging than it seems. Assigning the transactional ID incorrectly can lead to either application errors or LOSS OF EXACTLY-ONCE GUARANTEES."

The two key requirements:

  1. Consistent for the same instance of the application Between restarts
  2. Different for different instances of the application

► “otherwise the brokers will not be able to fence off zombie instances”

The pre-2.5 world: static partition mapping

"Until release 2.5, the only way to guarantee fencing was to STATICALLY MAP THE TRANSACTIONAL ID TO PARTITIONS. This guaranteed that each partition will always be consumed with the same transactional ID."

Why anything else breaks:

✗ loses connectivityNOT FENCEDproducer transactional.id = Aprocesses messages from topic TNEW producer transactional.id = Breplaces itproducer A comes backAS A ZOMBIEoutput topicboth A and B write — duplicates
  • · “Zombie A will NOT BE FENCED because THE ID DOESN’T MATCH that of the new producer B.”
  • · What we need: “producer A to ALWAYS BE REPLACED BY PRODUCER A” — the new A gets a HIGHER EPOCH, and zombie A is properly fenced.
Figure 8.5.1Why anything else breaks

The book flags its own example: "In those releases, the previous example would be INCORRECT — transactional IDs are assigned randomly to threads without making sure the same transactional ID is always used to write to the same partition."

KIP-447 (Kafka 2.5): fencing by consumer group metadata

"In Apache Kafka 2.5, KIP-447 introduced a second method of fencing based on CONSUMER GROUP METADATA, in addition to transactional IDs. We use the producer offset commit method and pass as an argument the CONSUMER GROUP METADATA rather than just the consumer group ID."

The scenario — and why the old way was wasteful:

Before the zombie (Figure 8-3)
T1 t-0consumer Aproducer Atxn.id = AT2 partition 0T1 t-1consumer Bproducer Btxn.id = BT2 partition 1Both consumers are in the same group.
After instance A becomes a zombie (Figure 8-4)
T1 t-0T1 t-1consumer Bboth partitionsproducer Btxn.id = BT2 partition 0T2 partition 1instance AZOMBIEno longer in the group — its writes must be rejected
The pre-2.5 problem

“If we want to guarantee that no zombies write to partition 0, consumer B CAN’T just start reading from partition 0 and writing to partition 0 with transactional ID B. Instead the application will need to INSTANTIATE A NEW PRODUCER, WITH TRANSACTIONAL ID A, to safely write to partition 0 and fence the old transactional ID A. THIS IS WASTEFUL.”

The KIP-447 solution
  • · Include the CONSUMER GROUP GENERATION in the transaction.
  • · producer B’s transactions show a NEWER GENERATION → they GO THROUGH
  • · zombie producer A’s transactions show an OLD GENERATION → FENCED
  • · ► No need to spin up a producer per partition-identity. B can just work.
Figure 8.5.2The scenario — and why the old way was wasteful

This is what makes the code example in §4 legal. With subscribe(), "partitions assigned to this instance can change at any point as a result of rebalance" — impossible to reconcile with static transactional-ID↔partition mapping. Group-generation fencing removes that constraint.


On this page