"transactions were added to Kafka to provide multipartition atomic WRITES (but not READS) and to fence zombie producers in stream processing applications. As a result, they provide exactly-once guarantees when used within chains of consume-process-produce stream processing tasks. In other contexts, transactions will either straight-out NOT WORK or will require ADDITIONAL EFFORT."
"The two main mistakes are assuming that exactly-once guarantees apply on ACTIONS OTHER THAN PRODUCING TO KAFKA, and that consumers ALWAYS READ ENTIRE TRANSACTIONS and have information about TRANSACTION BOUNDARIES."
"Let's say that the record processing step includes sending email to users. Enabling exactly-once semantics will NOT guarantee that the email will only be sent once. The guarantee only applies to records written to Kafka. Using sequence numbers to deduplicate records or using markers to abort or cancel a transaction works within Kafka, but IT WILL NOT UN-SEND AN EMAIL. The same is true for any action with external effects: calling a REST API, writing to a file, etc."
"the application is writing to an external database rather than to Kafka. In this scenario, THERE IS NO PRODUCER INVOLVED — records are written to the database using a database driver (likely JDBC) and offsets are committed to Kafka within the consumer. There is NO MECHANISM that allows writing results to an external database and committing offsets to Kafka within a single transaction."
The workaround — flip which system owns the transaction:
"Instead, we could manage offsets IN THE DATABASE (as explained in Chapter 4) and commit both data and offsets to the database in a single transaction — this would rely on the DATABASE's transactional guarantees rather than Kafka's."
"Microservices often need to update the database AND publish a message to Kafka within a single atomic transaction, so either both will happen or neither will. Kafka transactions will NOT do this."
Direction 1 — Kafka is the outbox:
microservice ──► publishes ONLY to a Kafka topic (the "outbox") │ ▼ a separate MESSAGE RELAY SERVICE │ ▼ updates the database ⚠ "Because Kafka won't guarantee an exactly-once update to the database, IT IS IMPORTANT TO MAKE SURE THE UPDATE IS IDEMPOTENT." ► Guarantees: "the message will EVENTUALLY make it to Kafka, the topic consumers, AND the database — OR TO NONE OF THOSE."
Direction 2 — a database table is the outbox:
microservice ──► writes data AND an outbox row in ONE DB transaction │ ▼ a relay service reads the outbox table │ ▼ produces to Kafka "This pattern is PREFERRED when built-in RDBMS constraints, such as UNIQUENESS and FOREIGN KEYS, are useful."
Reference: "The Debezium project published an in-depth blog post on the outbox pattern with detailed examples."
"It is very tempting to believe that we can build an app that will read data from a database, identify database transactions, write the records to Kafka, and from there write records to another database, still maintaining the original transactions from the source database. Unfortunately, Kafka transactions don't have the necessary functionality to support these kinds of end-to-end guarantees."
Two independent reasons:
the problem from ② above — can’t commit records and offsets to a DB and Kafka in one transaction
“read_committed guarantees in Kafka consumers are Too weak to preserve database transactions. Yes, a consumer will not see records that were not committed. But it is Not guaranteed to have seen All the records that were committed within the transaction because It could be lagging on some topics; it has no information to identify transaction boundaries, so it can’t know when a transaction began and ended, and whether it has seen Some, none, or all of its records.”
"This one is more subtle — it IS possible to support exactly-once guarantees when copying data from one Kafka cluster to another. There is a description of how this is done in the KIP for adding exactly-once capabilities in MirrorMaker 2.0. At the time of this writing, the proposal is still in draft, but the algorithm is clearly described. This proposal includes the guarantee that each record in the source cluster will be copied to the destination cluster exactly once."
"However, this does NOT guarantee that TRANSACTIONS WILL BE ATOMIC. If an app produces several records and offsets transactionally, and then MirrorMaker 2.0 copies them to another cluster, the transactional properties and guarantees WILL BE LOST during the copy process."*
Same root cause:"the consumer reading data from Kafka can't know or guarantee that it is getting ALL the events in a transaction. For example, it can replicate PART of a transaction if it is only subscribed to a subset of the topics."
exactly-once RECORD delivery ✓ possible (MM2 KIP) preserving TRANSACTION atomicity across clusters ✗ not possible
"We've discussed exactly-once in the context of the consume-process-produce pattern, but the publish/subscribe pattern is a very common use case. Using transactions in a pub/sub use case provides some guarantees: consumers configured with read_committed will not see records that were published as part of an aborted transaction. But those guarantees FALL SHORT of exactly-once. CONSUMERS MAY PROCESS A MESSAGE MORE THAN ONCE, DEPENDING ON THEIR OWN OFFSET COMMIT LOGIC."
The JMS comparison — and the important difference:
"The guarantees Kafka provides in this case are similar to those provided by JMS transactions but DEPEND ON CONSUMERS in read_committed mode to guarantee that uncommitted transactions will remain invisible. JMS brokers withhold uncommitted transactions from ALL consumers."
JMS
the Broker enforces invisibility for everyone.
Kafka
the Consumer must opt in (isolation.level), and the default opts Out. Visibility is a client-side decision.
"An important pattern to AVOID is publishing a message and then WAITING FOR ANOTHER APPLICATION TO RESPOND before committing the transaction. The other application WILL NOT RECEIVE THE MESSAGE UNTIL AFTER THE TRANSACTION WAS COMMITTED, resulting in a DEADLOCK."
Figure 8.4.2·⑤ Publish/subscribe pattern
Summary table:
Scenario
Exactly-once?
Consume → process → produce to Kafka
✅ Yes — this is what transactions are for
Any external side effect (email, REST, file)
❌ No — cannot be rolled back
Kafka → external database
❌ No producer involved → use the DB's transaction (offsets in the DB)
DB write + Kafka publish atomically
❌ No → outbox pattern (either direction)
DB → Kafka → DB preserving source transactions
❌ No — plus consumers have no transaction boundaries