8.6 Performance of transactions
Note that the producer's natural unit is "one transaction per poll() batch," which conveniently ties transaction size to max.poll.records — giving you one knob that moves both.
Producer side — overhead is per-transaction, not per-message
| Cost | How often |
|---|---|
| Register transactional ID | Once in the producer lifecycle |
| Register partitions in a transaction | At most once per partition per transaction |
| Commit request | One per transaction |
| Commit marker written | On each partition |
⚠ “Transactional INITIALIZATION and transaction COMMIT requests are SYNCHRONOUS, so NO DATA WILL BE SENT until they complete successfully, fail, or time out, which FURTHER INCREASES THE OVERHEAD.”
💡 The key optimization
"the overhead of transactions on the producer is INDEPENDENT OF THE NUMBER OF MESSAGES IN A TRANSACTION. So a larger number of messages per transaction will BOTH reduce the relative overhead AND reduce the number of synchronous stops, resulting in HIGHER THROUGHPUT overall."
Consumer side — no throughput cost, but a latency cost
- Some overhead reading COMMIT MARKERS.
- The key impact: “consumers in
read_committedmode will not return records that are part of an OPEN transaction. LONG INTERVALS BETWEEN TRANSACTION COMMITS mean the consumer will need to WAIT LONGER before returning messages, and as a result, END-TO-END LATENCY WILL INCREASE.” - Good news: “the consumer does NOT need to BUFFER messages that belong to open transactions. THE BROKER WILL NOT RETURN THOSE in response to fetch requests. Since there is NO EXTRA WORK for the consumer when reading transactions, THERE IS NO DECREASE IN THROUGHPUT EITHER.”
⚠️ The transaction-size tension
| Bigger transactions | Smaller transactions |
|---|---|
| LOWER producer overhead, HIGHER throughput. | LOWER consumer latency. |
But LONGER read_committed consumer latency — the LSO is held back longer. | But MORE synchronous commits, MORE markers, HIGHER relative overhead. |
► Transaction size is a THROUGHPUT-vs-LATENCY dial, exactly like linger.ms is for producers (Ch. 3) and fetch.min.bytes is for consumers (Ch. 4). Same shape of trade-off, different layer.
Note that the producer's natural unit is "one transaction per poll() batch," which conveniently ties transaction size to max.poll.records — giving you one knob that moves both.