3.4 Configuring producers — every important knob
A logical identifier for the client and the application it's used in.
4.1 client.id
A logical identifier for the client and the application it's used in. Any string. Used by brokers to identify messages from the client — in logging, metrics, and quotas.
The book's argument for taking this seriously is the most practical sentence in the chapter:
Choosing a good client name is "the difference between 'We are seeing a high rate of authentication failures from IP 104.27.155.134' and 'Looks like the Order Validation service is failing to authenticate — can you ask Laura to take a look?'"
4.2 acks — the durability knob
Controls how many partition replicas must receive the record before the producer considers the write successful.
Default: leader-only received (acks=1) — "release 3.0 of Apache Kafka is expected to change this default" (it changed to acks=all).
- sends as fast as the network will support → very high throughput
- if the broker never received it, THE PRODUCER WILL NOT KNOW. Message lost, silently.
- if the write to leader fails (e.g. leader crashed, no new leader yet), producer gets an ERROR and can retry → avoids potential data loss
- “The message can still get lost if the leader crashes and the latest messages were not yet replicated to the new leader.”
- SAFEST — more than one broker has the message; survives a crash
- higher latency than acks=1 (waiting for more than one broker)
4.3 💡 The single best insight in the chapter: producer latency ≠ end-to-end latency
*"You will see that with lower and less reliable
acksconfiguration, the producer will be able to send records faster. This means that you trade off reliability for producer latency.However, end-to-end latency is measured from the time a record was produced until it is available for consumers to read — and is IDENTICAL for all three options.
The reason is that, in order to maintain consistency, Kafka will not allow consumers to read records until they are written to all in-sync replicas. Therefore, if you care about end-to-end latency, rather than just the producer latency, *there is no trade-off to make: you will get the same end-to-end latency if you choose the most reliable option."
The only thing acks < all buys you is EARLIER FALSE CONFIDENCE on the producer side. The data isn’t readable any sooner — Kafka “will not allow consumers to read records until they are written to all in-sync replicas.”
Practical implication: if your SLO is "event visible to consumers within X ms" — which is almost always the real SLO — then acks=all is free. Choosing acks=1 for latency is usually optimizing a metric nobody cares about while accepting real data-loss risk.