Learn Labs
3. Kafka Producers: Writing Messages to Kafka

3.14 Deploy / monitor / scale — the producer's operational surface

There is no producer-side backup, but there is a producer-side durability boundary, and knowing where it sits is what matters:

Config recipes by requirement

Maximum durability — credit-card transactions: no loss, no duplicates

SettingValue
acksall
enable.idempotencetrue ← no dupes + ordering
max.in.flight.requests.per.connection≤ 5
retries> 0 (leave the near-infinite default)
delivery.timeout.ms> measured cluster recovery time
broker-side: min.insync.replicas2, RF ≥ 3 (Ch. 2/7)
send patternsend(record, callback) — never fire-and-forget

Recall: acks=all costs nothing in end-to-end latency.

Maximum throughput — click tracking: some loss/dupes tolerable

SettingValue
acks1 (or 0 if truly disposable)
linger.ms10–100 ← bigger batches, better compression
batch.sizeraised ← memory cost only, no latency cost
compression.typesnappy (balanced) | gzip (bandwidth-constrained)
buffer.memoryraised ← absorb bursts before blocking
max.in.flight2–5 (2 maximizes single-DC throughput)

Cross-datacenter producer

SettingValue
send.buffer.bytes / receive.buffer.bytesincreased (higher latency, lower bandwidth links)
linger.msraised — batching pays more when RTT is high

Monitoring the producer

SignalWhy
produce-throttle-time-avg / -maxYou are being quota-throttled (or request-time throttled)
Buffer available bytes / buffer exhausted ratePrecursor to send() blocking and TimeoutException
Record error rate / retry rateRetriable errors churning; possible leader instability
Batch size avg, records per requestWhether linger.ms/batch.size are actually batching
Compression rateWhether compression is achieving anything (see #13)
Request latency avg/maxvs request.timeout.ms headroom
Callback latency (your own metric)Guards against the blocking-callback trap (#3)
client.id on every metricThe whole point of naming clients well

Scaling the producer

GoalLever
More producer throughputLarger linger.ms + batch.size + compression; then more producer instances — a producer is thread-safe and shareable across threads, so first scale threads, then processes.
More partitionsMore parallelism at the broker — but see the key→partition warning before adding partitions to a keyed topic.
Avoid hot partitionsCustom partitioner, or UniformStickyPartitioner.
Protect the clusterQuotas — dynamic, per client-id or user.

"Backup" from the producer's perspective

There is no producer-side backup, but there is a producer-side durability boundary, and knowing where it sits is what matters:

the networkin your app’s heaprecord accumulator (buffer.memory) · batches waiting on linger.msin Kafkaleader disk (acks=1) · all ISR disks (acks=all)← LOST on process crash. acks can’t save this.← durability starts HERE
Figure 3.14.3'Backup' from the producer's perspective

Practical consequence: anything in the accumulator is unreplicated, unacked, and gone if the JVM dies. If you cannot tolerate that, you need either a durable local outbox before the producer, or send().get() at the boundary (accepting the throughput cost) — and the "errors file for later analysis" pattern the book mentions for the callback failure path.


On this page