3.14 Deploy / monitor / scale — the producer's operational surface
There is no producer-side backup, but there is a producer-side durability boundary, and knowing where it sits is what matters:
Config recipes by requirement
Maximum durability — credit-card transactions: no loss, no duplicates
| Setting | Value |
|---|---|
acks | all |
enable.idempotence | true ← no dupes + ordering |
max.in.flight.requests.per.connection | ≤ 5 |
retries | > 0 (leave the near-infinite default) |
delivery.timeout.ms | > measured cluster recovery time |
broker-side: min.insync.replicas | 2, RF ≥ 3 (Ch. 2/7) |
| send pattern | send(record, callback) — never fire-and-forget |
Recall: acks=all costs nothing in end-to-end latency.
Maximum throughput — click tracking: some loss/dupes tolerable
| Setting | Value |
|---|---|
acks | 1 (or 0 if truly disposable) |
linger.ms | 10–100 ← bigger batches, better compression |
batch.size | raised ← memory cost only, no latency cost |
compression.type | snappy (balanced) | gzip (bandwidth-constrained) |
buffer.memory | raised ← absorb bursts before blocking |
max.in.flight | 2–5 (2 maximizes single-DC throughput) |
Cross-datacenter producer
| Setting | Value |
|---|---|
send.buffer.bytes / receive.buffer.bytes | increased (higher latency, lower bandwidth links) |
linger.ms | raised — batching pays more when RTT is high |
Monitoring the producer
| Signal | Why |
|---|---|
produce-throttle-time-avg / -max | You are being quota-throttled (or request-time throttled) |
| Buffer available bytes / buffer exhausted rate | Precursor to send() blocking and TimeoutException |
| Record error rate / retry rate | Retriable errors churning; possible leader instability |
| Batch size avg, records per request | Whether linger.ms/batch.size are actually batching |
| Compression rate | Whether compression is achieving anything (see #13) |
| Request latency avg/max | vs request.timeout.ms headroom |
| Callback latency (your own metric) | Guards against the blocking-callback trap (#3) |
client.id on every metric | The whole point of naming clients well |
Scaling the producer
| Goal | Lever |
|---|---|
| More producer throughput | Larger linger.ms + batch.size + compression; then more producer instances — a producer is thread-safe and shareable across threads, so first scale threads, then processes. |
| More partitions | More parallelism at the broker — but see the key→partition warning before adding partitions to a keyed topic. |
| Avoid hot partitions | Custom partitioner, or UniformStickyPartitioner. |
| Protect the cluster | Quotas — dynamic, per client-id or user. |
"Backup" from the producer's perspective
There is no producer-side backup, but there is a producer-side durability boundary, and knowing where it sits is what matters:
Practical consequence: anything in the accumulator is unreplicated, unacked, and gone if the JVM dies. If you cannot tolerate that, you need either a durable local outbox before the producer, or send().get() at the boundary (accepting the throughput cost) — and the "errors file for later analysis" pattern the book mentions for the callback failure path.