3.5 Message delivery time — the timeout configs and how they interact
How long the producer may block when calling send() and when explicitly requesting metadata via partitionsFor().
"The producer has multiple configuration parameters that interact to control one of the behaviors that are of most interest to developers: how long will it take until a call to
send()will succeed or fail."The configs "were modified several times over the years" — this describes the latest implementation, introduced in Apache Kafka 2.1.
The two intervals
| Interval | What it covers | Governed by |
|---|---|---|
| 1 | Until the async send() call returns — the thread that called send() is blocked during this. | max.block.ms |
| 2 | From send() returning successfully until the callback fires (with success or failure) — that is, from “record placed in a batch” until Kafka responds with success, a non-retriable failure, or we run out of allocated time. | delivery.timeout.ms (the umbrella), which contains linger.ms + request.timeout.ms + retries |
Note: "If you use
send()synchronously, the sending thread will block for both time intervals continuously, and you won't be able to tell how much time was spent in each."
Detailed timeline
- linger.ms — wait for more msgs to batch
- request.timeout.ms — wait for the server’s reply to THIS request
- retry (retry.backoff.ms) ↺ repeat
The umbrella runs from “record placed in a batch” until Kafka responds with success, a non-retriable failure, or we run out of allocated time — including retries.
The configs, one by one
max.block.ms
How long the producer may block when calling send() and when explicitly requesting metadata via partitionsFor(). Those methods block when:
- the producer's send buffer is full, or
- metadata is not available.
When reached, a timeout exception is thrown.
delivery.timeout.ms — the one you should actually tune
Limits the time from the point a record is ready for sending (send() returned successfully and the record is in a batch) until either the broker responds or the client gives up — including time spent on retries.
- Must be greater than
linger.ms+request.timeout.ms. "If you try to create a producer with an inconsistent timeout configuration, you will get an exception." - "Messages can be successfully sent much faster than
delivery.timeout.ms, and typically will."
Which exception you get depends on where the clock ran out:
| Timeout exceeded... | Callback receives |
|---|---|
| while retrying | "the exception that corresponds to the error that the broker returned before retrying" |
| while the record batch was still waiting to be sent | a timeout exception |
💡 TIP — the right mental model for retries
*"Configure the delivery timeout to the maximum time you'll want to wait for a message to be sent, typically a few minutes, and then leave the default number of retries (virtually infinite). With this configuration, the producer will keep retrying for as long as it has time to keep trying (or until it succeeds). This is a much more reasonable way to think about retries.
Our normal process for tuning retries is: 'In case of a broker crash, it typically takes leader election 30 seconds to complete, so let's keep retrying for 120 seconds just to be on the safe side.' Instead of converting this mental dialog to number of retries and time between retries, *you just configure
delivery.timeout.msto 120[s]."
This reframing is the single most useful operational idea in the chapter: you reason in wall-clock recovery time, not in retry counts. Retry count × backoff is a derived quantity you should never hand-compute.
request.timeout.ms
How long the producer waits for a reply from the server when sending data.
"Note that this is the time spent waiting on each producer request before giving up; it does not include retries, time spent before sending, and so on."
If reached without a reply, the producer either retries sending or completes the callback with a TimeoutException.
retries and retry.backoff.ms
retries = how many times to retry before giving up. Default wait between retries: 100 ms (retry.backoff.ms).
"We recommend against using these parameters in the current version of Kafka. Instead, test how long it takes to recover from a crashed broker (i.e., how long until all partitions get new leaders), and set
delivery.timeout.mssuch that the total time spent retrying will be longer than the time it takes the cluster to recover from the crash — otherwise, the producer will give up too soon."
"Because the producer handles retries for you, there is no point in handling retries within your own application logic. You will want to focus your efforts on handling non-retriable errors or cases where retry attempts were exhausted."
TIP: "If you want to completely disable retries, setting
retries=0is the only way to do so."
linger.ms
Time to wait for additional messages before sending the current batch. KafkaProducer sends a batch either when the batch is full, or when linger.ms is reached.
Default behavior (linger.ms = 0): "the producer will send messages as soon as there is a sender thread available to send them, even if there's just one message in the batch."
Setting it > 0:
"increases latency a little and significantly increases throughput — the overhead per message is much lower, and compression, if enabled, is much better."
linger.ms = 0 | linger.ms = 10 |
|---|---|
[m1] → request — 40 requests | Wait up to 10 ms, collect: [m1 m2 m3 m4 m5 … m40] → 1 request |
| High per-message overhead. Compression sees 1 msg → poor ratio. | Far less per-message overhead. Compression sees 40 msgs of similar data → much better ratio. Cost: +≤10 ms latency. |
The compression point is the underrated one: compression ratio is a function of how much similar data is in the same block. linger.ms=0 largely defeats compression.
buffer.memory
Memory the producer uses to buffer messages waiting to be sent.
If the application sends faster than messages can be delivered, "the producer may run out of space, and additional
send()calls will block formax.block.msand wait for space to free up before throwing an exception.""Note that unlike most producer exceptions, this timeout is thrown by
send()and NOT by the resultingFuture."
That distinction matters: a try/catch around send() catches this one; a callback does not.
compression.type
Default: messages are sent uncompressed. Options: snappy, gzip, lz4, zstd.
| Codec | Character | Recommended when |
|---|---|---|
| snappy | "invented by Google to provide decent compression ratios with low CPU overhead and good performance" | both performance and bandwidth are a concern |
| gzip | "typically use more CPU and time but results in better compression ratios" | network bandwidth is more restricted |
| lz4, zstd | (also available) |
"By enabling compression, you reduce network utilization and storage, which is often a bottleneck when sending messages to Kafka."
(Cross-reference Ch. 2: the broker decompresses to validate checksums and assign offsets, then recompresses. Compression is not free on the broker either.)
batch.size
Memory in bytes (not messages!) used for each batch. When the batch is full, all messages in it are sent.
The crucial clarification most people get wrong:
"However, this does not mean that the producer will wait for the batch to become full. The producer will send half-full batches and even batches with just a single message in them. Therefore, setting the batch size too large will not cause delays in sending messages; it will just use more memory for the batches. Setting the batch size too small will add some overhead because the producer will need to send messages more frequently."
So: batch.size is a ceiling and a memory cost, not a trigger to wait for. linger.ms is the only thing that makes the producer wait.
max.in.flight.requests.per.connection
How many message batches the producer will send without receiving responses.
"Higher settings can increase memory usage while improving throughput. Apache's wiki experiments show that in a single-DC environment, the throughput is maximized with only 2 in-flight requests; however, the default value is 5 and shows similar performance."