10.8 Tuning MirrorMaker
8.1 Sizing the cluster
“MirrorMaker is HORIZONTALLY SCALABLE. Sizing depends on THE THROUGHPUT YOU NEED and THE LAG YOU CAN TOLERATE.”
| Tolerance | Size it like this |
|---|---|
| CAN’T tolerate lag | “size MirrorMaker with enough capacity to keep up with YOUR TOP THROUGHPUT” |
| CAN tolerate lag | “size MirrorMaker to be 75–80% UTILIZED 95–99% OF THE TIME. Then EXPECT SOME LAG TO DEVELOP when you are at peak throughput. Because MirrorMaker has SPARE CAPACITY MOST OF THE TIME, IT WILL CATCH UP ONCE THE PEAK IS OVER.” |
8.2 Finding the right tasks.max — an actual procedure
- Use
kafka-performance-producerto “generate load on a source cluster” - “connect MirrorMaker and start mirroring this load”
- Test with: 1, 2, 4, 8, 16, 24, 32 tasks
- “Watch where performance tapers off and SET
tasks.maxJUST BELOW THIS POINT.”
⚠ CPU watch: “If you are consuming or producing Compressed events (Recommended, since Bandwidth is the main bottleneck for cross-datacenter mirroring), Mirrormaker will have to decompress and recompress the events. this uses a lot of CPU, so keep an eye on CPU utilization as you increase the number of tasks.”
Then: if one worker isn’t enough, add Workers.
⚠ “If you are running MirrorMaker on an existing Connect cluster with other connectors, Make sure you also take the load from those connectors into account when sizing.”
(The decompress/recompress cost is the same broker-side cost from Ch. 6 §5.5 — and it's why Cluster Linking (§9) is faster: it avoids the round trip entirely.)
8.3 Isolate your sensitive topics
"you may want to separate sensitive topics — those that absolutely require LOW LATENCY and where the mirror must be as close to the source as possible — to a SEPARATE MIRRORMAKER CLUSTER. This will PREVENT A BLOATED TOPIC OR AN OUT-OF-CONTROL PRODUCER FROM SLOWING DOWN YOUR MOST SENSITIVE DATA PIPELINE."
8.4 TCP stack tuning (cross-datacenter)
- CLIENT-side buffers:
send.buffer.bytes/receive.buffer.bytes - BROKER-side buffers:
socket.send.buffer.bytes/socket.receive.buffer.bytes
“These should be COMBINED WITH OPTIMIZATION OF THE NETWORK CONFIGURATION IN LINUX:”
# increase TCP buffer sizes
net.core.rmem_default, net.core.rmem_max,
net.core.wmem_default, net.core.wmem_max, net.core.optmem_max
# enable AUTOMATIC WINDOW SCALING
sysctl -w net.ipv4.tcp_window_scaling=1
# (or add net.ipv4.tcp_window_scaling=1 to /etc/sysctl.conf)
# REDUCE THE TCP SLOW START TIME
set /proc/sys/net/ipv4/tcp_slow_start_after_idle to 0
"tuning the Linux network is a large and complex topic" — recommended reading: Performance Tuning for Linux Servers by Sandra K. Johnson et al. (IBM Press).
(tcp_slow_start_after_idle=0 is the WAN-specific one: it stops TCP from resetting its congestion window after idle periods, which matters enormously on high-latency links.)
8.5 💡 Diagnosing the bottleneck: producer or consumer?
- METHOD 1 — metrics: “If ONE PROCESS IS IDLE WHILE THE OTHER IS FULLY UTILIZED, you know which one needs tuning.” (
io-ratio,io-wait-ratio) - METHOD 2 — “do several THREAD DUMPS (using
jstack) and see if the MirrorMaker threads are spending most of the time in POLL or in SEND:- more time POLLING → THE CONSUMER is the bottleneck
- more time SENDING → THE PRODUCER is the bottleneck”
8.6 Tuning the producer
| Config | When to change it — based on a metric |
|---|---|
linger.ms / batch.size | "If monitoring shows the producer consistently sends partially empty batches (i.e., batch-size-avg and batch-size-max are lower than configured batch.size), you can increase throughput by introducing a bit of latency — increase linger.ms." Conversely: "If you are sending full batches and have memory to spare, increase batch.size." |
max.in.flight.requests.per.connection | See below — a real correctness/throughput trade |
⚠️ The MirrorMaker ordering caveat
*"Limiting the number of in-flight requests to 1 is currently THE ONLY WAY FOR MIRRORMAKER TO GUARANTEE THAT MESSAGE ORDERING IS PRESERVED if some messages require multiple retries before they are successfully acknowledged. But this means every request sent by the producer has to be acknowledged by the target cluster before the next message is sent. THIS CAN LIMIT THROUGHPUT, ESPECIALLY IF THERE IS SIGNIFICANT LATENCY before the brokers acknowledge.
If message order is NOT critical for your use case, using the default value of 5 can SIGNIFICANTLY INCREASE YOUR THROUGHPUT."*
Note the contrast with Ch. 3 §6:
- For Your own producers,
enable.idempotence=truegives you ordering With 5 in-flight requests — no trade needed. - For Mirrormaker, the book says in-flight=1 is “currently the only way.”
► Over a WAN, where RTT is high, in-flight=1 is Brutally expensive. You are choosing between ordering and throughput, explicitly.
8.7 Tuning the consumer
| Config | The metric that tells you to change it |
|---|---|
fetch.max.bytes | "If fetch-size-avg and fetch-size-max are close to fetch.max.bytes, the consumer is reading as much as it is allowed. If you have memory, increase it." |
fetch.min.bytes + fetch.max.wait.ms | "If fetch-rate is HIGH, the consumer is sending too many requests and not receiving enough data in each. Increase both so the consumer receives more per request and the broker waits until enough data is available." |
(Every tuning recommendation in this section is metric-driven. That's the model to copy: never tune a config without the metric that justifies it.)
10.7 Deploying MirrorMaker in production
A canary is the only listed monitor that measures the thing you actually care about — end-to-end delivery — rather than a proxy for it.
10.9 Other cross-cluster mirroring solutions
That last mechanism is elegant: it automatically trades throughput for availability exactly when needed, and trades back when the emergency ends — precisely the manual decision th…