2.7 OS tuning
This is a great example of a config whose semantics changed underneath a widely-copied best practice.
Most Linux distributions ship kernel-tuning defaults that work fairly well for most applications, but a few changes improve performance for a Kafka broker. These revolve around virtual memory, networking, and the disk mount point used for log segments. Typically configured in
/etc/sysctl.conf(check your distro docs).
7.1 Virtual memory
Swap — avoid at (almost) all costs
Two reasons:
- The cost of swapped-out pages "will show up as a noticeable impact on all aspects of performance in Kafka."
- Kafka makes heavy use of the page cache — "if the VM system is swapping to disk, there is not enough memory being allocated to page cache."
Should you disable swap entirely? You can — swap isn't a requirement — but:
It does provide a safety net if something catastrophic happens. Having swap can prevent the OS from abruptly killing a process due to an out-of-memory condition.
Therefore:
vm.swappiness = 1The parameter is a percentage of how likely the VM subsystem is to use swap space rather than dropping pages from the page cache. It is preferable to reduce the memory available for page cache rather than utilize any amount of swap memory.
Why not swappiness = 0? — the changed-semantics trap
The old recommendation was
0, which used to mean "do not swap unless there is an out-of-memory condition." The meaning changed as of Linux kernel 3.5-rc1, backported widely (RHEL kernels as of 2.6.32-303).0now means "never swap under any circumstances" — which removes the OOM safety net. Hence1is now the recommendation.
This is a great example of a config whose semantics changed underneath a widely-copied best practice. If you inherited a runbook that says vm.swappiness=0, it's wrong for the reason you think it's right.
Dirty pages
Kafka relies on disk I/O performance to give producers good response times — which is also why log segments go on a fast disk (SSD, or a subsystem with significant NVRAM for caching, e.g. RAID).
vm.dirty_background_ratio = 5 # default 10 — % of total system memoryLower it below the default of 10. 5 is appropriate in many situations. Do NOT set it to zero — that causes the kernel to continually flush pages, which eliminates the kernel's ability to buffer disk writes against temporary spikes in the underlying device performance.
vm.dirty_ratio = 60..80 # default 20 — % of total system memoryThis is the total dirty pages allowed before the kernel forces synchronous operations to flush them. Raise above the default of 20; between 60 and 80 is reasonable.
This introduces risk in two ways: the amount of unflushed disk activity (data in memory, not on disk), and the potential for long I/O pauses if synchronous flushes are forced. If a higher
vm.dirty_ratiois chosen, it is highly recommended that replication be used in the cluster to guard against system failures.
The tradeoff, stated plainly:
- High
dirty_ratio→ better write batching, higher throughput - → more data in volatile memory at any instant
- → bigger stall when a forced sync flush happens
⇒ You are trading durability-on-one-machine for speed, and paying for it with replication. The book says it directly: “if a higher vm.dirty_ratio is chosen, it is highly recommended that replication be used in the cluster to guard against system failures.”
Observe the actual behavior instead of guessing:
cat /proc/vmstat | egrep "dirty|writeback"
# nr_dirty 21845
# nr_writeback 0
# nr_writeback_temp 0
# nr_dirty_threshold 32715981
# nr_dirty_background_threshold 2726331Review the number of dirty pages over time while the cluster is under load, in production or simulated.
File descriptors / memory maps
Kafka uses file descriptors for log segments and open connections.
minimum needed ≈ (number_of_partitions) × (partition_size / segment_size)
+ number of connections the broker makesvm.max_map_count = 400000 # or 600000 — based on the calculation above
vm.overcommit_memory = 0Set
vm.max_map_countto a very large number based on the above calculation; 400,000 or 600,000 has generally been successful.vm.overcommit_memory = 0(the default) means the kernel determines the amount of free memory from an application. A non-zero value "could lead the operating system to grab too much memory, depriving memory for Kafka to operate optimally. This is common for applications with high ingestion rates."
7.2 Disk / filesystem
Outside of hardware and RAID choice, the filesystem has the next largest impact on performance.
| XFS (recommended) | Ext4 | |
|---|---|---|
| Status | Default FS for many Linux distros | — |
| Performance | Outperforms Ext4 for most workloads with minimal tuning | Can perform well |
| Tuning needed | Beyond the FS's own automatic tuning, none | Requires parameters considered less safe |
| Specific risk | Also uses delayed allocation, but generally safer than Ext4's | Longer commit interval to force less frequent flushes; delayed allocation of blocks → greater chance of data loss and filesystem corruption on system failure |
| Batching | More efficient when batching disk writes → better overall I/O throughput | — |
Mount options, regardless of filesystem:
mount -o noatime,largeio ...noatime — why it's safe here specifically. File metadata has three timestamps: ctime (creation), mtime (last modified), atime (last access). By default atime is updated every time a file is read → a large number of disk writes (writes generated by reads — the worst kind).
The
atimeattribute is generally of little use unless an app needs to know if a file was accessed since last modified (userelatimethen).atimeis not used by Kafka at all, so disabling it is safe.noatimeprevents these updates without affecting proper handling ofctimeandmtime.
(Important: Kafka does rely on mtime — that's how time-based retention works. noatime leaves mtime intact, which is why it's safe.)
largeio — helps improve efficiency when there are larger disk writes.
7.3 Networking
The kernel is not tuned by default for large, high-speed data transfers. The recommended changes for Kafka are the same as those suggested for most web servers and other networking applications.
Socket buffers (all sockets):
net.core.wmem_default = 131072 # 128 KiB
net.core.rmem_default = 131072 # 128 KiB
net.core.wmem_max = 2097152 # 2 MiB
net.core.rmem_max = 2097152 # 2 MiB"Keep in mind that the maximum size does not indicate that every socket will have this much buffer space allocated; it only allows up to that much if needed."
TCP-specific buffers — set separately, as three space-separated integers min default max:
net.ipv4.tcp_wmem = 4096 65536 2048000
net.ipv4.tcp_rmem = 4096 65536 2048000
# 4KiB 64KiB 2MiBThe maximum cannot be larger than
net.core.wmem_max/net.core.rmem_max. Based on actual workload, you may want to increase the maximums to allow greater buffering.
Other useful parameters:
| Parameter | Set to | Why |
|---|---|---|
net.ipv4.tcp_window_scaling | 1 | Clients transfer data more efficiently, and data can be buffered on the broker side |
net.ipv4.tcp_max_syn_backlog | > 1024 (default) | Allows a greater number of simultaneous connections to be accepted |
net.core.netdev_max_backlog | > 1000 (default) | Assists with bursts of network traffic, specifically at multigigabit speeds, by allowing more packets to be queued for the kernel to process |