Learn Labs
2. Installing Kafka

2.7 OS tuning

This is a great example of a config whose semantics changed underneath a widely-copied best practice.

Most Linux distributions ship kernel-tuning defaults that work fairly well for most applications, but a few changes improve performance for a Kafka broker. These revolve around virtual memory, networking, and the disk mount point used for log segments. Typically configured in /etc/sysctl.conf (check your distro docs).

7.1 Virtual memory

Swap — avoid at (almost) all costs

Two reasons:

  1. The cost of swapped-out pages "will show up as a noticeable impact on all aspects of performance in Kafka."
  2. Kafka makes heavy use of the page cache — "if the VM system is swapping to disk, there is not enough memory being allocated to page cache."

Should you disable swap entirely? You can — swap isn't a requirement — but:

It does provide a safety net if something catastrophic happens. Having swap can prevent the OS from abruptly killing a process due to an out-of-memory condition.

Therefore:

vm.swappiness = 1

The parameter is a percentage of how likely the VM subsystem is to use swap space rather than dropping pages from the page cache. It is preferable to reduce the memory available for page cache rather than utilize any amount of swap memory.

Why not swappiness = 0? — the changed-semantics trap

The old recommendation was 0, which used to mean "do not swap unless there is an out-of-memory condition." The meaning changed as of Linux kernel 3.5-rc1, backported widely (RHEL kernels as of 2.6.32-303). 0 now means "never swap under any circumstances" — which removes the OOM safety net. Hence 1 is now the recommendation.

This is a great example of a config whose semantics changed underneath a widely-copied best practice. If you inherited a runbook that says vm.swappiness=0, it's wrong for the reason you think it's right.

Dirty pages

Kafka relies on disk I/O performance to give producers good response times — which is also why log segments go on a fast disk (SSD, or a subsystem with significant NVRAM for caching, e.g. RAID).

vm.dirty_background_ratio = 5     # default 10 — % of total system memory

Lower it below the default of 10. 5 is appropriate in many situations. Do NOT set it to zero — that causes the kernel to continually flush pages, which eliminates the kernel's ability to buffer disk writes against temporary spikes in the underlying device performance.

vm.dirty_ratio = 60..80           # default 20 — % of total system memory

This is the total dirty pages allowed before the kernel forces synchronous operations to flush them. Raise above the default of 20; between 60 and 80 is reasonable.

This introduces risk in two ways: the amount of unflushed disk activity (data in memory, not on disk), and the potential for long I/O pauses if synchronous flushes are forced. If a higher vm.dirty_ratio is chosen, it is highly recommended that replication be used in the cluster to guard against system failures.

The tradeoff, stated plainly:

  • High dirty_ratio → better write batching, higher throughput
  • → more data in volatile memory at any instant
  • → bigger stall when a forced sync flush happens

⇒ You are trading durability-on-one-machine for speed, and paying for it with replication. The book says it directly: “if a higher vm.dirty_ratio is chosen, it is highly recommended that replication be used in the cluster to guard against system failures.”

Observe the actual behavior instead of guessing:

cat /proc/vmstat | egrep "dirty|writeback"
# nr_dirty 21845
# nr_writeback 0
# nr_writeback_temp 0
# nr_dirty_threshold 32715981
# nr_dirty_background_threshold 2726331

Review the number of dirty pages over time while the cluster is under load, in production or simulated.

File descriptors / memory maps

Kafka uses file descriptors for log segments and open connections.

 minimum needed ≈ (number_of_partitions) × (partition_size / segment_size)
                  + number of connections the broker makes
vm.max_map_count = 400000   # or 600000 — based on the calculation above
vm.overcommit_memory = 0

Set vm.max_map_count to a very large number based on the above calculation; 400,000 or 600,000 has generally been successful. vm.overcommit_memory = 0 (the default) means the kernel determines the amount of free memory from an application. A non-zero value "could lead the operating system to grab too much memory, depriving memory for Kafka to operate optimally. This is common for applications with high ingestion rates."

7.2 Disk / filesystem

Outside of hardware and RAID choice, the filesystem has the next largest impact on performance.

XFS (recommended)Ext4
StatusDefault FS for many Linux distros—
PerformanceOutperforms Ext4 for most workloads with minimal tuningCan perform well
Tuning neededBeyond the FS's own automatic tuning, noneRequires parameters considered less safe
Specific riskAlso uses delayed allocation, but generally safer than Ext4'sLonger commit interval to force less frequent flushes; delayed allocation of blocks → greater chance of data loss and filesystem corruption on system failure
BatchingMore efficient when batching disk writes → better overall I/O throughput—

Mount options, regardless of filesystem:

 mount -o noatime,largeio ...

noatime — why it's safe here specifically. File metadata has three timestamps: ctime (creation), mtime (last modified), atime (last access). By default atime is updated every time a file is read → a large number of disk writes (writes generated by reads — the worst kind).

The atime attribute is generally of little use unless an app needs to know if a file was accessed since last modified (use relatime then). atime is not used by Kafka at all, so disabling it is safe. noatime prevents these updates without affecting proper handling of ctime and mtime.

(Important: Kafka does rely on mtime — that's how time-based retention works. noatime leaves mtime intact, which is why it's safe.)

largeio — helps improve efficiency when there are larger disk writes.

7.3 Networking

The kernel is not tuned by default for large, high-speed data transfers. The recommended changes for Kafka are the same as those suggested for most web servers and other networking applications.

Socket buffers (all sockets):

net.core.wmem_default = 131072      # 128 KiB
net.core.rmem_default = 131072      # 128 KiB
net.core.wmem_max     = 2097152     # 2 MiB
net.core.rmem_max     = 2097152     # 2 MiB

"Keep in mind that the maximum size does not indicate that every socket will have this much buffer space allocated; it only allows up to that much if needed."

TCP-specific buffers — set separately, as three space-separated integers min default max:

net.ipv4.tcp_wmem = 4096 65536 2048000
net.ipv4.tcp_rmem = 4096 65536 2048000
#                   4KiB  64KiB  2MiB

The maximum cannot be larger than net.core.wmem_max / net.core.rmem_max. Based on actual workload, you may want to increase the maximums to allow greater buffering.

Other useful parameters:

ParameterSet toWhy
net.ipv4.tcp_window_scaling1Clients transfer data more efficiently, and data can be buffered on the broker side
net.ipv4.tcp_max_syn_backlog> 1024 (default)Allows a greater number of simultaneous connections to be accepted
net.core.netdev_max_backlog> 1000 (default)Assists with bursts of network traffic, specifically at multigigabit speeds, by allowing more packets to be queued for the kernel to process

On this page