Learn Labs
6. Kafka Internals

6.9 Deploy / monitor / scale / backup — through the internals lens

Ch. 6 finally makes precise what Kafka's durability actually is:

The metrics this chapter explains

AreaMetricWhat it tells you
ControllerActiveControllerCountmust be exactly 1 cluster-wide
controller epoch changesfrequent changes = ZooKeeper instability or GC pauses (see failure #1)
ReplicationUnderReplicatedPartitionsISR shrank → election ineligibility
IsrShrinksPerSec / IsrExpandsPerSecflapping = lag near the threshold
replica.lag.time.max.msalso bounds consumer visibility delay (§5.6) — not just durability
Request pipeline — the §5.1 diagramRequestQueueSizenetwork threads → I/O threads
ResponseQueueSizeI/O threads → network threads
NetworkProcessorAvgIdlePcttoo low → add network threads
RequestHandlerAvgIdlePcttoo low → add I/O threads
PurgatorySizeacks=all waits + delayed fetches
Correlation IDties a client timeout to a broker log line
Format / compatFetchMessageConversionsPerSecold clients burning broker CPU (KIP-188)
MessageConversionsTimeMshow much CPU time that conversion costs
Storageper-mount disk usageplacement is count-based, not size
open file descriptorsone handle per segment, always
log cleaner activity / errorsoffset-map memory failures

Tuning knobs and the mechanism each controls

KnobThe mechanism it actually touches
num.network.threadsProcessor threads in §5.1 — moving requests to/from queues
num.io.threadsRequest handler threads — actual request processing
replica.lag.time.max.msISR membership and the consumer visibility delay ceiling
min.insync.replicasThe third produce-request validation (§5.3)
acks=allWhether the response waits in purgatory
linger.msBatch fill → per-record overhead + compression ratio (§6.5)
log.segment.bytes / log.roll.msSegment count → retention precision, file handles, timestamp-seek precision
log.cleaner.enabled + offset-map memory + thread countWhether compaction can run at all
min/max.compaction.lag.msCompaction timing guarantees (compliance)
broker.rackThe rack-alternating broker list in partition allocation (§6.3)
client.rack + replica.selector.classFollower fetching, at the cost of HW-propagation delay
metadata.max.age.msClient metadata cache freshness → NotLeaderForPartition frequency

"Backup" — the internals answer

Ch. 6 finally makes precise what Kafka's durability actually is:

  1. Kafka does not fsync. “It relies on replication for message durability.” Your durability unit is “N brokers in N failure domains hold this in page cache,” not “it's on disk.”
  2. Therefore: RF + rack/AZ diversity is the durability design. Correlated power loss across all ISR members can lose acknowledged data. That is by design, not a bug.
  3. Indexes are derived — freely deletable, auto-regenerated. They never need backing up.
  4. Compacted topics are the closest thing Kafka has to a permanent store: “turn Kafka into a long-term data store.” That's what makes state recovery-by-replay viable.
  5. Tiered storage is the real long-term answer — and explicitly aims to eliminate “separate data pipelines to copy the data from Kafka to external stores, as done currently in many deployments.”

On this page