13. Monitoring Kafka
13.11 The recommended monitoring stack
Tier 1 — primary alerts (page a human)
- SLO burn rate (from client-side or synthetic-client SLIs) — the only alert that catches problems you've never seen before
- Xinfra Monitor: produce/consume availability per broker
- Burrow: consumer group status (not raw lag)
OfflinePartitionsCount > 0(“site down”)ActiveControllerCount(cluster sum) ≠ 1- Producer
record-error-rate > 0(client side; should always be zero) - Broker/host up-down health check (external port connect)
Tier 2 — secondary alerts / tickets
RequestHandlerAvgIdlePercent< 20% (warn) / < 10% (urgent)- Produce
TotalTimeMsp99.9 above baseline - Producer
request-latency-avgabove baseline - Consumer
commit-latency-avgabove baseline ControllerEventManager EventQueueSizerising and not drainingLeaderCount == 0on a live broker, or ≫ the 1/RF share- Disk free space / inodes — capacity, because it fails as a cliff
- Open FDs approaching max
- Network interface error count increasing
- Swap in use > 0
- Consumer
sync-rate≫ 0 (rebalance storm) - Throttle-time metrics > 0 (quota throttling — otherwise invisible)
Tier 3 — dashboard (no alerts; for diagnosis)
UnderReplicatedPartitionsper broker — find the common brokerPartitionCount,LeaderCount,LeaderCount ÷ PartitionCountper broker- BytesIn / BytesOut / MessagesIn per broker (balance and growth trend)
- All seven request timings per request type (phase localization)
- GC
CollectionTime/Count+LastGcInfo.duration - OS: load average (÷ CPUs), CPU breakdown (especially
wa,st), disk wait/util/queue, network in/out - Producer:
outgoing-byte-rate,record-send-rate,request-rate,batch-size-avg,record-queue-time-avg, per-broker latency - Consumer:
bytes-consumed-rate/records-consumed-rate,fetch-rate,fetch-size-avg,assigned-partitionsper instance
Logging baseline
- Everything at
INFO kafka.controller→ separate file,INFOkafka.server.ClientQuotaManager→ separate file,INFOkafka.log.LogCleaner/Cleaner/LogCleanerManager→DEBUG— cheap, and the only visibility into silently-halted compactionkafka.request.logger→ off unless actively debugging- Ch. 11:
kafka.authorizer.loggerfor security auditing
The diagnostic decision tree
- Host-level, in detail: disk (SMART, IPMI, controller and BBU,
wa% in CPU) · network interface errors, speed/duplex · another process eating resources (top) · config drift. - ⚠ Remember “one bad egg” — this will look cluster-wide from the producers' perspective.
- On the balanced-but-slow branch, also check for old clients forcing recompression.