Learn Labs
13. Monitoring Kafka

13.11 The recommended monitoring stack

Tier 1 — primary alerts (page a human)

  • SLO burn rate (from client-side or synthetic-client SLIs) — the only alert that catches problems you've never seen before
  • Xinfra Monitor: produce/consume availability per broker
  • Burrow: consumer group status (not raw lag)
  • OfflinePartitionsCount > 0 (“site down”)
  • ActiveControllerCount (cluster sum) ≠ 1
  • Producer record-error-rate > 0 (client side; should always be zero)
  • Broker/host up-down health check (external port connect)

Tier 2 — secondary alerts / tickets

  • RequestHandlerAvgIdlePercent < 20% (warn) / < 10% (urgent)
  • Produce TotalTimeMs p99.9 above baseline
  • Producer request-latency-avg above baseline
  • Consumer commit-latency-avg above baseline
  • ControllerEventManager EventQueueSize rising and not draining
  • LeaderCount == 0 on a live broker, or ≫ the 1/RF share
  • Disk free space / inodes — capacity, because it fails as a cliff
  • Open FDs approaching max
  • Network interface error count increasing
  • Swap in use > 0
  • Consumer sync-rate ≫ 0 (rebalance storm)
  • Throttle-time metrics > 0 (quota throttling — otherwise invisible)

Tier 3 — dashboard (no alerts; for diagnosis)

  • UnderReplicatedPartitions per broker — find the common broker
  • PartitionCount, LeaderCount, LeaderCount ÷ PartitionCount per broker
  • BytesIn / BytesOut / MessagesIn per broker (balance and growth trend)
  • All seven request timings per request type (phase localization)
  • GC CollectionTime/Count + LastGcInfo.duration
  • OS: load average (÷ CPUs), CPU breakdown (especially wa, st), disk wait/util/queue, network in/out
  • Producer: outgoing-byte-rate, record-send-rate, request-rate, batch-size-avg, record-queue-time-avg, per-broker latency
  • Consumer: bytes-consumed-rate/records-consumed-rate, fetch-rate, fetch-size-avg, assigned-partitions per instance

Logging baseline

  • Everything at INFO
  • kafka.controller → separate file, INFO
  • kafka.server.ClientQuotaManager → separate file, INFO
  • kafka.log.LogCleaner / Cleaner / LogCleanerManager → DEBUG — cheap, and the only visibility into silently-halted compaction
  • kafka.request.logger → off unless actively debugging
  • Ch. 11: kafka.authorizer.logger for security auditing

The diagnostic decision tree

Alert firesburn rate / availability / lag status① preferred replica electionrun it first — did it fix it? → done② ONE broker or ALL?outliers in BytesIn/Out, LeaderCount, latencyONE BROKERhost-level investigationhost-level checksdisk · network errors · other processes · config driftALL BROKERSis the cluster BALANCED?NOT balancedkafka-reassign-partitions / Cruise Controlbalanced, latency high, idle ratio lowOVERLOADED: reduce load or add brokers③ URP nonzero?list the URPs, look for the COMMON brokercommon broker foundthat broker is the problemno common brokercluster-wide problem④ “that’s REALLY WEIRD”suspect the CONTROLLERActiveControllerCount (sum) 0 or 2?EventQueueSize stuck high?⑤ which PHASE is slow?read the seven request timingsqueue → local (disk) → remote (replication)→ throttle → response
  • Host-level, in detail: disk (SMART, IPMI, controller and BBU, wa% in CPU) · network interface errors, speed/duplex · another process eating resources (top) · config drift.
  • ⚠ Remember “one bad egg” — this will look cluster-wide from the producers' perspective.
  • On the balanced-but-slow branch, also check for old clients forcing recompression.
Figure 13.11.2The diagnostic decision tree

On this page