13. Monitoring Kafka
13. Monitoring Kafka
Chapter 13 of Kafka: The Definitive Guide — 12 sections.
Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 13 The problem this chapter solves: "The Apache Kafka applications have numerous measurements — so many, in fact, that it can easily become confusing as to what is important to watch and what can be set aside. ... They provide a detailed view into every operation in the broker, but they can also make you the bane of whoever is responsible for managing your monitoring system."
Note: this edition reverses the previous edition's central advice. The old guidance ("alert on under-replicated partitions") is explicitly retracted in favor of SLO-based alerting. See §3.2.
Sections
- 13.12Metric basics
- 13.24Service-level objectives
- 13.36Kafka broker metricsThat's an important nuance: disk space isn't a gradual-degradation signal — it's a cliff.
- 13.48The broker metrics reference
- 13.51JVM and OS monitoringVia java.lang:type=OperatingSystem:
- 13.61Logging
- 13.72Client monitoringThis is the metric that tells you whether linger.ms is actually costing you what you think (Ch. 3 §5).
- 13.82Lag monitoring
- 13.91End-to-end monitoring
- 13.10Failure catalogWhat actually breaks in production — Ch. 13 consolidated
- 13.112The recommended monitoring stack
- 13.12Self-testSelf-test