13.3 Kafka broker metrics
That's an important nuance: disk space isn't a gradual-degradation signal — it's a cliff.
⚠️ WHO WATCHES THE WATCHERS?
*"Many organizations use Kafka for collecting application metrics, system metrics, and logs for consumption by a central monitoring system. This is an excellent way to decouple the applications from the monitoring system, BUT IT PRESENTS A SPECIFIC CONCERN FOR KAFKA ITSELF. If you use this same system for monitoring Kafka itself, IT IS VERY LIKELY THAT YOU WILL NEVER KNOW WHEN KAFKA IS BROKEN BECAUSE THE DATA FLOW FOR YOUR MONITORING SYSTEM WILL BE BROKEN AS WELL.
Two fixes:
- "use a SEPARATE monitoring system for Kafka that does not have a dependency on Kafka"
- "if you have multiple datacenters, make sure that the metrics for the Kafka cluster in datacenter A are produced to datacenter B, and VICE VERSA"
"However you decide to handle it, MAKE SURE THAT THE MONITORING AND ALERTING FOR KAFKA DOES NOT DEPEND ON KAFKA WORKING."
3.1 The three categories of cluster problem
| Category | How it presents | What to monitor / do |
|---|---|---|
| ① Single-broker problems — “by far the easiest to diagnose” | “Show up as outliers in the metrics for the cluster,” and are “frequently related to slow or failing storage devices or compute restraints from other applications on the system.” | Individual server availability and storage device status, via OS metrics. “Absent a problem identified at the OS or hardware level, however, the cause is almost always an imbalance in the load.” “While Kafka attempts to keep the data evenly spread, this does not mean that client access to that data is evenly distributed. It also does not detect issues such as hot partitions.” So it is “highly recommended that you utilize an external tool for keeping the cluster balanced at all times” — e.g. Cruise Control. |
| ② Overloaded clusters — “also easy to detect” | “If the cluster is balanced, and many of the brokers are showing elevated latency for requests or a low request handler pool idle ratio, you are reaching the limits of your brokers.” | “You may find that a client has changed its request pattern… Even when this happens, there may be little you can do about changing the client. The solutions are either reduce the load or increase the number of brokers.” |
| ③ Controller problems — “much more difficult to diagnose” | “Often fall into the category of bugs in Kafka itself.” Manifests as broker metadata being out of sync, offline replicas when the brokers appear to be fine, and topic control actions like creation not happening properly. | Active controller count and controller queue size. “If you're scratching your head over a problem and saying ‘that's really weird,’ there is a very good chance that it is because the controller did something unpredictable and bad.” |
💡 PREFERRED REPLICA ELECTIONS — always do this first
"THE FIRST STEP before trying to diagnose a problem further is to ensure that you have RUN A PREFERRED REPLICA ELECTION recently. Kafka brokers do not automatically take partition leadership back (unless auto leader rebalance is enabled) after they have released leadership... THE PREFERRED REPLICA ELECTION IS SAFE AND EASY TO RUN, SO IT'S A GOOD IDEA TO DO THAT FIRST AND SEE IF THE PROBLEM GOES AWAY."
3.2 ⚠️ Under-replicated partitions — and the retracted advice
- Metric name
- Under-replicated partitions
- JMX MBean
kafka.server:type=ReplicaManager,name=UnderReplicatedPartitions- Value range
- Integer, zero or greater
"gives a count of the number of partitions for which the broker is the LEADER replica, where the FOLLOWER replicas are NOT CAUGHT UP."
⚠️ THE URP ALERTING TRAP — the book retracting its own advice
*"In the PREVIOUS EDITION of this book, as well as in many conference talks, the authors have spoken at length about the fact that the URP metric SHOULD BE YOUR PRIMARY ALERTING METRIC because of how many problems it describes.
THIS APPROACH HAS A SIGNIFICANT NUMBER OF PROBLEMS, not the least of which is that the URP metric CAN FREQUENTLY BE NONZERO FOR BENIGN REASONS. This means that you will receive FALSE ALERTS, WHICH LEAD TO THE ALERT BEING IGNORED. It also requires A SIGNIFICANT AMOUNT OF KNOWLEDGE to be able to understand what the metric is telling you.
FOR THIS REASON, WE NO LONGER RECOMMEND THE USE OF URP FOR ALERTING. Instead, you should depend on SLO-BASED ALERTING to detect unknown problems."*
URP remains excellent for diagnosis — the interpretation guide:
- A steady URP count on many brokers “normally indicates that one of the brokers in the cluster is offline.” “The count across the entire cluster will equal the number of partitions that are assigned to that broker, and the broker that is down will not report a metric.”
- “If the URPs are on a single broker, then that broker is typically the problem. The error shows that other brokers are having a problem replicating messages from that one.”
- “If several brokers have URPs, it could be a cluster problem, but it might still be a single broker — because a single broker is having problems replicating messages from everywhere.”
- The technique: list the URPs and look for a common broker.
The common-broker technique, worked:
kafka-topics.sh --bootstrap-server kafka1.example.com:9092/kafka-cluster \
--describe --under-replicated
Topic: topicOne Partition: 5 Leader: 1 Replicas: 1,2 Isr: 1
Topic: topicOne Partition: 6 Leader: 3 Replicas: 2,3 Isr: 3
Topic: topicTwo Partition: 3 Leader: 4 Replicas: 2,4 Isr: 4
Topic: topicTwo Partition: 7 Leader: 5 Replicas: 5,2 Isr: 5
Topic: topicSix Partition: 1 Leader: 3 Replicas: 2,3 Isr: 3
... ▲
BROKER 2 APPEARS IN EVERY ROW
but is never in the Isr"In this example, the common broker is number 2. This indicates that this broker is having a problem with message replication and will lead us to focus our investigation on that one broker. If there is NO common broker, there is likely a CLUSTER-WIDE problem."
3.3 Cluster-level problems
Problem A: unbalanced load — "the easiest to FIND even though FIXING it can be an involved process"
Five metrics to compare across brokers:
- Partition count
- Leader partition count
- All topics messages in rate
- All topics bytes in rate
- All topics bytes out rate
What "balanced" looks like (Table 13-4):
| Broker | Partitions | Leaders | Messages in | Bytes in |
|---|---|---|---|---|
| 1 | 100 | 50 | 13130 msg/s | 3.56 MBps |
| 2 | 101 | 49 | 12842 msg/s | 3.66 MBps |
| 3 | 100 | 50 | 13086 msg/s | 3.23 MBps |
“all the brokers are taking approximately the same amount of traffic”
"Assuming you have ALREADY RUN A PREFERRED REPLICA ELECTION, a large deviation indicates that the traffic is not balanced. To resolve this, you will need to move partitions from the heavily loaded brokers to the less heavily loaded brokers using
kafka-reassign-partitions.sh(Ch. 12)."
HELPERS FOR BALANCING CLUSTERS
"The Kafka broker itself does NOT provide for automatic reassignment of partitions. This means that balancing traffic can be a MIND-NUMBING PROCESS of manually reviewing long lists of metrics and trying to come up with a replica assignment that works." Tools:
kafka-assignerin LinkedIn's open source kafka-tools repo; "Some enterprise offerings for Kafka support also provide this feature."
Problem B: resource exhaustion
The bottlenecks: "CPU, disk IO, and network throughput are a few of the most common."
⚠️ The disk-utilization exception: "DISK UTILIZATION is NOT one of them, as the brokers will operate properly RIGHT UP UNTIL THE DISK IS FILLED, AND THEN THIS DISK WILL FAIL ABRUPTLY."
That's an important nuance: disk space isn't a gradual-degradation signal — it's a cliff. Alert on free space as a capacity metric, not a performance one.
OS-level metrics to track:
- CPU utilization
- Inbound network throughput
- Outbound network throughput
- Disk average wait time
- Disk percent utilization
💡 The key insight: "Exhausting ANY of these resources will typically show up as THE SAME PROBLEM: under-replicated partitions. It's critical to remember that THE BROKER REPLICATION PROCESS OPERATES IN EXACTLY THE SAME WAY THAT OTHER KAFKA CLIENTS DO. IF YOUR CLUSTER IS HAVING PROBLEMS WITH REPLICATION, THEN YOUR CUSTOMERS ARE HAVING PROBLEMS WITH PRODUCING AND CONSUMING MESSAGES AS WELL."
(Ch. 6 §4.3: followers use the same Fetch requests consumers use. Replication health is a proxy for client health.)
"It makes sense to develop a BASELINE for these metrics when your cluster is operating correctly and then SET THRESHOLDS THAT INDICATE A DEVELOPING PROBLEM LONG BEFORE YOU RUN OUT OF CAPACITY. ... All Topics Bytes In Rate is a good guideline to show cluster usage."
3.4 Host-level problems
Four categories: "Hardware failures · Networking · Conflicts with another process · Local configuration differences"
Hardware — the soft failures are the dangerous ones
"Hardware failures are sometimes obvious, like when the server just stops working, but it's the LESS OBVIOUS problems that cause performance issues. These are usually SOFT FAILURES that ALLOW THE SYSTEM TO KEEP RUNNING BUT DEGRADE OPERATION. This could be a bad bit of memory, where the system has detected the problem and BYPASSED THAT SEGMENT (reducing the overall available memory). The same can happen with a CPU failure."
Tools: "the facilities that your hardware provides, such as an intelligent platform management interface (IPMI)"; "looking at the kernel ring buffer using dmesg will help you to see log messages that are getting thrown to the system console."
💡 Disk failure — and why one bad disk ruins everything
"The MORE COMMON type of hardware failure that leads to a performance degradation in Kafka is A DISK FAILURE. Kafka is dependent on the disk for persistence, and PRODUCER PERFORMANCE IS DIRECTLY TIED TO HOW FAST YOUR DISKS COMMIT THOSE WRITES. Any deviation will show up as problems with the performance of the producers AND the replica fetchers. The latter is what leads to under-replicated partitions."
⚠️ ONE BAD EGG
"A SINGLE DISK FAILURE ON A SINGLE BROKER CAN DESTROY THE PERFORMANCE OF AN ENTIRE CLUSTER. This is because producer clients will connect to ALL brokers that lead partitions for a topic, and if you have followed best practices, those partitions will be EVENLY SPREAD over the entire cluster. If ONE broker starts performing poorly and slowing down produce requests, THIS WILL CAUSE BACK PRESSURE IN THE PRODUCERS, SLOWING DOWN REQUESTS TO ALL BROKERS."
- Good partition balance means every producer talks to every broker, which means one sick broker degrades every producer.
- This is why single-broker outliers are urgent, not cosmetic.
Disk monitoring checklist:
- Hardware status from IPMI or your hardware's interface.
- SMART (Self-Monitoring, Analysis and Reporting Technology) tools — “to both monitor and test the disks on a regular basis. This will alert you to a failure that is about to happen.”
- ⚠ The disk controller, “especially if it has RAID functionality, whether you are using hardware RAID or not. Many controllers have an onboard cache that is only used when the controller is healthy and the battery backup unit (BBU) is working. A failure of the BBU can result in the cache being disabled, degrading disk performance.”
The BBU failure mode is a classic: nothing is "broken," no alert fires, and write latency quietly triples because the controller silently switched from write-back to write-through.
Networking
- Hardware
- “a bad network cable or connector”
- Configuration
- “a change in the Speed or Duplex settings for the connection, either on the server side Or upstream on the networking hardware”
- OS
- “having the Network buffers undersized or Too many network connections taking up too much of the overall memory footprint”
💡 Key indicator: “the Number of errors detected on the network interfaces. if the error count is increasing, there is probably an unaddressed issue.”
Process conflicts and config drift
"another common problem to look for is another application running on the system that is consuming resources. This could be something that was installed in error, or a process that is SUPPOSED to be running, such as A MONITORING AGENT, but is having problems. Use tools such as
top."
"If the other options have been exhausted... a CONFIGURATION DIFFERENCE has likely crept in. ... This is why it is CRUCIAL that you utilize a CONFIGURATION MANAGEMENT SYSTEM, such as Chef or Puppet, in order to maintain consistent configurations across your OSes and applications (including Kafka)."