Learn Labs
13. Monitoring Kafka

13.1 Metric basics

1.1 Where the metrics live

All Kafka metrics are exposed via JMX. Three collection approaches:

  1. A Separate process on the host connecting to the JMX interface — e.g. Nagios XI check_jmx plug-in, jmxtrans.
  2. A JMX agent running Inside the Kafka process, exposing metrics over HTTP — e.g. Jolokia, MX4J. ◄ Recommended (see the security note).
  3. Monitoring-as-a-service — “many companies offer monitoring agents, metrics collection points, storage, graphing, and alerting in a services package”.

⚠️ FINDING THE JMX PORT — and why you probably shouldn't open it

*"the broker sets the configured JMX port in the broker information stored in ZooKeeper. The /brokers/ids/<ID> znode contains JSON-formatted data including hostname and jmx_port keys.

However, REMOTE JMX IS DISABLED BY DEFAULT in Kafka FOR SECURITY REASONS. If you are going to enable it, you must properly configure security for the port. THIS IS BECAUSE JMX NOT ONLY ALLOWS A VIEW INTO THE STATE OF THE APPLICATION, IT ALSO ALLOWS CODE EXECUTION.

It is HIGHLY RECOMMENDED that you use a JMX metrics agent that is LOADED INTO THE APPLICATION."*

JMX is a remote-code-execution surface. That single sentence should settle the architecture decision: use an in-process agent (option ②), not an open remote JMX port.

1.2 💡 The five metric sources — ordered by objectivity

"the LOWER in the list, the more OBJECTIVE a view of Kafka they provide."

CategoryDescription
Application metrics"from Kafka itself, from the JMX interface"
Logs"Also from Kafka itself. Because it is some form of text or structured data, and not just a number, it requires a little more processing"
Infrastructure metrics"from systems that you have in front of Kafka but are still within the request path and under your control. An example is a load balancer"
Synthetic clients"external to your Kafka deployment, just like a client, but under your direct control and typically not performing the same work as your clients. An external monitor like Kafka Monitor"
Client metrics"exposed by the Kafka clients that connect to your cluster"

The website analogy that makes the point:

"The web server is running properly, and all of the metrics IT is reporting say that it is working. HOWEVER, there is a problem with the NETWORK between your web server and your external users, which means that NONE OF YOUR USERS CAN REACH THE WEB SERVER. A synthetic client running outside your network would detect this."

"relying on metrics from your brokers will SUFFICE AT THE START, but later on you will want a more objective view."

app metricslogsinfrastructuresynthetic clientsclient metricsSUBJECTIVE — “the broker thinks it’s fine”OBJECTIVE — “the client actually succeeded”good for DEBUGGINGgood for ALERTING and SLIs
Figure 13.1.11.2 💡 The five metric sources — ordered by objectivity

1.3 Alerting vs debugging vs historical — three different data lifecycles

PurposeRetentionCharacterConsumer
Alerting"a very short period... hours, or maybe days""important for these metrics to be more OBJECTIVE, as a problem that does not impact clients is far less critical than one that does""automation that responds to known problems, as well as the human operators"
Debugging"days or weeks past when it is collected""more SUBJECTIVE measurements, or data from the Kafka application itself"Humans, diagnosing
Historical"measured in YEARS""resources used, including compute, storage, and network" + "additional METADATA to put the metrics into context, such as when brokers were added to or removed from the cluster"Capacity management

💡 You don't have to collect debugging data continuously: "it is NOT ALWAYS NECESSARY to collect this data into a monitoring system. If the metrics are used for debugging problems in place, it is sufficient that the metrics are available when needed. YOU DO NOT NEED TO OVERWHELM THE MONITORING SYSTEM BY COLLECTING TENS OF THOUSANDS OF VALUES ON AN ONGOING BASIS."

1.4 💡 Automation vs humans — the "Check Engine light" principle

Consumed by automation:
“they should be Very specific. It's OK to have a Large number of metrics, each describing Small details, because this is why computers exist… The more specific the data is, the Easier it is to create automation.”
Consumed by humans:

“presenting a large number of metrics will be Overwhelming.”

⚠ “It is far too easy to succumb to ‘Alert fatigue,’ where there are so many alerts going off that It is difficult to know how severe the problem is. It is also Hard to properly define thresholds for every metric And keep them up-to-date. When the alerts are overwhelming or often incorrect, We begin to not trust that the alerts are correctly describing the state of our applications.”

The car analogy:

"To properly adjust the ratio of air to fuel, the computer needs a number of measurements of air density, fuel, exhaust... These measurements would be OVERWHELMING to the human operator. Instead, we have a 'CHECK ENGINE' LIGHT. A SINGLE INDICATOR tells you that there is a problem, and there is a way to find out more detailed information to tell you exactly what the problem is."

1.5 Application health checks

  1. An external process that reports whether the broker is up or down — for the broker, “simply connecting to the external port (the same port that clients use) to check that it responds.”
  2. Alerting on the lack of metrics being reported — “stale metrics.”

⚠ “Though the second method works, it can make it difficult to differentiate between a failure of the Kafka broker and a failure of the monitoring system itself.”


On this page