Learn Labs
13. Monitoring Kafka

13.9 End-to-end monitoring

"Consumer and producer clients have metrics that CAN INDICATE that there might be a problem with the Kafka cluster, BUT THIS CAN BE A GUESSING GAME as to whether increased latency is due to A PROBLEM WITH THE CLIENT, THE NETWORK, OR KAFKA ITSELF. In addition, it means that if you are responsible for running the Kafka cluster, AND NOT THE CLIENTS, YOU WOULD NOW HAVE TO MONITOR ALL OF THE CLIENTS AS WELL."

The two questions you actually need answered:

 ① Can I PRODUCE messages to the Kafka cluster?
 ② Can I CONSUME messages from the Kafka cluster?

Xinfra Monitor (formerly Kafka Monitor):

"open sourced by the Kafka team at LinkedIn, [it] continually produces and consumes data from a topic that is SPREAD ACROSS ALL BROKERS in a cluster. It measures the AVAILABILITY of both produce and consume requests ON EACH BROKER, as well as the TOTAL PRODUCE-TO-CONSUME LATENCY."

"In an ideal world, you would be able to monitor this for every topic individually. However, in most situations IT IS NOT REASONABLE TO INJECT SYNTHETIC TRAFFIC INTO EVERY TOPIC. We can, however, at least provide those answers for every BROKER in the cluster."

"This type of monitoring is INVALUABLE to be able to EXTERNALLY VERIFY that the Kafka cluster is operating as intended, since — JUST LIKE CONSUMER LAG MONITORING — THE KAFKA BROKER CANNOT REPORT WHETHER OR NOT CLIENTS ARE ABLE TO USE THE CLUSTER PROPERLY."

is the cluster usable?end-to-end availabilityare consumers keeping up?lagare clients happy?SLIsthe broker’s own metricscannot answer any of thesesomething OUTSIDE Kafkaall three require it

The recurring theme of this chapter: broker metrics are for diagnosis, not for knowing whether you have a problem.

Figure 13.9.1End-to-end monitoring