Learn Labs
13. Monitoring Kafka

13.7 Client monitoring

This is the metric that tells you whether linger.ms is actually costing you what you think (Ch. 3 §5).

7.1 Producer metrics

"The Kafka producer client has greatly compacted the metrics available by making them available as attributes on a small number of JMX MBeans. In contrast, the previous (unsupported) version used a larger number of MBeans but had MORE DETAIL (a greater number of percentile measurements and different moving averages). As a result, the overall number of metrics covers a WIDER SURFACE AREA, but IT CAN BE MORE DIFFICULT TO TRACK OUTLIERS."

Three beans (Table 13-19):

ScopeJMX MBean
Overallkafka.producer:type=producer-metrics,client-id=CLIENTID
Per-brokerkafka.producer:type=producer-node-metrics,client-id=CLIENTID,node-id=node-BROKERID
Per-topickafka.producer:type=producer-topic-metrics,client-id=CLIENTID,topic=TOPICNAME

"while we will discuss several metrics that are averages (ending in -avg), there are also maximum values for each metric (ending in -max) that have LIMITED USEFULNESS."

The two producer metrics to ALERT on
MetricWhy it deserves an alert
record-error-rate“One attribute that you will definitely want to set an alert for. This metric should always be zero, and if it is anything greater, the producer is dropping messages it is trying to send.” “The producer has a configured number of retries and a backoff, and once that has been exhausted, the messages will be dropped.” “There is also a record-retry-rate attribute that can be tracked, but it is less critical because retries are normal.”
request-latency-avg“The average amount of time a produce request sent to the brokers takes. You should be able to establish a baseline value for what this number should be in normal operations, and set an alert threshold above that.” “An increase means produce requests are getting slower. This could be networking issues, or problems on the brokers. Either way, it's a performance issue that will cause back pressure and other problems in your producing application.”

(Ch. 7 §6.3 named exactly these two: "the two metrics most important for reliability are error-rate and retry-rate per record." Note record-error-rate also counts benign idempotent-duplicate rejections — Ch. 8 §2.3.)

The three views of traffic volume

"A single REQUEST contains one or more BATCHES. A single BATCH contains one or more MESSAGES. And, of course, each MESSAGE is made up of some number of BYTES. These metrics are all useful to have on an application dashboard."

outgoing-byte-rate
“the messages in Absolute size in bytes per second”
record-send-rate
“the traffic in terms of the Number of messages produced per second”
request-rate
“the number of Produce requests sent to the brokers per second”
The size metrics
request-size-avg
“average size of the produce Requests in bytes”
batch-size-avg
“average size of a single message Batch (which, by definition, is comprised of messages for A single topic partition) in bytes”
record-size-avg

“average size of a single Record in bytes”

“For a Single-topic producer, this provides useful information. For Multiple-topic producers, such as Mirrormaker, It is less informative.”

records-per-request-avg
“average number of messages in a single produce request”
💡 record-queue-time-avg — the linger.ms tuning metric

"the average amount of time, in milliseconds, that a single message WAITS IN THE PRODUCER, AFTER THE APPLICATION SENDS IT, BEFORE IT IS ACTUALLY PRODUCED TO KAFKA."

After send(), the producer waits until One of two things happens:

  1. “It has enough messages to Fill a batch based on batch.size”
  2. “It has been long enough since the last batch was sent based on linger.ms”

“The easiest way to understand it is that For busy topics, the first condition will apply, whereas For slow topics, the second will apply.”

“record-queue-time-avg will indicate how long messages take to be produced, and therefore is Helpful when tuning these two configurations to meet the latency requirements for your application.”

This is the metric that tells you whether linger.ms is actually costing you what you think (Ch. 3 §5).

Per-broker and per-topic producer metrics

"useful for debugging problems in some cases, but they are NOT metrics that you are going to want to review on an ongoing basis. All of the attributes... are the same as the overall producer beans."

Per-broker — one metric stands out:

"The most useful metric provided by the per-broker producer metrics is request-latency-avg. This is because this metric will be MOSTLY STABLE (given stable batching) and CAN STILL SHOW A PROBLEM WITH CONNECTIONS TO A SPECIFIC BROKER. The other attributes... tend to VARY depending on WHAT PARTITIONS EACH BROKER IS LEADING. This means that what these measurements 'should' be can quickly change, depending on the state of the Kafka cluster."

Per-broker request-latency-avg is how you find the "one bad egg" from the client side.

Per-topic:

"only useful for producers working with MORE than one topic... only usable on a regular basis if the producer is NOT working with a LOT of topics. For example, a MirrorMaker could be producing HUNDREDS, OR THOUSANDS, of topics. It is difficult to review all of those metrics, and NEARLY IMPOSSIBLE TO SET REASONABLE ALERT THRESHOLDS on them." "record-send-rate and record-error-rate can be used to ISOLATE DROPPED MESSAGES TO A SPECIFIC TOPIC (or validated to be across all topics)." Plus a byte-rate per topic.

7.2 Consumer metrics

Five beans (Table 13-20):

ScopeJMX MBean
Overall consumerkafka.consumer:type=consumer-metrics,client-id=CLIENTID
Fetch managerkafka.consumer:type=consumer-fetch-manager-metrics,client-id=CLIENTID
Per-topic…consumer-fetch-manager-metrics,client-id=CLIENTID,topic=TOPICNAME
Per-brokerkafka.consumer:type=consumer-node-metrics,client-id=CLIENTID,node-id=node-BROKERID
Coordinatorkafka.consumer:type=consumer-coordinator-metrics,client-id=CLIENTID

"the overall consumer metric bean is LESS USEFUL for us because the metrics of interest are located in the FETCH MANAGER beans instead. The overall consumer bean has metrics regarding lower-level network operations, but the fetch manager bean has metrics regarding bytes, request, and record rates."

💡 "Unlike the producer client, the metrics provided by the CONSUMER are USEFUL TO LOOK AT BUT NOT USEFUL FOR SETTING UP ALERTS ON."

fetch-latency-avg — and why it's a poor alert

"As with request-latency-avg in the producer, this tells us how long fetch requests take. THE PROBLEM WITH ALERTING ON THIS METRIC IS THAT THE LATENCY IS GOVERNED BY THE CONSUMER CONFIGURATIONS fetch.min.bytes AND fetch.max.wait.ms. A SLOW TOPIC WILL HAVE ERRATIC LATENCIES, as sometimes the broker will respond quickly (when there are messages available), and sometimes it will not respond for fetch.max.wait.ms (when there are none). When consuming topics that have more regular, and abundant, message traffic, this metric may be more useful."

⚠️ WAIT! NO LAG?

*"The best advice for all consumers is that you MUST monitor the consumer lag. So why do we not recommend monitoring the records-lag-max attribute on the fetch manager bean? This metric shows the current lag for THE PARTITION THAT IS THE MOST BEHIND.

The problem is TWOFOLD:

  1. IT ONLY SHOWS THE LAG FOR ONE PARTITION, and
  2. IT RELIES ON PROPER FUNCTIONING OF THE CONSUMER.

If you have NO OTHER OPTION, use this attribute for lag and set up alerting for it. BUT THE BEST PRACTICE IS TO USE EXTERNAL LAG MONITORING."* (→ §8)

Traffic metrics — and a warning about alerting on them
 bytes-consumed-rate    bytes per second consumed by this client instance
 records-consumed-rate  messages per second

⚠️ "Some users set MINIMUM thresholds on these metrics for alerting so they are notified if the consumer is NOT DOING ENOUGH WORK. YOU SHOULD BE CAREFUL WHEN DOING THIS, HOWEVER. Kafka is intended to DECOUPLE the consumer and producer clients... The rate at which the consumer is able to consume messages IS OFTEN DEPENDENT ON WHETHER OR NOT THE PRODUCER IS WORKING CORRECTLY, so monitoring these metrics ON THE CONSUMER MAKES ASSUMPTIONS ABOUT THE STATE OF THE PRODUCER. THIS CAN LEAD TO FALSE ALERTS."

Relationship metrics:

fetch-rate
“number of Fetch requests per second”
fetch-size-avg
“average Size of those fetch requests in bytes”
records-per-request-avg
“average number of Messages in each fetch request”

⚠ “the consumer does Not provide an equivalent to the producer's record-size-avg. If this is important, You will need to infer it from the other metrics or Capture it in your application.”

💡 Consumer coordinator metrics — the rebalance detector

"The BIGGEST problem that consumers can run into due to coordinator activities is A PAUSE IN CONSUMPTION WHILE THE CONSUMER GROUP SYNCHRONIZES. This is when the consumer instances negotiate which partitions will be consumed by which client. Depending on the number of partitions, THIS CAN TAKE SOME TIME."

MetricMeaning and use
sync-time-avg"the average amount of time, in milliseconds, that the sync activity takes"
sync-rate"the number of group syncs that happen every second. FOR A STABLE CONSUMER GROUP, THIS NUMBER SHOULD BE ZERO MOST OF THE TIME."
commit-latency-avg"the average amount of time that offset commits take. You should monitor this value JUST AS YOU WOULD THE REQUEST LATENCY IN THE PRODUCER. It should be possible to establish a baseline and set reasonable thresholds for alerting." — because "these commits are essentially just produce requests (though they have their own request type), in that the offset commit is a message produced to a special topic"
assigned-partitions"a count of the number of partitions that the consumer client (as a single instance in the group) has been assigned. This is helpful because, WHEN COMPARED TO THIS METRIC FROM OTHER CONSUMER CLIENTS IN THE GROUP, IT IS POSSIBLE TO SEE THE BALANCE OF LOAD ACROSS THE ENTIRE CONSUMER GROUP. We can use this to identify imbalances that might be caused by problems in the algorithm used by the consumer coordinator."

sync-rate is the single best rebalance-storm detector (Ch. 4 §2), and assigned-partitions across instances is how you detect an unbalanced assignor like RangeAssignor (Ch. 4 §6.5).

7.3 Quota throttling metrics

"The Kafka broker DOES NOT USE ERROR CODES in the response to indicate that the client is being throttled. This means that IT IS NOT OBVIOUS TO THE APPLICATION THAT THROTTLING IS HAPPENING WITHOUT MONITORING THE METRICS."

ClientMBeanAttribute
Consumerkafka.consumer:type=consumer-fetch-manager-metrics,client-id=CLIENTIDfetch-throttle-time-avg
Producerkafka.producer:type=producer-metrics,client-id=CLIENTIDproduce-throttle-time-avg

💡 "Quotas are not enabled by default, but IT IS SAFE TO MONITOR THESE METRICS IRRESPECTIVE of whether you are currently using quotas. Monitoring them is a good practice as THEY MAY BE ENABLED AT SOME POINT IN THE FUTURE, and IT'S EASIER TO START WITH MONITORING THEM AS OPPOSED TO ADDING METRICS LATER."

This is the "my client is mysteriously slow and the broker looks fine" diagnosis. Throttling is invisible without these two metrics.


On this page