Learn Labs
13. Monitoring Kafka

13.2 Service-level objectives

"One area of monitoring that is especially critical for INFRASTRUCTURE SERVICES, such as Kafka... This is how we communicate to our clients what level of service they can expect. The clients want to be able to treat services like Kafka as an OPAQUE SYSTEM: they do not want or need to understand the internals — only the interface they are using and knowing it will do what they need it to do."

2.1 The definitions — used incorrectly constantly

"Frequently, you will hear engineers, managers, executives, and everyone else use terms in the 'service-level' space INCORRECTLY, which leads to confusion about what is actually being talked about."

TermWhat it isExample
SLI — service-level indicator“A metric that describes one aspect of a service's reliability.” It “should be closely aligned with your client's experience, so it is usually true that the more objective these measurements are, the better.” “In a request processing system, such as Kafka, it is usually best to express these as a ratio between the number of good events and the total number of events.”“The proportion of requests to a web server that return a 2xx, 3xx, or 4xx response.”
SLO — service-level objective (a.k.a. SLT, service-level threshold)“Combines an SLI with a target value.” “A common way to express the target is by the number of nines (99.9% is ‘three nines’), though it is by no means required.” “The SLO should also include a time frame that it is measured over, frequently on the scale of days.”“99% of requests must return a 2xx/3xx/4xx response over 7 days.”
SLA — service-level agreement“A contract between a service provider and a client.” It “usually includes several SLOs, as well as details about how they are measured and reported, how the client seeks support, and penalties that the service provider will be subject to.”—

OPERATIONAL-LEVEL AGREEMENT (OLA)

"Less frequently used. It describes agreements between MULTIPLE INTERNAL SERVICES or support providers in the overall delivery of an SLA. The goal is to assure that the multiple activities necessary to fulfill the SLA are properly described and accounted for in day-to-day operations."

💡 The practical correction: "It is very common to hear people talk about SLAs when they really mean SLOs. ... it is RARE that the engineers running the applications are responsible for anything more than the performance of that service WITHIN THE SLOs. ... those who only have internal clients generally do not have SLAs with those internal customers. THIS SHOULD NOT PREVENT YOU FROM SETTING AND COMMUNICATING SLOs, HOWEVER, as doing that will lead to FEWER ASSUMPTIONS BY CUSTOMERS as to how they think Kafka should be performing."

2.2 What makes a good SLI

*"In general, the metrics for your SLIs should be gathered using something EXTERNAL to the Kafka brokers. ... YOUR CLIENTS DO NOT CARE IF YOU THINK YOUR SERVICE IS RUNNING CORRECTLY; IT IS THEIR EXPERIENCE (IN AGGREGATE) THAT MATTERS.

This means: infrastructure metrics are OK, synthetic clients are GOOD, and CLIENT-SIDE METRICS ARE PROBABLY THE BEST for most of your SLIs."*

The five common SLI types (Table 13-2):

TypeQuestion
Availability"Is the client able to make a request and get a response?"
Latency"How quickly is the response returned?"
Quality"Does the response include a proper response?"
Security"Are the request and response appropriately protected, whether that is authorization or encryption?"
Throughput"Can the client get enough data, fast enough?"

⚠️ 2.3 Why quantiles make bad SLIs

*"it is usually better for your SLIs to be based on A COUNTER OF EVENTS THAT FALL INSIDE THE THRESHOLDS of the SLO. This means that ideally, each event would be individually checked to see if it meets the threshold.

THIS RULES OUT QUANTILE METRICS AS GOOD SLIs, as those will only tell you that 90% of your events were below a given value WITHOUT ALLOWING YOU TO CONTROL WHAT THAT VALUE IS."*

💡 The bucket compromise: "aggregating values into BUCKETS (e.g., 'less than 10 ms,' '10–50 ms,' '50–100 ms') can be useful, ESPECIALLY WHEN YOU ARE NOT YET SURE WHAT A GOOD THRESHOLD IS. This will give you a view into the DISTRIBUTION of events within the range of the SLO, and you can configure the buckets so that the boundaries are reasonable values for the SLO threshold."

A quantile tells youA bucketed counter tells you
p90 = 47 ms<10 ms: 994,231 · 10–50 ms: 5,600 · 50–100 ms: 150 · >100 ms: 19
You can't ask “how many were over 10 ms?” — the number 10 isn't in it.“999 events breached the 10 ms threshold” — directly answerable, and you can change the threshold later without changing instrumentation.

CUSTOMERS ALWAYS WANT MORE

*"There are some SLOs that your customers may be interested in that are important to them but NOT WITHIN YOUR CONTROL. For example, they may be concerned about the CORRECTNESS or FRESHNESS of the data produced to Kafka.

DO NOT AGREE TO SUPPORT SLOs THAT YOU ARE NOT RESPONSIBLE FOR, as that will only lead to taking on work that DILUTES THE CORE JOB of keeping Kafka running properly. Make sure to connect them with the proper group."*

💡 2.4 Using SLOs for alerting — the burn-rate technique

*"SLOs should inform your PRIMARY ALERTS. ... Generally speaking, IF A PROBLEM DOES NOT IMPACT YOUR CLIENTS, IT DOES NOT NEED TO WAKE YOU UP AT NIGHT.

SLOs will also tell you about the problems that YOU DON'T KNOW HOW TO DETECT because you've never seen them before. THEY WON'T TELL YOU WHAT THOSE PROBLEMS ARE, BUT THEY WILL TELL YOU THAT THEY EXIST."*

The problem with alerting on the SLO directly:

*"SLOs are best for LONG TIMESCALES, such as a week, as we want to report them to management and customers... In addition, BY THE TIME THE SLO ALERT FIRES, IT'S TOO LATE — YOU'RE ALREADY OPERATING OUTSIDE OF THE SLO.

The BEST way to approach using SLOs for alerting is to OBSERVE THE RATE AT WHICH YOU ARE BURNING THROUGH YOUR SLO over its timeframe."*

The worked example — follow the arithmetic:

alerts go offnormal — 0.1%/hourTue 10:00 → 0.4%/hour: a ticket2%/hour — would breach by Friday lunchtimeback to 0.4%/hour after ~4 hoursburn rate, % of budget per hourthe SLO week
  • Setup. 1,000,000 requests/week; SLO: 99.9% of requests send the first byte of response within 10 ms → budget: up to 1,000 slow requests per week. Normal is ~1 slow request/hour ≈ 168/week → normal burn rate 0.1% per hour.
  • Tuesday 10:00 — 0.4%/hour. “This isn't great, but it's still not a problem because you'll be well within the SLO by the end of the week.” Open a ticket, go back to higher-priority work. No page.
  • Wednesday 14:00 — 2%/hour. “Your alerts go off. You know that at this rate, you'll breach the SLO by lunchtime on Friday.” Drop everything and diagnose; after ~4 hours the burn rate is back to 0.4%/hour and stays there.
  • “By using the burn rate, you were able to avoid breaching the SLO.”
Figure 13.2.3The worked example — follow the arithmetic
SLO burn rate% of budget per hour0.1%/hr — normalno action0.4%/hr — a ticketwill not breach2.0%/hr — a PAGEwill breach in ~2 days

The same metric produces three different responses depending on the rate, not the value. A static threshold cannot express any of this.

Figure 13.2.42.4 Using SLOs for alerting — the burn-rate technique

Further reading the book recommends: Site Reliability Engineering and The Site Reliability Workbook, both ed. Betsy Beyer et al. (O'Reilly).


On this page