Learn Labs
13. Monitoring Kafka

13.12 Self-test

Self-test56 questions

—/56
  1. Why is an in-process JMX agent recommended over an open remote JMX port?

  2. List the five metric sources in order of objectivity. Which are best for SLIs, and which for debugging?

  3. Give the website analogy for why broker metrics alone are insufficient.

  4. Contrast the retention and character of alerting, debugging, and historical metrics. What extra metadata does the last one need?

  5. Why is "collect everything continuously" wrong for debugging metrics?

  6. Explain the "Check Engine light" principle and the specific failure mode it prevents.

  7. Give both application health-check methods and the drawback of the second.

  8. Define SLI, SLO, and SLA precisely. What's an OLA?

  9. Why does the book say most engineers should be setting SLOs, not SLAs — and why bother if you have no external customers?

  10. What makes a good SLI? Why do quantile metrics fail as SLIs, and what's the compromise?

  11. Name the five SLI types.

  12. Why shouldn't you accept an SLO on data freshness?

  13. Why can't you alert on an SLO directly? Explain the burn-rate technique with the 0.1% / 0.4% / 2% example.

  14. "Who watches the watchers?" — state the problem and both solutions.

  15. Name the three categories of cluster problem. Which produces "that's really weird," and which two metrics cover it?

  16. What single action should you always take before diagnosing further, and why is it safe?

  17. Why does the book retract its previous advice about URP alerting? What is URP still good for?

  18. A steady URP count reported by many brokers means what? A fluctuating one?

  19. Describe the "common broker" technique. What does no common broker tell you?

  20. List the five metrics for detecting cluster imbalance, and what balanced looks like.

  21. Why is disk utilization not a performance bottleneck metric for Kafka?

  22. Why does exhausting any resource show up as under-replicated partitions? What does that imply about your clients?

  23. Explain the "one bad egg" cascade in full. Why does good partition balance make it worse?

  24. What is a BBU, and what's its silent failure mode?

  25. Name three soft hardware failures that degrade rather than break.

  26. What does ActiveControllerCount summing to 2 mean? To 0? What's the fix for each, and what complication should you expect?

  27. What are the two thresholds for request handler idle ratio? Why is this pool more critical than the network threads?

  28. How many request handler threads should you configure, and why is that enough given how many requests exist?

  29. Describe the pre-0.10 recompression tax and what the v2 format changed. Why is fixing it "one of the single largest performance improvements"?

  30. Give the seven attributes of a rate metric. Which is best for spikes? Which should you not alert on?

  31. Why can bytes-out equal bytes-in with zero consumers? Give the formula.

  32. Why is there no "messages out" metric?

  33. Why is leader count more alert-worthy than partition count? What derived percentage should you compute, and what's the expected value?

  34. Why does OfflinePartitionsCount read 0 on most brokers, and how must you aggregate it?

  35. Name all seven request timing phases and what a spike in each one points to.

  36. Which request type's p99.9 total time makes the best broker-side latency alert, and why not Fetch?

  37. What does a size discrepancy between partitions of the same topic indicate?

  38. Why isn't LogEndOffset − LogStartOffset the message count?

  39. Which two JVM OperatingSystem attributes matter, and what consumes file descriptors?

  40. What exactly does Linux system load average count? What value means 100% on a 24-CPU box?

  41. Why is heap memory relatively unimportant for a broker, and what should you still watch?

  42. Which three loggers should you enable at DEBUG by default, and what invisible failure do they cover?

  43. Which two loggers should get their own files, and why?

  44. Name the two producer metrics that deserve alerts and what each one means.

  45. What are the three views of producer traffic volume, and how do requests, batches, messages, and bytes nest?

  46. What does record-queue-time-avg measure, and which two configs does it help you tune?

  47. Why is per-broker request-latency-avg the most useful per-broker producer metric?

  48. Why are consumer metrics "useful to look at but not useful for alerting"?

  49. Give both reasons records-lag-max is a poor lag metric.

  50. Why is alerting on a minimum records-consumed-rate risky?

  51. Name the four coordinator metrics and what each detects. What should sync-rate normally be?

  52. Why is quota throttling invisible without specific metrics, and which two are they?

  53. Describe the correct architecture for lag monitoring, and the two reasons the CLI approach doesn't scale.

  54. How does Burrow avoid needing thresholds?

  55. What two questions does end-to-end monitoring answer, and why can't the broker answer them itself?

  56. State the recurring theme: which three critical facts about Kafka cannot come from broker metrics?

    Previous: Chapter 12 — Administering Kafka Next: Chapter 14 — Stream Processing