Learn Labs
7. Reliable Data Delivery

7.9 Self-test

Self-test31 questions

—/31
  1. State Kafka's reliability guarantees precisely. What are the qualifying clauses in the ordering guarantee and the durability guarantee?

  2. What exactly does "committed" mean — and what does it explicitly not mean?

  3. Give all three conditions for a replica to be in sync. Which one surprises people, and why?

  4. A replica falls out of the ISR and your p99 latency improves. Explain what just happened and why it's dangerous.

  5. Give the formula for replication traffic. Why is this the number Ch. 2 said people forget?

  6. Someone proposes RF=2 because the SAN already triple-replicates blocks. What's right and what's wrong about that reasoning?

  7. RF=3 and a single top-of-rack switch failure took a partition offline. How, and what's the fix?

  8. Describe both scenarios that leave a partition with no in-sync replica. Which one should change your default configuration, and how?

  9. Walk through the offset 100–200 inconsistency caused by unclean leader election. What happens to the old leader's data when it returns?

  10. Why does acks=all provide almost no protection with min.insync.replicas=1?

  11. What exactly happens when in-sync replicas drop below min.insync.replicas? What can consumers still do?

  12. Explain why "a produce request that fails loudly" is better than "an acknowledged write you'll lose."

  13. Kafka doesn't fsync on every write. State the bet it's making, and the condition under which that bet is bad.

  14. replica.lag.time.max.ms went from 10 s to 30 s in 2.5.0. Name the benefit and the hidden cost.

  15. With perfect broker configuration, describe the two ways a producer can still lose data.

  16. Why is acks=0 popular in benchmarks and misleading as a latency result?

  17. LEADER_NOT_AVAILABLE vs INVALID_CONFIG: which is retriable, and what determines the category?

  18. Producer retries give you which delivery guarantee? What upgrades it, and how?

  19. Name four error categories the producer will not handle for you.

  20. When is autocommit actually safe, and what single change makes it unsafe?

  21. Why are offset commits more expensive than they look, and why does the cost concentrate on one broker?

  22. Record 30 failed, 31 succeeded. Why can't you commit 31? Give both remediation patterns and their trade-offs.

  23. Your consumer restarts at the correct offset but produces wrong aggregates. What's missing, and what are the two proper solutions?

  24. What do VerifiableProducer/VerifiableConsumer do, and what's the pass criterion for a failure test?

  25. List the four configuration-validation scenarios. Which one produces a number you must feed into a producer config?

  26. What is a "brown out," and why is it harder to handle than an outright failure?

  27. Why is "no more than 1,000 duplicate values" a better test expectation than "no duplicates"?

  28. Why are static thresholds on consumer lag a bad alert design, and what's the alternative?

  29. What must you instrument to prove no messages were lost end to end? What's the state of open-source tooling for it?

  30. FailedProduceRequestsPerSec is rising. Name one benign cause and one serious one, and how you'd tell them apart.

  31. Draw the reliability dependency chain from replication.factor through to transactions. Pick any one link and explain what breaks if you omit it.

    Previous: Chapter 6 — Kafka Internals Next: Chapter 8 — Exactly-Once Semantics