6.10 Self-test
Self-test35 questions
Name the three causes of a broker's ZK ephemeral node disappearing. Why does it matter that they're indistinguishable?
A broker dies permanently. What's the fastest way to restore its partition assignments, and why does it work?
Walk the controller election protocol. What ZooKeeper primitive guarantees a single controller?
Trace the zombie-controller scenario end to end. What mechanism fences it, and what ZooKeeper operation makes the epoch safe?
Why does controller restart get slower as a cluster grows? Which KRaft change fixes it?
List the four problems that motivated KRaft, and which one is a human problem rather than a technical one.
In KRaft, do brokers get pushed metadata or pull it? Name two benefits of that direction.
What is the "fenced" broker state, and which specific bug does it eliminate?
Followers use the same Fetch requests consumers use. What does the leader infer from a fetch at offset N, and why is that elegant?
State both conditions under which a replica is declared out of sync. What config controls them?
Which replica is the preferred leader, and how do you identify it from
kafka-topics.shoutput? Why does this matter when reassigning replicas manually?Reading from followers saves money. What does it cost, and what mechanism causes that cost?
Draw the broker's threading model: acceptor, processor/network, request queue, I/O/handler, response queue, purgatory. What two situations put a response in purgatory?
Why must produce and fetch requests go to the leader? How does a client find out where that is, and what happens when it's wrong?
Name the three validations a leader runs on a produce request.
Complete the sentence and explain the implication: "Kafka does not wait for the data to get persisted to disk — it relies on ___ for message durability."
Explain zero-copy, and state the precondition in the storage format that makes it possible.
What is the high-water mark? Give the consistency argument for why consumers can't read past it, in terms of what a consumer would observe after a leader crash.
What is the worst-case extra consume latency caused by slow replication, and what bounds it?
How does the high-water mark explain Ch. 3's claim that end-to-end latency is identical for all
acksvalues?What is the fetch session cache for? What happens when it can't serve you, and how would you find out?
Why must you upgrade brokers before clients? Give the version-negotiation walkthrough.
What is message-format down-conversion, why is it expensive, and which two metrics reveal it?
What is the hard upper bound on a single partition's size, and why?
State the three goals of partition allocation, and how the rack-alternating list achieves the third.
Kafka picks a directory for a new partition by counting what? What does that mean the moment you add a fresh disk?
Why are partitions split into segments at all? Which segment can never be deleted, and what retention bug does that cause?
Why is the batch header large but the per-record header tiny? What does that imply about
linger.msand about partition fan-out per producer?What are control batches, who consumes them, and what does their existence explain about offsets?
Are index files worth protecting? What's the recovery procedure if one is corrupt, and what does it cost?
Give the offset-map arithmetic for compaction (bytes per entry, and the 1 GB segment example). Why might reducing cleaner threads fix a compaction failure?
What is a tombstone, and what's the precise failure if a consumer is offline for its entire retention window?
Contrast tombstones with
deleteRecords: mechanism, granularity, downstream visibility, and which topics each works on.Which two configs together make a compacted topic GDPR-capable, and what does each guarantee?
Why does tiered storage improve p99 latency in the historical-read scenario when it slightly worsens it in the steady-state one?
Previous: Chapter 5 — Managing Kafka Programmatically Next: Chapter 7 — Reliable Data Delivery