14.10 Self-test
Self-test52 questions
Define a data stream. What single word is doing most of the work?
Give the three attributes of event streams beyond unboundedness, and contrast each with a database table.
Explain the deposit/withdrawal example. What does it prove about ordering?
A canceled transaction — what happens to the original event? What's the redo-log analogy?
Why does the book credit replayability specifically for Kafka's success in stream processing?
What does the definition of stream processing deliberately say nothing about?
Contrast the three programming paradigms on latency, blocking behavior, and database analogue.
What is the "gap" that stream processing fills, stated in concrete numbers?
A job runs at 2 a.m., reads 500 records, outputs a result, and exits. Is it stream processing? Why not?
Name the three notions of time. Which matters most, which should be avoided, and why?
When is log-append time an acceptable substitute for event time?
Give all four rules Kafka Streams uses to assign output timestamps.
Why must the whole pipeline share one time zone, and what do you do if it can't?
Why is storing state in a local variable unreliable? (The book admits doing this — where?)
Contrast local and external state on speed, size, sharing, and availability.
What single design consequence follows from local state's memory limit?
State the stream-table duality. What can a table answer that a stream can't, and vice versa?
How do you convert a table to a stream? A stream to a table? What's the name for the second operation?
Name the three window dimensions. What's the trade-off in window size?
Define hopping, tumbling, and session windows.
What's a grace period, and what question does it answer?
Why does key-based partitioning make local state correct? Name the two Kafka guarantees involved.
Name the three problems local state creates and how Kafka Streams solves the second one (two mechanisms).
Why is log compaction essential to the changelog-topic approach?
Why can't a local-state aggregate compute the daily top 10? Describe the multiphase solution and why phase 2 can be single-instance.
What does Kafka Streams do that MapReduce didn't, regarding multiple reduce phases?
Give the three problems with per-record external lookups, with numbers.
State the caching dilemma, and how CDC resolves it.
Why is a table-table join never windowed?
Why is a stream-stream join necessarily windowed? What must be true of the two topics?
How does Kafka Streams guarantee one task sees both sides of a join for a given key?
Give three real-world causes of out-of-sequence events.
List the four things an app must do to handle late events. Which one is the fundamental difference from batch?
How does Kafka Streams implement late-window correction? What Kafka feature does it rely on?
Give both reprocessing variants and the three steps of the recommended one. Why is it "much safer"?
What are interactive queries, and when are they worth it?
What does
APPLICATION_ID_CONFIGcontrol besides coordination?What does
groupByKey()actually do? When is it a no-op?Why do you aggregate sum-and-count rather than average?
Why does the windowed Serde need the window size even though it isn't serialized?
Why is "just start multiple instances" a significant claim? What does it save you compared with other frameworks?
Name the three processor kinds in a topology. What must a topology start and end with?
Give the three execution steps of a Kafka Streams app. Which one admits optimization, and what's the easy mistake when enabling it?
What does
TopologyTestDrivernot simulate? Which integration framework is recommended, and why?What determines the number of tasks? What are the two ways to scale, and what caps both?
How does Kafka Streams implement a shuffle, and why is that architecturally significant?
Why can the two sides of a repartition run fully in parallel despite one depending on the other?
Kafka Streams always recovers. So what is the actual availability problem, and what are the two specific fixes?
Why does segment size affect recovery time? (Reference the active-segment rule.)
What is beaconing, and why is it a good fit for stream processing?
For each of the four application types, say whether stream processing is the right answer and what to look for.
Give the four global framework-selection criteria. Which one is about leaky abstractions?
Previous: Chapter 13 — Monitoring Kafka Next: Appendices A & B · Back to the index