9.12 Self-test
Self-test38 questions
What are the two pipeline shapes involving Kafka, and why does the book insist you consider the bigger picture?
How does Kafka decouple timeliness requirements? Give the batch-consumer example.
Where does back pressure come from in a Kafka pipeline, and what mechanism does the broker use?
When can a Kafka-based pipeline be exactly-once end to end? What does Connect contribute?
Why does Kafka-as-buffer eliminate the need for a complex back-pressure mechanism?
Kafka and Connect are "completely agnostic" about data formats. What component makes that true?
Describe a "great pipeline"'s behavior when someone adds a column in MySQL.
Give one example each of push vs pull sources, and append-only vs updatable sinks.
Define ETL and ELT. Give the main drawback of each, with a concrete cost.
Which transformations should happen in the pipeline? State the test.
List the five security questions for a data pipeline. Which one is specific to crossing datacenters?
Why shouldn't connector credentials live in config files, and what does Connect provide instead?
Which failure-handling question determines your retention setting, and why?
Name the three ways coupling sneaks into a pipeline, with the concrete failure each produces.
State the decision rule: Connect vs producer/consumer clients.
No connector exists for your datastore. Why is Connect still recommended? Give the "day or two vs a few months" argument.
Draw a Connect cluster: workers, connectors, tasks, converters, internal topics.
Two workers have the same
group.id. What does that mean? What happens when one crashes, and which protocol does that use?Explain
plugin.pathlayout. What's the one thing that "will not work," and why is the classpath alternative discouraged?When would you use standalone mode instead of distributed?
What differs between the FileStreamSource and FileStreamSink configs? Why is one plural?
Why must you not use FileStream connectors in production?
How do you discover a connector's available configuration options without reading documentation?
Contrast JDBC polling with log-based CDC on four dimensions. What does the book recommend, and why?
Why is
key.ignore=trueneeded for the Elasticsearch sink, and what does it make possible?What are SMTs for, and where does the line fall between SMTs and Kafka Streams?
Give the
transforms.*configuration naming convention.Which SMT would you use to remove PII? To route by timestamp? To detect tombstones?
What is
error.tolerance, and which connectors can use it?State the three responsibilities of a connector (as opposed to a task). How does the JDBC source decide its task count?
What does the source task context provide? What does the sink task context provide, and why is one of those items critical for exactly-once?
State the one-sentence division of labor between connectors/tasks and workers.
What are logical partitions and offsets? Give the file and JDBC examples.
Why is the choice of source partitioning and offset tracking the most important design decision in a source connector?
In what order does the worker send data and store offsets, and why does the order matter?
Name the three internal Connect topics and what each holds. How should they be configured?
When would you choose Flume/Logstash over Connect? A GUI ETL tool? A stream processing framework — and what do you give up?
What single feature does the chapter say matters most in any integration system, and what does it tell you to do about it?
Previous: Chapter 8 — Exactly-Once Semantics Next: Chapter 10 — Cross-Cluster Data Mirroring