9. Building Data Pipelines (Kafka Connect)
9. Building Data Pipelines (Kafka Connect)
Chapter 9 of Kafka: The Definitive Guide — 12 sections.
Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 9 Why Connect exists (added in Kafka 0.9): "We noticed that there were specific challenges in integrating Kafka into data pipelines that EVERY ORGANIZATION had to solve, and decided to add APIs to Kafka that solve some of those challenges rather than force every organization to figure them out from scratch."
Sections
- 9.11What problem does Kafka solve in a data pipeline?This is Ch. 1's N×M problem restated at the integration tooling layer.
- 9.213The eight considerations when building data pipelinesThis is exactly the workaround Ch. 8 §3.6 described ("manage offsets in the database") — Connect gives connectors a first-class hook for it.
- 9.31Connect vs producer/consumer clients — the decision rule
- 9.42Kafka Connect architecture
- 9.51Example: file source → file sinkResponse includes "tasks": [{"connector":"load-kafka-config","task":0}] and "type":"source".
- 9.62Example: MySQL → Kafka → ElasticsearchThree options: Confluent Hub client, download from Confluent Hub (or wherever the connector is hosted), or build from source:
- 9.72Single Message Transformations (SMTs)
- 9.86A deeper look at Connect internalsNote: "The JSON converter can be configured to either include a schema in the result record or not — so we can support both structured and semistructured data."
- 9.9Alternatives to Kafka Connect
- 9.10Failure catalogWhat actually breaks in production — Ch. 9 consolidated
- 9.112Deploy / monitor / scale / recover
- 9.12Self-testSelf-test