9.3 Connect vs producer/consumer clients — the decision rule
| Use Kafka clients (producer/consumer) when… | Use Kafka Connect when… |
|---|---|
| “you can modify the code of the application that you want to connect … and when you want to either push data into Kafka or pull data from Kafka.” | “to connect Kafka to datastores that you did not write and whose code or APIs you cannot or will not modify.” |
| ⇒ the client is embedded in your own application | ⇒ “users of Kafka Connect only need to write configuration files” |
And if no connector exists yet? Still prefer Connect.
"Connect is recommended because it provides out-of-the-box features like:
- configuration management
- offset storage
- parallelization
- error handling
- support for different data types
- standard management REST APIs
Writing a small app that connects Kafka to a datastore SOUNDS SIMPLE, but there are MANY LITTLE DETAILS you will need to handle concerning data types and configuration that make the task nontrivial. What's more, you will need to MAINTAIN this pipeline app and DOCUMENT it, and your TEAMMATES WILL NEED TO LEARN HOW TO USE IT."
The most persuasive version of this argument appears later in the chapter:
"Experienced developers know that writing code that reads data from Kafka and inserts it into a database takes maybe A DAY OR TWO, but if you need to handle configuration, errors, REST APIs, monitoring, deployment, scaling up and down, and handling failures, IT CAN TAKE A FEW MONTHS to get everything right. And most data integration pipelines involve more than just the one source or target. So now consider that effort spent on bespoke code for just a database integration, repeated many times for other technologies."