9. Building Data Pipelines (Kafka Connect)
9.9 Alternatives to Kafka Connect
| Alternative | When it makes sense | The drawback |
|---|---|---|
| Ingest frameworks for other datastores — Flume (Hadoop), Logstash / Fluentd (Elasticsearch) | "If you are actually building a Hadoop-centric or Elastic-centric system and Kafka is just one of many inputs into that system" | "We recommend Kafka's Connect API when Kafka is an integral part of the architecture and when the goal is to connect large numbers of sources and sinks." |
| GUI-based ETL tools — Informatica, Talend, Pentaho, Apache NiFi, StreamSets | "if you are already using these systems... you may not be interested in adding another data integration system just for Kafka. They also make sense if you are using a GUI-based approach." | "usually built for involved workflows and will be a somewhat heavy and involved solution if all you want to do is get data in and out of Kafka. We believe that data integration should focus on FAITHFUL DELIVERY OF MESSAGES UNDER ALL CONDITIONS, while most ETL tools add UNNECESSARY COMPLEXITY." |
| Stream processing frameworks | "If your destination system is supported and you already intend to use that framework to process events from Kafka" — "often saves a step (no need to store processed events in Kafka — just read them out and write them to another system)" | "it can be more difficult to troubleshoot things like LOST AND CORRUPTED MESSAGES." |
The bigger framing:
"We do encourage you to look at Kafka as a platform that can handle data integration (with Connect), application integration (with producers and consumers), and stream processing. Kafka could be a viable REPLACEMENT for an ETL tool that only integrates data stores."