Learn Labs
9. Building Data Pipelines (Kafka Connect)

9.1 What problem does Kafka solve in a data pipeline?

This is Ch. 1's N×M problem restated at the integration tooling layer.

Two shapes of pipeline

① Kafka as one end point
KafkaS3MongoDBKafka
② Kafka as an intermediary between two other systems
TwitterKafkaElasticsearch
Figure 9.1.1Two shapes of pipeline

The core value proposition, in one sentence

"The main value Kafka provides to data pipelines is its ability to serve as a very large, reliable BUFFER between various stages in the pipeline. This effectively decouples producers and consumers of data within the pipeline and allows use of the SAME DATA from the source in MULTIPLE target applications and systems, all with DIFFERENT timeliness and availability requirements."

PUTTING DATA INTEGRATION IN CONTEXT

*"Some organizations think of Kafka as an end point of a pipeline. They look at questions such as 'How do I get data from Kafka to Elastic?' This is a valid question to ask... But we are going to start the discussion by looking at the use of Kafka within a larger context that includes at least two (and possibly many more) end points that are not Kafka itself.

We encourage anyone faced with a data-integration problem to consider the bigger picture and not focus only on the immediate end points. FOCUSING ON SHORT-TERM INTEGRATIONS IS HOW YOU END UP WITH A COMPLEX AND EXPENSIVE-TO-MAINTAIN DATA INTEGRATION MESS."*

This is Ch. 1's N×M problem restated at the integration tooling layer. Each ad-hoc "get data from A to B" tool is a point-to-point connection in disguise.


On this page