Learn Labs
3. Kafka Producers: Writing Messages to Kafka

3.0 The motivating example (keep this in your head)

A credit-card transaction processing system:

A credit-card transaction processing system:

transaction / approve–denytransaction / responseboth streamsOnline store(producer)KAFKARules engineapprove / denyAnalytics app → databaseanalysts improve the rules

The transaction goes to the rules engine, the approve/deny response comes back through Kafka to the online store, and both streams land in the analytics database.

Figure 3.0.1A credit-card transaction processing system

Its requirements, stated precisely:

  • Never lose a single message. Never duplicate any message.
  • Latency low, but up to 500 ms tolerable.
  • Very high throughput — up to a million messages/second.

Contrast: website click tracking.

  • Some loss and a few duplicates are tolerable.
  • Latency can be high, as long as the user experience isn't affected — "we don't mind if it takes a few seconds for the message to arrive at Kafka, as long as the next page loads immediately after the user clicks."
  • Throughput depends on anticipated activity.

This is the chapter's core lesson: "The different requirements will influence the way you use the producer API to write messages to Kafka and the configuration you use." There is no universally correct producer config. There is only a config that matches a stated requirement.

The questions to answer before configuring a producer:

  1. Is every message critical, or can we tolerate loss?
  2. Are we OK with accidentally duplicating messages?
  3. Are there strict latency or throughput requirements?

Third-party clients

Kafka has a binary wire protocol — applications can read/write simply by sending the correct byte sequences to Kafka's network port. Multiple clients implement it in C++, Python, Go, and many more. These are not part of the Apache Kafka project; a list of non-Java clients is maintained in the project wiki.


On this page