Learn Labs
4. Kafka Consumers: Reading Data from Kafka

4.1 What problem do consumer groups solve?

Group assignment

assignor
0
idle consumers
1
load spread
consumer 0
P0P1P2
consumer 1
P3P4P5
consumer 2
P6P7
Safe

Balanced: every consumer owns within one partition of every other.

Partitions are the unit of parallelism: a partition is read by exactly one consumer in the group, so consumers beyond the partition count sit idle no matter how much hardware you add.

The starting point: an application reads from a topic, runs validations, writes results to another data store. One consumer object, one subscription. Works fine.

Then: "what if the rate at which producers write messages to the topic exceeds the rate at which your application can validate them? If you are limited to a single consumer reading and processing the data, your application may fall further and further behind, unable to keep up with the rate of incoming messages."

Why this is the normal case, not an edge case:

"It is common for Kafka consumers to do high-latency operations such as write to a database or a time-consuming computation on the data. In these cases, a single consumer can't possibly keep up with the rate data flows into a topic."

The solution: "Just like multiple producers can write to the same topic, we need to allow multiple consumers to read from the same topic, splitting the data among them."

The scaling ladder — memorize these four states

① One consumer
p0p1p2p3C1gets ALL
② Two consumers
p0p1p2p3C1p0, p2C2p1, p3e.g. p0, p2 → C1; p1, p3 → C2
③ Four consumers — optimal
p0p1p2p3C1C2C3C4
④ Five consumers — one wasted
p0p1p2p3C1C2C3C4C5IDLE — gets NOTHING
Figure 4.1.1The scaling ladder — memorize these four states

"The main way we scale data consumption from a Kafka topic is by adding more consumers to a consumer group. ... This is a good reason to create topics with a large number of partitions — it allows adding more consumers when the load increases. Keep in mind that there is no point in adding more consumers than you have partitions in a topic — some of the consumers will just be idle."

This closes the loop with Ch. 2: partition count is your permanent ceiling on consumer parallelism, and partitions can only be added (which breaks keyed routing, per Ch. 3). Hence: size partitions for future consumer parallelism, up front.

Multiple groups = multiple independent readers

Group G1 — two consumers
p0p1p2p3C1C2the four partitions are split between the two members
Group G2 — one consumer
p0p1p2p3C1gets ALL 4one member, so it owns every partition

Each group independently gets ALL messages. G2 gets everything regardless of what G1 is doing.

Figure 4.1.2Multiple groups = multiple independent readers

"One of the main design goals in Kafka was to make the data produced to Kafka topics available for many use cases throughout the organization. ... To make sure an application gets all the messages in a topic, ensure the application has its own consumer group. Unlike many traditional messaging systems, Kafka scales to a large number of consumers and consumer groups without reducing performance."

The rule, stated once

  • NEW consumer GROUP → for each APPLICATION that needs ALL messages.
  • NEW consumer in a group → to SCALE reading/processing within one app (each additional member gets a SUBSET).

On this page