4. Kafka Consumers: Reading Data from Kafka
4. Kafka Consumers: Reading Data from Kafka
Chapter 4 of Kafka: The Definitive Guide — 15 sections.
Source: Kafka: The Definitive Guide, 2nd Ed., Ch. 4 Learning goal: "Reading data from Kafka is a bit different than reading data from other messaging systems." The two things that make it different — consumer groups + rebalancing and offset commits — are also the two things that cause virtually every consumer bug in production. Master those and the API is trivial.
Sections
- 4.13What problem do consumer groups solve?
- 4.24Rebalancing — the mechanism, and why it hurts
- 4.32Static group membership — avoiding rebalance on restart
- 4.4Creating a consumerThat last point often blocks regex subscriptions outright in locked-down multi-tenant clusters — you cannot grant least-privilege topic access and use regex subscription.
- 4.52The poll loopThe timeout parameter controls how long poll() will block if data is not available in the consumer buffer.
- 4.68Configuring consumersMinimum data the consumer wants to receive from the broker per fetch.
- 4.76Commits and offsets — where consumers actually breakGet this wrong by one and you reprocess one record per commit forever (harmless-ish) or skip one record per commit (data loss).
- 4.8Rebalance listenersPass a ConsumerRebalanceListener to subscribe().
- 4.9Consuming from specific offsets① Map every partition assigned to this consumer (consumer.assignment()) to the target timestamp.
- 4.102Exiting cleanly — `wakeup()` and `close()`Skip close() and you pay session.timeout.ms of dead air on every deploy — across every partition that consumer owned.
- 4.11Deserializers① KafkaAvroDeserializer deserializes the Avro messages.
- 4.122Standalone consumer — `assign()` instead of `subscribe()`
- 4.13Failure catalogWhat actually breaks in production — Ch. 4 consolidated
- 4.143Deploy / monitor / scale / recoverThe consumer's recovery story is offset manipulation:
- 4.15Self-testSelf-test