1.6 What actually breaks in production (Ch. 1 foundations)
Chapter 1 is conceptual, but it plants several landmines that detonate later.
Chapter 1 is conceptual, but it plants several landmines that detonate later. Naming them now makes the rest of the book read as consequences rather than trivia.
6.1 Assuming topic-wide ordering
Symptom: events processed out of order; state machines end up in impossible states; "the update arrived before the create."
Cause: ordering holds per partition, not per topic. Default producer behavior spreads messages evenly across all partitions.
Fix: key by the entity whose ordering you need (user_id, account_id, trace_id) so all its events land in one partition.
6.2 Changing partition count on a keyed topic
Symptom: after adding partitions, a key's messages start landing in a different partition than its history. Ordering silently breaks; stateful consumers see split history; compacted-topic semantics get weird.
Cause: hash(key) mod N — change N, change the mapping for most keys. The guarantee is explicitly conditional: same partition "provided that the partition count does not change."
Also: partition count can only be increased, never decreased.
Fix: size partitions for expected future throughput, not current. This is why Ch. 2 says to calculate based on future usage when keying.
6.3 Assuming offsets are contiguous
Symptom: gap-detection logic false-alarms; expected == last + 1 assertions fail.
Cause: offsets increase but "not necessarily monotonically greater" — gaps are legal (compaction, transaction markers, aborted transactions).
Fix: treat offsets as opaque, increasing cursors. Never arithmetic.
6.4 Schema-less topics → coupled deployments
Symptom: you cannot add a field without a coordinated multi-team deploy in a strict order; consumers crash on unknown fields; a producer rollback breaks consumers. Cause: no schema contract; writing and reading are tightly coupled. Fix: schema registry + a format with real compatibility rules (Avro). This is the lesson from LinkedIn's XML tracking system breaking "constantly due to changing schemas."
6.5 Retention shorter than your recovery time
Symptom: a consumer is down for maintenance longer than retention; on restart, its committed offset no longer exists; it either jumps to latest (silent data loss) or to earliest (re-processes everything). Cause: retention defines a minimum window, and it is a deletion policy, not an archive. Fix: retention ≥ worst-case consumer outage + recovery, with margin. Alert on consumer lag approaching the retention edge.
6.6 Treating MirrorMaker as intra-cluster replication
Symptom: people expect cross-DC failover to be as seamless as broker failover. It isn't — offsets are not identical across clusters, and MirrorMaker is asynchronous. Cause: Kafka's replication is explicitly intra-cluster only. MirrorMaker is a consumer+producer pair, i.e. an application with its own lag, its own failure modes, and its own offsets. Fix: design cross-cluster topologies deliberately (Ch. 10), and never assume offset equivalence between clusters.
6.7 The ActiveMQ lesson — coupling telemetry to serving
Symptom (historical, and still a live risk): the broker pauses → client connections back up → the application can no longer serve user requests. Cause: a messaging system that applies backpressure into a synchronous serving path. Fix / design rule: producers must be async and bounded; a telemetry path must be allowed to drop rather than block the request thread. Kafka's push-pull split exists precisely so slow consumers can't stall producers — but you can still recreate the bug yourself with a synchronous, unbounded, blocking producer call in a request handler.
6.8 Colocating Kafka with other memory-hungry apps
Foreshadowed here, explicit in Ch. 2: Kafka's read performance comes from the OS page cache. Anything else on the box competing for page cache degrades consumer performance. (Details in Ch. 2.)
1.5 Use cases (with the "why Kafka specifically" for each)
The win: avoids duplicating this logic in every app, and enables aggregation that would not otherwise be possible (you can't batch a user's notifications if each app sends its own…
1.7 Deployment / monitoring / scaling / backup — what Ch. 1 establishes
These are covered properly in Ch. 2, 7, 10, 12, 13, but Ch. 1 sets the frame: