Learn Labs
1. Meet Kafka

1.6 What actually breaks in production (Ch. 1 foundations)

Chapter 1 is conceptual, but it plants several landmines that detonate later.

Chapter 1 is conceptual, but it plants several landmines that detonate later. Naming them now makes the rest of the book read as consequences rather than trivia.

6.1 Assuming topic-wide ordering

Symptom: events processed out of order; state machines end up in impossible states; "the update arrived before the create." Cause: ordering holds per partition, not per topic. Default producer behavior spreads messages evenly across all partitions. Fix: key by the entity whose ordering you need (user_id, account_id, trace_id) so all its events land in one partition.

6.2 Changing partition count on a keyed topic

Symptom: after adding partitions, a key's messages start landing in a different partition than its history. Ordering silently breaks; stateful consumers see split history; compacted-topic semantics get weird. Cause: hash(key) mod N — change N, change the mapping for most keys. The guarantee is explicitly conditional: same partition "provided that the partition count does not change." Also: partition count can only be increased, never decreased. Fix: size partitions for expected future throughput, not current. This is why Ch. 2 says to calculate based on future usage when keying.

6.3 Assuming offsets are contiguous

Symptom: gap-detection logic false-alarms; expected == last + 1 assertions fail. Cause: offsets increase but "not necessarily monotonically greater" — gaps are legal (compaction, transaction markers, aborted transactions). Fix: treat offsets as opaque, increasing cursors. Never arithmetic.

6.4 Schema-less topics → coupled deployments

Symptom: you cannot add a field without a coordinated multi-team deploy in a strict order; consumers crash on unknown fields; a producer rollback breaks consumers. Cause: no schema contract; writing and reading are tightly coupled. Fix: schema registry + a format with real compatibility rules (Avro). This is the lesson from LinkedIn's XML tracking system breaking "constantly due to changing schemas."

6.5 Retention shorter than your recovery time

Symptom: a consumer is down for maintenance longer than retention; on restart, its committed offset no longer exists; it either jumps to latest (silent data loss) or to earliest (re-processes everything). Cause: retention defines a minimum window, and it is a deletion policy, not an archive. Fix: retention ≥ worst-case consumer outage + recovery, with margin. Alert on consumer lag approaching the retention edge.

6.6 Treating MirrorMaker as intra-cluster replication

Symptom: people expect cross-DC failover to be as seamless as broker failover. It isn't — offsets are not identical across clusters, and MirrorMaker is asynchronous. Cause: Kafka's replication is explicitly intra-cluster only. MirrorMaker is a consumer+producer pair, i.e. an application with its own lag, its own failure modes, and its own offsets. Fix: design cross-cluster topologies deliberately (Ch. 10), and never assume offset equivalence between clusters.

6.7 The ActiveMQ lesson — coupling telemetry to serving

Symptom (historical, and still a live risk): the broker pauses → client connections back up → the application can no longer serve user requests. Cause: a messaging system that applies backpressure into a synchronous serving path. Fix / design rule: producers must be async and bounded; a telemetry path must be allowed to drop rather than block the request thread. Kafka's push-pull split exists precisely so slow consumers can't stall producers — but you can still recreate the bug yourself with a synchronous, unbounded, blocking producer call in a request handler.

6.8 Colocating Kafka with other memory-hungry apps

Foreshadowed here, explicit in Ch. 2: Kafka's read performance comes from the OS page cache. Anything else on the box competing for page cache degrades consumer performance. (Details in Ch. 2.)


On this page