12. Administering Kafka
12.10 What actually breaks in production — Ch. 12 consolidated
| # | Symptom | Root cause | Fix |
|---|---|---|---|
| 1 | Anyone with shell access changed prod topics; no audit trail | "Default configurations DO NOT RESTRICT the use of these tools" | Restrict tool access to administrators; ACLs alone don't cover ZK-writing tools |
| 2 | A CLI tool corrupted cluster state | Tool version ≠ broker version; some tools write to ZooKeeper | Run tools on the brokers themselves, using the deployed version |
| 3 | Two topics merged in dashboards | Periods in topic names become underscores in metrics | Never use . in topic names |
| 4 | Confusion with Kafka internals | Topic name starts with __ | Reserved by convention for internal topics |
| 5 | Automation silently skipped a config change | --if-exists on --alter masked a missing topic | "Using it is NOT RECOMMENDED" — it hides real problems |
| 6 | Producers failing with NotEnoughReplicas; nobody noticed the buildup | No monitoring on the ISR ladder | Alert on --at-min-isr-partitions (before it becomes under-min-ISR) |
| 7 | A partition is completely offline | No leader available | --unavailable-partitions; possibly unclean election (Ch. 7 §3.2) |
| 8 | URP alerts fire constantly during deploys | URPs are expected during maintenance/rebalance | "Isn't necessarily bad" — alert on duration/trend, not presence |
| 9 | Keyed consumers break after adding partitions | hash(key) % N changed | "Set the number of partitions ONCE... and avoid resizing" |
| 10 | Need fewer partitions; can't | Reducing partitions is impossible | Delete + re-create, or create topic-v2 and migrate producers |
| 11 | --delete appeared to do nothing | delete.topic.enable=false — the request is ignored | Enable it (accepting the risk), or delete manually with full downtime |
| 12 | Controller overwhelmed; cluster sluggish after a cleanup script | Deleted many topics at once | "NOT MORE THAN ONE OR TWO TOPICS AT A TIME"; consider controller_mutations_rate |
| 13 | Deleted the wrong topic | "NOT A REVERSIBLE OPERATION" — and no success/failure output | delete.topic.enable=false as a guardrail; verify with --list/--describe |
| 14 | Cluster slow with thousands of unused topics | Even empty topics consume disk, filehandles, memory, and controller metadata | Delete unused topics (carefully, per #12) |
| 15 | --delete --group failed: "The group is not empty" | Group has active members | Shut down all consumers first |
| 16 | Offsets reset accidentally | Ran the export command without --dry-run | The export/destroy commands differ by one flag — script it carefully |
| 17 | Imported offsets had no effect | Consumers were running and overwrote them | "ALL consumers in the group are STOPPED" first |
| 18 | A client's quota is 5× smaller than expected | Quotas are per-broker; all leadership landed on one broker | Keep leadership balanced (preferred leader election / Cruise Control) |
| 19 | Two unrelated consumer groups share a quota | Same client.id across groups | "Best practice to set the client ID for each consumer group to something unique" |
| 20 | Automation misread a topic's effective config | --describe shows only overrides, never cluster defaults | Keep separate knowledge of defaults, or use AdminClient's describeConfigs (Ch. 5), which reports isDefault() |
| 21 | Leadership badly skewed after a rolling restart | Original leaders do not automatically resume leadership | auto.leader.rebalance.enable, or run preferred leader election, or Cruise Control |
| 22 | A wrapper script around the console consumer lost messages | "Difficult to interact with the console consumer in a way that does not lose messages" | Use the real client libraries |
| 23 | Console producer messages all have null keys | No tab separator, or parse.key=false | Set key.separator and parse.key via --property (not --producer-property) |
| 24 | A config passed to the client had no effect | Used --property (formatter) instead of --producer-property/--consumer-property (client) | Know which is which |
| 25 | Cluster performance tanked during a reassignment | Reassignment copies whole partitions; disrupts page cache + network + disk I/O | --throttle, combinable with --additional to throttle an in-flight move |
| 26 | Reassignment off a decommissioning broker is glacially slow | That broker is the leader for everything being copied — one NIC serving all reads | Move leadership off first (Cruise Control demotion, or just bounce the broker, temporarily disabling auto-rebalance) |
| 27 | Can't verify reassignment progress | Lost the JSON file used in --execute | Keep both generated files (revert- and expand-) |
| 28 | A reassignment proposal is impossible to satisfy | Rack-awareness constraints | --disable-rack-aware (understanding the durability cost) |
| 29 | RF didn't increase after changing the cluster default | "Existing topics will NOT automatically be increased" | Reassign with an extra broker ID in each replica set |
| 30 | A cancelled reassignment left the cluster worse off | --cancel reverts to the prior replica set — which may include a dead/overloaded broker; and replica order isn't guaranteed (so preferred leaders may change) | Understand before cancelling; re-run preferred leader election afterward |
| 31 | A poison-pill message breaks a consumer and you can't see it | — | kafka-dump-log.sh --print-data-log on the right segment |
| 32 | Consumption errors that look like corruption | Corrupt index file | --index-sanity-check / --verify-index-only; indexes regenerate |
| 33 | Replicas silently diverged (gaps in a follower) | "previously replicated log segments can get deleted from a broker, and the follower WILL NOT FILL IN THE GAPS" | kafka-replica-verification.sh — but see #34 |
| 34 | Running replica verification destroyed cluster performance | It reads all messages from the oldest offset, from all replicas, in parallel, in a loop | Use sparingly, off-peak, scoped by topic regex |
| 35 | Controller alive but non-functional | Controller thread hit an exception | Delete /admin/controller znode → forces resignation and re-election (cannot choose the successor) |
| 36 | A topic is stuck "marked for deletion" forever | Deletion requested with deletion disabled, or replicas went offline mid-delete | Delete /admin/delete_topic/<topic> (not the parent), then force a controller move to clear cached requests |
| 37 | Cluster unstable after editing ZooKeeper | Modified topic metadata while brokers were online | "NEVER attempt to delete or modify topic metadata in ZooKeeper while the cluster is online" — manual deletion requires full shutdown |