Learn Labs
12. Administering Kafka

12.10 What actually breaks in production — Ch. 12 consolidated

Production failure catalog
0 rows
#SymptomRoot causeFix
1Anyone with shell access changed prod topics; no audit trail"Default configurations DO NOT RESTRICT the use of these tools"Restrict tool access to administrators; ACLs alone don't cover ZK-writing tools
2A CLI tool corrupted cluster stateTool version ≠ broker version; some tools write to ZooKeeperRun tools on the brokers themselves, using the deployed version
3Two topics merged in dashboardsPeriods in topic names become underscores in metricsNever use . in topic names
4Confusion with Kafka internalsTopic name starts with __Reserved by convention for internal topics
5Automation silently skipped a config change--if-exists on --alter masked a missing topic"Using it is NOT RECOMMENDED" — it hides real problems
6Producers failing with NotEnoughReplicas; nobody noticed the buildupNo monitoring on the ISR ladderAlert on --at-min-isr-partitions (before it becomes under-min-ISR)
7A partition is completely offlineNo leader available--unavailable-partitions; possibly unclean election (Ch. 7 §3.2)
8URP alerts fire constantly during deploysURPs are expected during maintenance/rebalance"Isn't necessarily bad" — alert on duration/trend, not presence
9Keyed consumers break after adding partitionshash(key) % N changed"Set the number of partitions ONCE... and avoid resizing"
10Need fewer partitions; can'tReducing partitions is impossibleDelete + re-create, or create topic-v2 and migrate producers
11--delete appeared to do nothingdelete.topic.enable=false — the request is ignoredEnable it (accepting the risk), or delete manually with full downtime
12Controller overwhelmed; cluster sluggish after a cleanup scriptDeleted many topics at once"NOT MORE THAN ONE OR TWO TOPICS AT A TIME"; consider controller_mutations_rate
13Deleted the wrong topic"NOT A REVERSIBLE OPERATION" — and no success/failure outputdelete.topic.enable=false as a guardrail; verify with --list/--describe
14Cluster slow with thousands of unused topicsEven empty topics consume disk, filehandles, memory, and controller metadataDelete unused topics (carefully, per #12)
15--delete --group failed: "The group is not empty"Group has active membersShut down all consumers first
16Offsets reset accidentallyRan the export command without --dry-runThe export/destroy commands differ by one flag — script it carefully
17Imported offsets had no effectConsumers were running and overwrote them"ALL consumers in the group are STOPPED" first
18A client's quota is 5× smaller than expectedQuotas are per-broker; all leadership landed on one brokerKeep leadership balanced (preferred leader election / Cruise Control)
19Two unrelated consumer groups share a quotaSame client.id across groups"Best practice to set the client ID for each consumer group to something unique"
20Automation misread a topic's effective config--describe shows only overrides, never cluster defaultsKeep separate knowledge of defaults, or use AdminClient's describeConfigs (Ch. 5), which reports isDefault()
21Leadership badly skewed after a rolling restartOriginal leaders do not automatically resume leadershipauto.leader.rebalance.enable, or run preferred leader election, or Cruise Control
22A wrapper script around the console consumer lost messages"Difficult to interact with the console consumer in a way that does not lose messages"Use the real client libraries
23Console producer messages all have null keysNo tab separator, or parse.key=falseSet key.separator and parse.key via --property (not --producer-property)
24A config passed to the client had no effectUsed --property (formatter) instead of --producer-property/--consumer-property (client)Know which is which
25Cluster performance tanked during a reassignmentReassignment copies whole partitions; disrupts page cache + network + disk I/O--throttle, combinable with --additional to throttle an in-flight move
26Reassignment off a decommissioning broker is glacially slowThat broker is the leader for everything being copied — one NIC serving all readsMove leadership off first (Cruise Control demotion, or just bounce the broker, temporarily disabling auto-rebalance)
27Can't verify reassignment progressLost the JSON file used in --executeKeep both generated files (revert- and expand-)
28A reassignment proposal is impossible to satisfyRack-awareness constraints--disable-rack-aware (understanding the durability cost)
29RF didn't increase after changing the cluster default"Existing topics will NOT automatically be increased"Reassign with an extra broker ID in each replica set
30A cancelled reassignment left the cluster worse off--cancel reverts to the prior replica set — which may include a dead/overloaded broker; and replica order isn't guaranteed (so preferred leaders may change)Understand before cancelling; re-run preferred leader election afterward
31A poison-pill message breaks a consumer and you can't see it—kafka-dump-log.sh --print-data-log on the right segment
32Consumption errors that look like corruptionCorrupt index file--index-sanity-check / --verify-index-only; indexes regenerate
33Replicas silently diverged (gaps in a follower)"previously replicated log segments can get deleted from a broker, and the follower WILL NOT FILL IN THE GAPS"kafka-replica-verification.sh — but see #34
34Running replica verification destroyed cluster performanceIt reads all messages from the oldest offset, from all replicas, in parallel, in a loopUse sparingly, off-peak, scoped by topic regex
35Controller alive but non-functionalController thread hit an exceptionDelete /admin/controller znode → forces resignation and re-election (cannot choose the successor)
36A topic is stuck "marked for deletion" foreverDeletion requested with deletion disabled, or replicas went offline mid-deleteDelete /admin/delete_topic/<topic> (not the parent), then force a controller move to clear cached requests
37Cluster unstable after editing ZooKeeperModified topic metadata while brokers were online"NEVER attempt to delete or modify topic metadata in ZooKeeper while the cluster is online" — manual deletion requires full shutdown