Learn Labs
5. Managing Apache Kafka Programmatically (AdminClient)

5.9 What actually breaks in production — Ch. 5 consolidated

Production failure catalog
0 rows
#SymptomRoot causeFix
1createTopics().get() succeeds, then listTopics() doesn't show itEventual consistency — the Future completes when the controller is updated; the read went to a broker that hasn't got the metadata yetRetry with backoff; never assert immediately after a mutation
2deleteTopics().get() succeeds, topic still describedSame — "due to the async nature of deletes, it is possible that at this point the topic still exists"Same
3Catch block never matches; you can't identify the errorYou checked the ExecutionException type instead of e.getCause()Always inspect getCause() — all AdminClient results wrap errors this way
4TopicExistsException on startup, intermittentlyRace: two app instances both describe (not found), both createCatch it and treat as success — then describe to validate config
5Irrecoverable data loss from a wrong topic name"deletion of topics is FINAL — no recycle bin, no checks that the topic is empty"Broker-side delete.topic.enable=false; confirmation + audit in any tooling
6HTTP/API server thread pool exhausted while Kafka is slowBlocking .get() inside a request handlerKafkaFuture.whenComplete() + per-call Options().timeoutMs()
7Service fails to start because Kafka is slowStartup topic validation with the 120 s default request.timeout.ms, treated as fatalLower the timeout and start anyway, validating later or skipping
8SASL authentication fails via a DNS aliasClient authenticates the alias; server principal is the real hostname → SASL treats the mismatch as a possible MITMclient.dns.lookup=resolve_canonical_bootstrap_servers_only
9Client can't connect even though brokers are healthy (K8s/LB)Client tries only the first resolved IP; that LB IP is downclient.dns.lookup=use_all_dns_ips
10Consumer group offsets won't update; UnknownMemberIdExceptionThe group is still active — Kafka blocks offset edits to live groups because consumers would overwrite themShut the consuming application down first (no admin command exists for this)
11Stateful stream app double-counts after an offset resetOffsets reset, but the state store still holds the old aggregateReset both; in dev, delete the state store entirely first
12Deleting offsets produced unpredictable behaviorPost-delete behavior is decided by the consumer's auto.offset.reset, which the operator may not knowSet offsets explicitly (e.g. to earliest) instead of deleting
13listConsumerGroups throws on one bad group and returns nothingUsed .all() — throws on the first errorUse .valid() (+ .errors() to inspect), for tooling that must degrade
14Group listing empty / describe failsAuthorization, or the group's coordinator is unavailableCheck ACLs and coordinator health
15Lag monitoring broke after a Kafka upgradeCustom code parsed __consumer_offsets internal messages — "Kafka does not guarantee compatibility of the internal message formats"listConsumerGroupOffsets + listOffsets
16createPartitions created the wrong number of partitionsThe argument is the TOTAL after expansion, not the count to adddescribe first, then increaseTo(current + n)
17Multi-topic expansion partially applied"some of the topics will be successfully expanded, while others will fail"Make tooling idempotent and re-runnable; check each result
18Keyed consumers break after adding partitionshash(key) % N changes (Ch. 3 §9.4)"check that the operation will not break any application that consumes from the topic"
19Regulator finds 90-day-old data on a 30-day-retention topic"retention policies were not implemented in a way that guarantees legal compliance" — a single unclosed segment retains everythinglistOffsets(forTimestamp) + deleteRecords(beforeOffset); also fix segment rolling (Ch. 2)
20Records "deleted" but disk usage unchanged"Full cleanup from disk will happen asynchronously" — deletion first only makes records inaccessibleExpected; don't gate disk-space alarms on it
21Cluster network saturated; replication falls behind after a reassignmentReplica reassignment copies large amounts of data with no throttleThrottle replication using quotas (broker config, editable via AdminClient)
22Leadership didn't move after a reassignmentThe first element of the replica list is the preferred leader; you kept the old broker firstOrder the list intentionally; then run preferred leader election
23ElectionNotNeededExceptionCluster healthy; the preferred leader already is the leaderNot an error — handle it as a no-op
24Partition permanently unavailable, no eligible leaderLeader down; all other replicas are missing data so are ineligibleEither wait for the old leader, or accept unclean leader election and permanent silent data loss
25Reassignment/election results look inconsistent right after the callAsync metadata propagationPoll listPartitionReassignments() / re-describe over time
26Broker config file destroyed during an upgrade, no backupNo config backup processA surviving broker IS your backup: describeConfigs against it (the book's war story)
27A topic silently stopped being compacted and data aged outConfig drift on a topic your app depends onPeriodically validate topic config from the app — "more frequently than the default retention period, just to be safe"
28UnsupportedOperationException: Not implemented yet in unit testsMockAdminClient doesn't mock everything (e.g. incrementalAlterConfigs ≤ 2.5)Inject your own implementation (Mockito doReturn)
29MockAdminClient not found on the test classpathIt ships in a test jarAdd <classifier>test</classifier> to the dependency
30Admin tooling broke on a ZooKeeper-less (KRaft) clusterCode manipulated ZooKeeper directly"NEVER use ZooKeeper directly" — AdminClient's API survives the migration
31All mutations fail while all reads succeedWrites go to the controller; reads go to any (least-loaded) broker → a sick controller shows exactly this asymmetryCheck controller health (describeCluster().controller())
32Ran a destructive tool against the wrong clusterNo cluster identity checkCompare cluster.clusterId() (a GUID) before destructive operations