5. Managing Apache Kafka Programmatically (AdminClient)
5.10 Deploy / monitor / scale / backup — AdminClient's operational role
This chapter is, quietly, where lag monitoring comes from:
AdminClient as the monitoring primitive
This chapter is, quietly, where lag monitoring comes from:
| What you're monitoring | Calls | What they give you |
|---|---|---|
| Consumer lag, computed correctly | listConsumerGroupOffsets(group)listOffsets(OffsetSpec.latest()) | Committed offset per partition; log end offset per partition — lag = latest - committed. |
| Group health | describeConsumerGroups(…) | Members, hosts, assignments, assignment algorithm, coordinator host. |
| Cluster health | describeCluster() | clusterId, nodes, controller — controller identity and count is a top-tier alert (see Ch. 2/13). |
The important point: this is the supported way. Parsing __consumer_offsets is unsupported and breaks on upgrade (#15).
AdminClient as the recovery toolkit
| Incident | AdminClient response |
|---|---|
| Broker config lost | describeConfigs on a surviving broker |
| App must reprocess (bad output, DR failover) | alterConsumerGroupOffsets — stop the group first; reset the state store too |
| Leadership imbalanced | electLeaders(PREFERRED) |
| Partition leaderless | electLeaders(UNCLEAN) ← accepts data loss |
| Broker overloaded / decommissioning | alterPartitionReassignments (+ throttle via quotas) |
| Reassignment gone wrong | Optional.empty() to cancel it |
| GDPR deletion request | listOffsets(forTimestamp) + deleteRecords(beforeOffset) |
| Topic misconfigured | incrementalAlterConfigs (SET / DELETE / APPEND / SUBTRACT) |
| Throughput ceiling hit | createPartitions (⚠ breaks keyed apps) |
"Backup" — what this chapter contributes
Kafka still has no backup command, but Ch. 5 adds two genuinely useful things:
- Configuration is recoverable from any running broker.
describeConfigsturns every live broker into a config backup. Consider dumping it to version control on a schedule — cheap insurance, per the war story. - Deletion is not recoverable.
deleteTopicsis final;deleteRecordsis one-way. The only protections aredelete.topic.enable=falseand your own tooling discipline.
And one anti-pattern to retire: do not back up or restore Kafka state by touching ZooKeeper. That path is being removed.
Scaling AdminClient usage itself
| Your situation | How to call it |
|---|---|
| Admin operations are rare | Blocking get() is fine |
| You're building an admin service | whenComplete() + per-call timeoutMs |
| Admin ops on the critical path | Low timeout + a degraded fallback |
| Many resources to inspect | Batch them — describeConfigs accepts multiple resources of multiple types |
| Lists you mutate concurrently | APPEND / SUBTRACT, not SET |