Learn Labs
5. Managing Apache Kafka Programmatically (AdminClient)

5.10 Deploy / monitor / scale / backup — AdminClient's operational role

This chapter is, quietly, where lag monitoring comes from:

AdminClient as the monitoring primitive

This chapter is, quietly, where lag monitoring comes from:

What you're monitoringCallsWhat they give you
Consumer lag, computed correctlylistConsumerGroupOffsets(group)
listOffsets(OffsetSpec.latest())
Committed offset per partition; log end offset per partition — lag = latest - committed.
Group healthdescribeConsumerGroups(…)Members, hosts, assignments, assignment algorithm, coordinator host.
Cluster healthdescribeCluster()clusterId, nodes, controller — controller identity and count is a top-tier alert (see Ch. 2/13).

The important point: this is the supported way. Parsing __consumer_offsets is unsupported and breaks on upgrade (#15).

AdminClient as the recovery toolkit

IncidentAdminClient response
Broker config lostdescribeConfigs on a surviving broker
App must reprocess (bad output, DR failover)alterConsumerGroupOffsets — stop the group first; reset the state store too
Leadership imbalancedelectLeaders(PREFERRED)
Partition leaderlesselectLeaders(UNCLEAN) ← accepts data loss
Broker overloaded / decommissioningalterPartitionReassignments (+ throttle via quotas)
Reassignment gone wrongOptional.empty() to cancel it
GDPR deletion requestlistOffsets(forTimestamp) + deleteRecords(beforeOffset)
Topic misconfiguredincrementalAlterConfigs (SET / DELETE / APPEND / SUBTRACT)
Throughput ceiling hitcreatePartitions (⚠ breaks keyed apps)

"Backup" — what this chapter contributes

Kafka still has no backup command, but Ch. 5 adds two genuinely useful things:

  1. Configuration is recoverable from any running broker. describeConfigs turns every live broker into a config backup. Consider dumping it to version control on a schedule — cheap insurance, per the war story.
  2. Deletion is not recoverable. deleteTopics is final; deleteRecords is one-way. The only protections are delete.topic.enable=false and your own tooling discipline.

And one anti-pattern to retire: do not back up or restore Kafka state by touching ZooKeeper. That path is being removed.

Scaling AdminClient usage itself

Your situationHow to call it
Admin operations are rareBlocking get() is fine
You're building an admin servicewhenComplete() + per-call timeoutMs
Admin ops on the critical pathLow timeout + a degraded fallback
Many resources to inspectBatch them — describeConfigs accepts multiple resources of multiple types
Lists you mutate concurrentlyAPPEND / SUBTRACT, not SET

On this page