Learn Labs
12. Administering Kafka

12.9 ⚠️ Unsafe operations — "Here be dragons"

"There are some administrative tasks that are technically possible to do but SHOULD NOT BE ATTEMPTED EXCEPT IN THE MOST EXTREME SITUATIONS. Often this is when you are diagnosing a problem and have RUN OUT OF OPTIONS, or you have found a specific bug that you need to work around temporarily. These tasks are usually UNDOCUMENTED, UNSUPPORTED, and POSE SOME AMOUNT OF RISK."

⚠️ DANGER: HERE BE DRAGONS

"The operations in this section often involve working with the cluster metadata stored in ZooKeeper DIRECTLY. This can be a VERY DANGEROUS operation, so you must be very careful to NOT MODIFY the information in ZooKeeper directly, EXCEPT AS NOTED."

9.1 Moving the cluster controller

When: "when troubleshooting a misbehaving cluster or broker, it may be useful to forcibly move the controller to a different broker WITHOUT SHUTTING DOWN THE HOST. One such example is when the controller has suffered AN EXCEPTION or other problem that has left it RUNNING BUT NOT FUNCTIONAL."

How: "deleting the ZooKeeper znode at /admin/controller manually will cause the current controller to RESIGN, and the cluster will randomly select a new controller."

⚠️ "There is currently NO WAY to specify a SPECIFIC broker to be controller in Apache Kafka." "Moving the controller in these situations does not normally have a high risk, but as it is not a normal task, it should not be performed regularly."

(Ch. 6 §2.2: the resigning controller's epoch is superseded, so its stale messages get fenced. That's why this is relatively safe.)

9.2 Unsticking topic deletion

Two scenarios where deletion gets stuck:

  1. “A requester has No way of knowing whether topic deletion is enabled in the cluster and can request deletion of a topic from a cluster in which Deletion is disabled.”
  2. “A Very large topic is requested to be deleted, but before the request is handled, One or more of the replica sets goes offline due to hardware failures, and the deletion cannot complete As the controller cannot ack that the deletion was completed successfully.”

The fix:

  1. Delete the znode /admin/delete_topic/<topic>
    ⚠ “Deleting the topic ZooKeeper nodes (BUT NOT THE PARENT /admin/delete_topic NODE) will remove the pending requests.”
  2. “If the deletion is Re-queued by cached requests in the controller, it may be necessary to Also forcibly move the controller (as shown above) immediately after removing the topic znode to ensure that no cached requests are pending in the controller.”

9.3 Deleting topics manually

When: "If you are running a cluster with delete topics disabled, or if you find yourself needing to delete some topics outside of the normal flow."

⚠️ SHUT DOWN BROKERS FIRST

"Modifying the cluster metadata in ZooKeeper when the cluster is ONLINE is a VERY DANGEROUS operation and can put the cluster into an UNSTABLE STATE. NEVER attempt to delete or modify topic metadata in ZooKeeper WHILE THE CLUSTER IS ONLINE."

  1. Shut down all brokers in the cluster.
  2. Remove the ZooKeeper path /brokers/topics/<topic>
    ⚠ “Note that this node Has child nodes that must be deleted first.”
  3. Remove the Partition directories from the log directories on Each broker. “These will be named <topic>-<int>, where <int> is the partition ID.”
  4. Restart all brokers.

(Requires full cluster downtime. This is the definition of a last resort.)


On this page