12.5 Partition management
Also supports --topic + --partition directly.
Leader distribution
- 3
- balanced
- +1
- skew
Skewed by +1 partitions per surviving broker. When a broker comes back its partitions rejoin the ISR as followers but do not reclaim leadership, so the survivors keep serving all the traffic. The cluster looks healthy — every partition has a leader — while a subset of brokers carries the load.
5.1 Preferred replica election
The problem it solves:
"Leadership is defined within Kafka as THE FIRST IN-SYNC REPLICA IN THE REPLICA LIST. However, when a broker is stopped or loses connectivity, leadership is transferred to another in-sync replica, and THE ORIGINAL DOES NOT RESUME LEADERSHIP OF ANY PARTITIONS AUTOMATICALLY. THIS CAN CAUSE WILDLY INEFFICIENT BALANCE AFTER A DEPLOYMENT ACROSS A FULL CLUSTER if automatic leader balancing is not enabled."
► END STATE: leadership piled unevenly. Every deploy makes it worse. A restarted broker “DOES NOT RESUME LEADERSHIP OF ANY PARTITIONS AUTOMATICALLY” — so the drift only ever accumulates unless automatic leader balancing is enabled.
"it is recommended to ensure that this setting is enabled or to use other open source tooling such as Cruise Control to ensure that a good balance is maintained at all times."
The fix — a cheap, safe operation:
"a lightweight, GENERALLY NON-IMPACTING procedure called preferred leader election. This tells the cluster controller to select the ideal leader for partitions. Clients can track leadership changes automatically, so they will be able to move to the new broker."
# all topics
kafka-leader-election.sh --bootstrap-server localhost:9092 \
--election-type PREFERRED --all-topic-partitions
# specific partitions from a JSON file
kafka-leader-election.sh --bootstrap-server localhost:9092 \
--election-type PREFERRED --path-to-json-file partitions.json{ "partitions": [
{ "partition": 1, "topic": "my-topic" },
{ "partition": 2, "topic": "foo" }
] }Also supports --topic + --partition directly.
"An older version of this tool called
kafka-preferred-replica-election.shis also available but has been DEPRECATED in favor of the new tool, which allows for more customization, such as specifying whether we want a 'PREFERRED' or 'UNCLEAN' election type."
5.2 Reassigning replicas — kafka-reassign-partitions.sh
Four reasons you'd do it:
- “uneven load on brokers that the automatic leader distribution is not correctly handling”
- “a broker is taken offline and the partition is under replicated”
- “a new broker is added and we want to more quickly balance new partitions”
- “You want to ADJUST THE REPLICATION FACTOR of a topic”
The three-step process:
Step 1 — generate a proposal
Scenario: "a four-broker cluster. You've recently added two new brokers, bringing the total up to six, and you want to move two of your topics onto brokers 5 and 6."
// topics.json
{ "topics": [ { "topic": "foo1" }, { "topic": "foo2" } ], "version": 1 }kafka-reassign-partitions.sh --bootstrap-server localhost:9092 \
--topics-to-move-json-file topics.json --broker-list 5,6 --generateOutput gives two JSON blocks:
Save both. The current assignment “can be used to MOVE PARTITIONS BACK to where they were originally if you need to ROLL BACK for some reason”; the proposed one is the file you pass to --execute, and --verify needs that same file again.
⚠️ "You'll notice in the output that there isn't a good balance of leadership, as the proposal will result in ALL LEADERSHIP MOVING TO BROKER 5. We will ignore this for now and presume the cluster automatic leadership balancing is enabled, which will help distribute it later."
💡 "the first step can be SKIPPED if you know exactly where you want to move your partitions to and you manually craft the JSON."
Step 2 — execute
kafka-reassign-partitions.sh --bootstrap-server localhost:9092 \
--reassignment-json-file expand-cluster-reassignment.json --execute"Save this to use as the
--reassignment-json-fileoption during rollback" — the tool reminds you.
What actually happens under the hood:
- Controller Adds the new replicas to the replica list for each partition
► “which will Temporarily increase the replication factor of these topics” - New replicas Copy all existing messages for each partition from the current leader
⚠ “Depending on the size of the partitions on disk, This can take a significant amount of time as the data is copied across the network” - Controller Removes the old replicas by reducing the RF back to the original size
Three useful flags:
| Flag | Purpose |
|---|---|
--additional | "allows you to add to the EXISTING reassignments so they can continue without interruption and WITHOUT the need to wait until the original movements have completed in order to start a new batch" |
--disable-rack-aware | "There may be times when, due to rack awareness settings, the end-state of a proposal may not be POSSIBLE. This can be overridden" |
--throttle | bytes/sec. "Partition reassignments have a big impact on the performance of your cluster, as they will cause changes in the consistency of the MEMORY PAGE CACHE and use network and disk I/O. ... Can be combined with --additional to throttle an ALREADY-STARTED reassignment process that may be causing issues." |
That last point is the emergency brake: --throttle + --additional lets you slow down a reassignment you already started and which is hurting the cluster.
(And note the page-cache remark — reassignment reads cold data from the beginning of partitions, evicting the hot pages that serve your real consumers. Ch. 6 §6.2's tiered-storage isolation argument, again.)
💡 IMPROVING NETWORK UTILIZATION WHEN REASSIGNING REPLICAS
*"When removing many partitions from a single broker — such as if that broker is being removed from the cluster — it may be useful to REMOVE ALL LEADERSHIP FROM THE BROKER FIRST.
Doing that manually "is arduous." Options:
- Cruise Control includes broker "demotion," which "safely moves leadership off a broker and is probably the simplest way to do this."
- Without such tools: A SIMPLE RESTART OF A BROKER WILL SUFFICE. "As a broker is preparing to shut down, all leadership for its partitions will move to other brokers. This can SIGNIFICANTLY INCREASE THE PERFORMANCE of reassignments and REDUCE THE IMPACT on the cluster, as the replication traffic will be DISTRIBUTED TO MANY BROKERS."
- ⚠️ "However, if automatic leader reassignment is enabled after the broker is bounced, LEADERSHIP MAY RETURN to this broker, so it may be beneficial to TEMPORARILY DISABLE this feature."
The copy is a replication read, and a replication read is served by the leader.
If broker 4 is the LEADER of everything being moved off it, then broker 4 alone must serve all the replication reads. Bouncing it first (or demoting it with Cruise Control) moves that read load onto the other five machines.
Step 3 — verify
kafka-reassign-partitions.sh --bootstrap-server localhost:9092 \
--reassignment-json-file expand-cluster-reassignment.json --verify
Reassignment of partition [foo1,0] completed successfully
Reassignment of partition [foo1,1] is in progress
..."This will show which reassignments are currently in progress, which have completed, and (if there was an error) which have failed. To do this, you MUST HAVE THE FILE with the JSON object that was used in the execute step."
5.3 Changing the replication factor
Why: "a partition was created with the wrong RF, you want increased redundancy as you expand your cluster, or you want to decrease redundancy for cost savings."
💡 "One clear example is that if a cluster RF DEFAULT setting is adjusted, EXISTING TOPICS WILL NOT AUTOMATICALLY BE INCREASED."
Method: craft the JSON with an extra broker ID in the replica set. RF 2 → RF 3 by adding broker 4 to [5,6]:
{ "version":1,
"partitions":[{"topic":"foo1","partition":1,"replicas":[5,6,4]},
{"topic":"foo1","partition":2,"replicas":[5,6,4]},
{"topic":"foo1","partition":3,"replicas":[5,6,4]}] }Then --execute and confirm via --verify or kafka-topics.sh --describe:
Topic:foo1 PartitionCount:3 ReplicationFactor:3 Configs:
Topic: foo1 Partition: 0 Leader: 5 Replicas: 5,6,4 Isr: 5,6,45.4 Canceling reassignments
"Canceling a replica reassignment in the past was a DANGEROUS process that required unsafe manual manipulation of ZooKeeper nodes (deleting the
/admin/reassign_partitionsznode). Fortunately, this is NO LONGER THE CASE."
--cancel "will cancel the active reassignments that are ongoing in a cluster... designed to RESTORE THE REPLICA SET TO THE ONE IT WAS PRIOR TO reassignment being initiated."
⚠ TWO CAVEATS on cancelling a reassignment:
- “if replicas are being removed from a DEAD BROKER or an OVERLOADED BROKER, IT MAY LEAVE THE CLUSTER IN AN UNDESIRABLE STATE.” You get reverted to a replica set that included a broken broker.
- “There is also NO GUARANTEE THAT THE REVERTED REPLICA SET WILL BE IN THE SAME ORDER as it was previously.” ► i.e. THE PREFERRED LEADER MAY CHANGE. (Ch. 6 §4.4)