Learn Labs
2. Installing Kafka

2.2 Environment setup

Kafka is a Java application and runs on Windows, macOS, Linux, and others.

2.1 Operating system

Kafka is a Java application and runs on Windows, macOS, Linux, and others. Linux is the recommendation for the general use case. (Windows/macOS install → Appendix A.)

2.2 Java

  • Needed before ZooKeeper or Kafka.
  • Works with all OpenJDK-based implementations, including Oracle JDK.
  • Latest Kafka versions support Java 8 and Java 11.
  • Runtime (JRE) is enough to run, but the full JDK is recommended when developing tools and applications.
  • Install the latest released patch version — older versions may carry security vulnerabilities.
  • Book's examples assume /usr/java/jdk-11.0.10.

2.3 ZooKeeper — what it's for and why it's shaped that way

What Kafka stores in ZooKeeper: metadata about the Kafka cluster (brokers, topics, partitions), and (historically) consumer client details.

ZooKeeper is "a centralized service for maintaining configuration information, naming, providing distributed synchronization, and providing group services."

Tested extensively with the stable 3.5 release; the book uses 3.5.9.

Note: you can run a ZooKeeper server from scripts inside the Kafka distribution, but "it is trivial to install a full version of ZooKeeper from the distribution" — do that instead for anything real.

Standalone server (dev only)
tar -zxf apache-zookeeper-3.5.9-bin.tar.gz
mv apache-zookeeper-3.5.9-bin /usr/local/zookeeper
mkdir -p /var/lib/zookeeper
cat > /usr/local/zookeeper/conf/zoo.cfg << EOF
tickTime=2000
dataDir=/var/lib/zookeeper
clientPort=2181
EOF
export JAVA_HOME=/usr/java/jdk-11.0.10
/usr/local/zookeeper/bin/zkServer.sh start

Verify with a four-letter command over the client port:

telnet localhost 2181
srvr
# Zookeeper version: 3.5.9-...
# Latency min/avg/max: 0/0/0
# Received: 1
# Sent: 0
# Connections: 1
# Outstanding: 0
# Zxid: 0x0
# Mode: standalone      ← this is what you're checking
# Node count: 5
Ensemble (production)

ZooKeeper is designed to run as a cluster — an ensemble — for high availability.

Why an odd number of servers? Because of the balancing/consensus algorithm, a majority of members (a quorum) must be working for ZooKeeper to respond to requests.

Ensemble sizeQuorumFailures tolerated
321
532
743

The count is odd because of the consensus algorithm: a majority of members (a quorum) must be working for ZooKeeper to respond. The argument for 5 over 3 is maintenance — you reload nodes one at a time, so if the ensemble “cannot tolerate more than one node being down, doing maintenance work introduces additional risk.” Do not run more than seven nodes; add observer nodes instead for read-only traffic.

Sizing guidance — and the reasoning matters more than the number:

Consider a five-node ensemble. To make configuration changes to the ensemble, including swapping a node, you reload nodes one at a time. If your ensemble cannot tolerate more than one node being down, doing maintenance work introduces additional risk.

That is the whole argument for 5 over 3: a 3-node ensemble tolerates 1 failure, and routine maintenance consumes that entire budget — so a hardware failure during a rolling restart takes you down.

Do not run more than seven nodes — performance degrades due to the nature of the consensus protocol. If 5–7 nodes can't support the load due to too many client connections, add observer nodes to balance read-only traffic (observers participate in reads but not in the quorum vote).

Ensemble configuration — shared config on all nodes:

tickTime=2000
dataDir=/var/lib/zookeeper
clientPort=2181
initLimit=20
syncLimit=5
server.1=zoo1.example.com:2888:3888
server.2=zoo2.example.com:2888:3888
server.3=zoo3.example.com:2888:3888
ParamMeaning
tickTimeThe base time unit (ms). Everything else is a multiple of this.
initLimitHow long to allow followers to connect with a leader. 20 × 2000ms = 40s
syncLimitHow long out-of-sync followers may lag the leader. 5 × 2000ms = 10s
server.X=host:peerPort:leaderPortX = server ID (integer; need not be zero-based or sequential), peerPort (2888) = inter-ensemble communication, leaderPort (3888) = leader election

Plus, per server: a file named myid in dataDir containing that server's ID number, matching the config.

Port reachability rules — a classic firewall bug:

WhoPorts that must be reachable
Clientsonly the clientPort (2181)
Ensemble membersall three ports to each other — 2181, 2888 (peer), 3888 (leader election)

Single-machine ensemble for testing: set all hostnames to localhost, give each instance unique peerPort/leaderPort, and a separate zoo.cfg per instance with unique dataDir and clientPort. Testing purposes only — not recommended for production. (It shares a single failure domain, which defeats the point.)


On this page