11. Securing Kafka
11.10 What actually breaks in production — Ch. 11 consolidated
| # | Symptom | Root cause | Fix |
|---|---|---|---|
| 1 | Anyone can read/write; no identity in logs | PLAINTEXT listener → principal is User:ANONYMOUS | SSL or SASL_SSL listener |
| 2 | Clients on an SSL listener show as User:ANONYMOUS | ssl.client.auth=requested (not required) and the client has no key store | Use required if you need identity |
| 3 | Man-in-the-middle possible | Hostname verification disabled to "fix" a cert problem | "Should NOT be disabled in production." Fix the SAN/CN, or use client.dns.lookup (Ch. 5) |
| 4 | TLS handshake failures appear overnight | Certificates expired | Rotate before expiry; broker stores are dynamically updatable via configs tool / Admin API |
| 5 | Inter-broker TLS fails after enabling client auth | Broker trust store lacks the client CA (or the broker CA) | "Broker trust stores should include the CA of the broker certificates AS WELL AS the CA of the client certificates" |
| 6 | Throughput drops 20–30% after enabling SSL | Zero-copy is not supported for SSL | Expected. Consider where encryption is truly needed (Ch. 10 §7.2) |
| 7 | Broker network threads saturated; clients can't connect | TLS handshake DoS — handshakes run on network threads | Connection quotas/limits + connection.failed.authentication.delay.ms |
| 8 | Private keys readable by other users | Key stores are plain files on disk | Filesystem permissions on all key/trust stores and keytabs |
| 9 | Kerberos auth fails intermittently | Forward/reverse DNS mismatch | rdns=false in client krb5.conf; secure DNS is a requirement |
| 10 | All clients fail to authenticate at once | KDC or DNS outage (or DoS) | "It is NECESSARY to monitor the availability of these services" |
| 11 | Kerberos replay detection breaks / auth fails after clock drift | Clock skew beyond configured variability | Secure, monitored NTP — clock sync is part of the security perimeter |
| 12 | Adding one user requires restarting every broker | SASL/PLAIN's built-in JAAS password store | Custom server callback handler → external password server |
| 13 | Passwords visible in logs | Config not declared as PASSWORD type | Use ConfigDef PASSWORD type; externalize/encrypt |
| 14 | Password rotation causes an outage | Server accepts only one password at a time | Callback that accepts old and new for an overlap window + reauthentication |
| 15 | Credentials stolen off the wire | SASL/PLAIN or SASL/SCRAM over SASL_PLAINTEXT | Always SASL_SSL — PLAIN sends clear text; SCRAM exposes hashed keys during handshake |
| 16 | SCRAM credentials stolen from ZooKeeper | ZooKeeper not SSL-enabled / disk not encrypted | Both are stated requirements for production SCRAM |
| 17 | A deleted user keeps working | Existing connections survive user deletion | connections.max.reauth.ms; Deny ACL for immediate effect |
| 18 | Compromised user still active after removal | No reauth interval; SSL renegotiation is not supported so SSL connections never re-verify | Deny ACL — "the quickest way to disable access" — plus reauth config |
| 19 | Compromised super user can't be revoked quickly | super.users cannot be denied and requires a broker restart to change | Don't use super.users in production; grant explicit ACLs |
| 20 | super.users list parsed wrongly | Used commas; DNs contain commas | Semicolon-separated |
| 21 | Adding an ACL unexpectedly revoked others' access | allow.everyone.if.no.acl.found=true, and a new prefix/wildcard ACL made no.acl.found false | Don't use it in production |
| 22 | New topics silently world-accessible | Same config | Same fix |
| 23 | OAuth "works" in staging but is insecure | Built-in OAUTHBEARER uses unsecured JWTs and does not validate tokens | Custom login + server validator callbacks against a real OAuth server |
| 24 | Connections outlive their OAuth tokens | No reauthentication | connections.max.reauth.ms + token revocation |
| 25 | A delegation token was used to mint more tokens | It can't be | "Clients authenticated using delegation tokens CANNOT create other delegation tokens" |
| 26 | All delegation tokens broke | Master key rotated — requires restarting all brokers and deleting existing tokens | Plan the rotation: delete tokens → update key on all brokers → restart → recreate |
| 27 | Idempotent producer fails authorization | Missing Cluster:IdempotentWrite (non-transactional only) | Grant it |
| 28 | Transactional producer fails authorization | Missing TransactionalId:Write and/or Group:Read | Grant both (Ch. 8) |
| 29 | Consumer can fetch but not join a group | Has Topic:Read but not Group:Read | Grant Group:Read |
| 30 | A client was granted unintended broker powers | Cluster:ClusterAction granted to a non-broker | "Should ONLY be granted to brokers" |
| 31 | Unmanageable ACL sprawl | Per-resource literal ACLs at scale | Prefixed ACLs by department + group/role principals via a custom authorizer |
| 32 | Departing employee's credentials still power a service | Application used a personal principal | "Long-running applications can be configured with SERVICE credentials" |
| 33 | A reused principal name inherits old access | Principal reuse | "Reuse of principals must be AVOIDED" |
| 34 | No record of who accessed what | Grants log at DEBUG, only denials at INFO | Enable DEBUG on kafka.authorizer.logger if you need a full trail |
| 35 | Sensitive data found in a broker heap dump | TLS + disk encryption don't cover broker memory | End-to-end encryption (serializer/deserializer + KMS) |
| 36 | Cloud provider / platform admin could read customer data | Broker sees plaintext | End-to-end encryption — "brokers never see the unencrypted contents" |
| 37 | Compression gives no benefit and adds CPU | Compressing after encryption (high-entropy data) | Compress before encrypting; disable Kafka compression |
| 38 | Partitioning and compaction break after encrypting keys | Encrypted keys aren't hash-stable | Message key = secure hash of the original; encrypted key in header/payload via a producer interceptor |
| 39 | Key rotation needs a maintenance window | Compacted topics retain old-key messages indefinitely; re-encryption requires producers and consumers offline | Plan it; keep old keys available for the retention period |
| 40 | ZooKeeper ACLs written by one broker exclude others | Full Kerberos principals differ per broker | kerberos.removeHostFromPrincipal=true + kerberos.removeRealmFromPrincipal=true |
| 41 | Unexpected ZooKeeper access granted | ZK with SASL and SSL client auth associates multiple principals; any may grant | Understand the model; audit ZK ACLs |
| 42 | Anyone can read Kafka metadata from ZooKeeper | zookeeper.set.acl not enabled | Enable it — metadata becomes broker-writable only; sensitive paths (SCRAM) are not world-readable |
| 43 | DIGEST-MD5 used in production | It has "known security vulnerabilities" | Kerberos or TLS mutual auth |
| 44 | Flag-day protocol migration caused an outage | Changed the listener protocol in place | Add a new listener on a new port, migrate clients, then remove the old |