Learn Labs
1. Meet Kafka

1.9 Self-test — can you answer these without looking?

Self-test12 questions

—/12
  1. Why is "we built a pub/sub system" the inevitable outcome of a growing data platform, and what specific cost does it eliminate?

  2. LinkedIn had a monitoring system and a tracking system. Give three concrete reasons neither could absorb the other's workload.

  3. ActiveMQ was rejected for two reasons. One was scale. What was the other, and why is it the more important lesson?

  4. What exactly does a key guarantee, and under what condition does that guarantee evaporate?

  5. Why does batching require messages to share both topic and partition?

  6. Kafka treats payloads as opaque bytes. Explain the deploy-ordering problem that creates, and how a schema registry fixes it.

  7. State the ordering guarantee precisely. What's the cost of getting global ordering?

  8. Who must producers connect to? Who may consumers connect to? Why the asymmetry?

  9. What's the difference between a delete-retention topic and a compacted topic, and which use case needs which?

  10. Why is Kafka's replication not a cross-datacenter solution, and what fills that gap?

  11. Why is "how do I back up Kafka?" the wrong question — and what are the five things that actually cover the intent?

  12. Why are Connect and Streams deliberately libraries rather than a YARN-style platform?

    Next: Chapter 2 — Installing Kafka — hardware selection, broker config, OS tuning, and the production concerns that make or break a cluster.