Learn Labs
1. Meet Kafka

1.5 Use cases (with the "why Kafka specifically" for each)

The win: avoids duplicating this logic in every app, and enables aggregation that would not otherwise be possible (you can't batch a user's notifications if each app sends its own…

Activity tracking — the original LinkedIn use case. Frontends generate messages about user actions: passive (page views, click tracking) or complex (profile updates). Published to topics; consumed by backends generating reports, feeding ML, updating search results.

Messaging (e.g. user notifications/emails) — the interesting part is why Kafka helps. Producing apps emit "notify this user" without knowing formatting or delivery. One application reads all of them and handles consistently:

  • Formatting/decorating with a common look and feel
  • Collecting multiple messages into a single notification
  • Applying the user's delivery preferences

The win: avoids duplicating this logic in every app, and enables aggregation that would not otherwise be possible (you can't batch a user's notifications if each app sends its own emails).

Metrics and logging — where "multiple applications producing the same type of message" shines. Apps publish metrics to a topic; consumed by monitoring/alerting and by an offline system like Hadoop for long-term analysis (growth projections). Logs route to Elasticsearch or security analysis. Key benefit: when the destination system changes (time to replace the log store), you do not alter the frontend applications or the aggregation mechanism. The producers are insulated from sink churn.

Commit log — publish DB changes to Kafka; apps monitor the stream for live updates. Uses:

  • Replicate DB updates to a remote system
  • Consolidate changes from multiple applications into a single database view
  • Durable retention buffers the changelog → replay after a consumer-side failure
  • Log-compacted topics give longer retention by keeping only one change per key

(This is CDC — see Ch. 9 for Connect-based implementations.)

Stream processing — applications giving map/reduce-like functionality but on data in real time, as quickly as messages are produced, versus Hadoop's hours/days aggregation windows. Tasks: counting metrics, repartitioning messages for efficient downstream processing, transforming messages using data from multiple sources.