12.8 Decision cheat sheet
Yes if consumers must be able to bootstrap from the log without a separate snapshot — which is the whole point of "rebuild a derived system from scratch." No if events are intent-…
Which broker style?
| If… | Then… |
|---|---|
| Messages are expensive to process, you want per-message parallelism, order is unimportant, and you want task-queue semantics | AMQP/JMS (RabbitMQ, SQS). Add a dead-letter queue from day one. |
| High throughput, each message is fast, order matters, and you want replay plus multiple independent consumers | Log-based (Kafka, Kinesis, Pulsar). Partition key = whatever must stay ordered, usually the entity ID. |
How do I keep N systems in sync? Never dual-write. Pick one system of record and make everything else a follower via CDC (existing mutable app) or event sourcing (new app, intent matters, auditability required). Use the outbox pattern if you don't want your internal schema to become a public contract.
Do I need log compaction? Yes if consumers must be able to bootstrap from the log without a separate snapshot — which is the whole point of "rebuild a derived system from scratch." No if events are intent-level (event sourcing), because later events don't supersede earlier ones.
Event time or processing time? Event time, essentially always, because it's the only choice that is deterministic under reprocessing. Use processing time only when the delay is negligibly short and you don't care about replay. Budget for watermarks, allowed lateness, and a dropped-event metric.
Which window? Fixed reporting intervals → tumbling. Smoothed trends → hopping. "Within N minutes of each other" → sliding (costly: buffers events). Per-user activity bursts → session.
Which fault-tolerance mechanism?
| Situation | Use |
|---|---|
| All effects stay inside the framework | Checkpointing / microbatching — free exactly-once |
| Writing to an external store that supports conditional writes | Idempotence with the offset as a dedup key — cheapest |
| Effects span the processor and one specific system that cooperates | Internal atomic commit (Kafka transactions, Dataflow) |
| Effects are truly external and irreversible (emails, payments) | Idempotency keys at the external boundary — no framework can help you |
Local state or remote? Local + periodic replication by default (remote lookup per message is slow). Remote only if state is enormous or shared. And know your restore time — that's your recovery objective.