Learn Labs
8. Exactly-Once Semantics

8.1 Why at-least-once isn't always enough

Ch. 7 delivered at-least-once — "the guarantee that Kafka will not lose messages that it acknowledged as committed.

Ch. 7 delivered at-least-once — "the guarantee that Kafka will not lose messages that it acknowledged as committed. This still leaves open the possibility of duplicate messages."

When duplicates are fine:

"In simple systems where messages are produced and then consumed by various applications, duplicates are an annoyance that is fairly easy to handle. Most real-world applications contain unique identifiers that consuming applications can use to deduplicate."

When they are not — and this is the whole argument for this chapter:

"Things become more complicated when we look at stream processing applications that AGGREGATE events. When inspecting an application that consumes events, computes an average, and produces the results, it is often IMPOSSIBLE for those who check the results to DETECT that the average is incorrect because an event was processed twice."

Single-record transform / filterAggregation
dup in → dup outdup in → SILENTLY WRONG NUMBER out
Detectable (same unique ID).UNDETECTABLE from the output.
Fixable (dedupe downstream).“Impossible to correct the result WITHOUT REPROCESSING THE INPUT.”

This is the key asymmetry: aggregation destroys the evidence. A duplicated record leaves a fingerprint you can dedupe; a duplicated contribution to a sum leaves nothing.