12. Stream Processing
12.7 Production failure catalog for this chapter
| Symptom | Underlying mechanism |
|---|---|
| Metrics quietly wrong; nobody noticed | UDP/at-most-once delivery; dropped messages aren't visible |
| Messages processed out of order | Load balancing + redelivery (§1.4), or cross-partition ordering assumed |
| One bad message loops forever, blocking the queue | Poison message with no DLQ |
| A new consumer can't see historical data | AMQP/JMS destructive consumption |
| Consumer falls behind, then silently skips data | Retention expiry past the consumer offset |
| Search index and database permanently disagree | Dual-write race (§3.1) |
| Half a write succeeded | Dual write partial failure — the atomic commit problem |
| Source database's disk fills up | Inactive CDC replication slot pinning WAL |
| Dropping a column caused a customer-facing outage | CDC turned the schema into a public API |
| Cannot rebuild a derived system without a snapshot | No log compaction on the change topic |
| Event-sourced log can't be compacted | Events express intent, not final state |
| GDPR erasure impossible | Immutability + copies everywhere; crypto-shredding not designed in |
| A traffic "spike" that never happened | Windowing by processing time after a redeploy backlog |
| Window results wrong after a network blip | Straggler events arriving after the window closed |
| Mobile events arrive days late with wrong timestamps | Untrusted device clock; no three-timestamp correction |
| Job emits nothing; looks healthy | Watermark stalled by an idle partition |
| Checkpoints grow until the job can't progress | Unbounded state — no TTL, too-wide join window |
| Stream job OOMs at high throughput | Sliding windows / joins buffering events |
| Reprocessing produces different results | Nondeterministic join across streams (time dependence) |
| Emails sent twice after a crash | Side effect outside the framework — microbatching/checkpointing can't undo it |
| Counter double-incremented on retry | Non-idempotent operation + at-least-once delivery |
| Zombie processor writes stale state | No fencing on failover (Ch 9) |
| Rebalance takes 10 minutes | Large local state restore from the changelog |