11. Batch Processing
11.6 Production failure catalog for this chapter
| Symptom | Underlying mechanism |
|---|---|
| Bad data written to the DB; rolling back the code doesn't fix it | Online systems lack human fault tolerance — batch's rerun-from-immutable-input does |
| One byte changed → the whole dataset reprocessed | Batch's fundamental inefficiency (→ stream processing, Ch 12) |
| Job succeeded but produced no/garbage data | Success ≠ correctness; no monitoring job comparing to the previous run |
| 199 tasks finish in seconds, one runs for hours | Data skew on the shuffle key |
| Shuffle spills to disk and the job takes 10× longer | Working set exceeds executor memory |
| Job restarts from scratch after one node dies | Intermediate data not checkpointed / lineage lost |
| Whole upstream stage recomputed repeatedly on spot instances | Preemption + in-memory shuffle output with no external shuffle service |
| Output is 40,000 tiny Parquet files | No repartition before write → small-file problem |
| Job "commit" step takes hours on S3 | Rename-based commit protocol on a store with nonatomic rename |
| NameNode out of memory | Millions of small files — metadata is per-file |
| List-then-process silently misses new files | Object-store listing semantics assumed to be a directory listing |
| Production DB melts when the nightly job runs | Writing directly to the production database from parallel tasks |
| Partially-complete job's output visible to users | External side effects break the all-or-nothing guarantee |
| Duplicate records after a task retry | Same — a restarted task duplicates external writes |
| Airflow scheduler falls behind | Heavy top-level code in DAG files re-executed on every parse |
| A backfill takes down the source database | Hundreds of concurrent DagRuns |
| Job processed yesterday's data — or tomorrow's | execution_date / timezone semantics |
| Cluster deadlocks with everything half-scheduled | Gang scheduling holding partial allocations |
| Large jobs never run | Starvation by a stream of small jobs |
Pandas code is instant, Spark code hangs on .count() | Eager vs lazy evaluation |
collect() kills the driver | Pulling a distributed dataset into one JVM |