Learn Labs
11. Batch Processing

11.6 Production failure catalog for this chapter

Production failure catalog
0 rows
SymptomUnderlying mechanism
Bad data written to the DB; rolling back the code doesn't fix itOnline systems lack human fault tolerance — batch's rerun-from-immutable-input does
One byte changed → the whole dataset reprocessedBatch's fundamental inefficiency (→ stream processing, Ch 12)
Job succeeded but produced no/garbage dataSuccess ≠ correctness; no monitoring job comparing to the previous run
199 tasks finish in seconds, one runs for hoursData skew on the shuffle key
Shuffle spills to disk and the job takes 10× longerWorking set exceeds executor memory
Job restarts from scratch after one node diesIntermediate data not checkpointed / lineage lost
Whole upstream stage recomputed repeatedly on spot instancesPreemption + in-memory shuffle output with no external shuffle service
Output is 40,000 tiny Parquet filesNo repartition before write → small-file problem
Job "commit" step takes hours on S3Rename-based commit protocol on a store with nonatomic rename
NameNode out of memoryMillions of small files — metadata is per-file
List-then-process silently misses new filesObject-store listing semantics assumed to be a directory listing
Production DB melts when the nightly job runsWriting directly to the production database from parallel tasks
Partially-complete job's output visible to usersExternal side effects break the all-or-nothing guarantee
Duplicate records after a task retrySame — a restarted task duplicates external writes
Airflow scheduler falls behindHeavy top-level code in DAG files re-executed on every parse
A backfill takes down the source databaseHundreds of concurrent DagRuns
Job processed yesterday's data — or tomorrow'sexecution_date / timezone semantics
Cluster deadlocks with everything half-scheduledGang scheduling holding partial allocations
Large jobs never runStarvation by a stream of small jobs
Pandas code is instant, Spark code hangs on .count()Eager vs lazy evaluation
collect() kills the driverPulling a distributed dataset into one JVM