11.4 Batch Use Cases
You'll find batch jobs WHEREVER THERE'S A LOT OF DATA AND DATA FRESHNESS ISN'T IMPORTANT. This might sound limiting, but IT TURNS OUT THAT A SIGNIFICANT AMOUNT OF DATA PROCESSING TASKS FIT THIS MODEL:
- Accounting and inventory reconciliation — verifying transactions line up with bank accounts and inventory
- Demand forecasting in manufacturing — a periodic batch job
- Ecommerce, media, and social media train their RECOMMENDATION MODELS with batch jobs
"Many financial systems are batch-based; for example, THE US BANKING NETWORK RUNS ALMOST ENTIRELY ON BATCH JOBS."
4.1 ETL / ELT
Why batch fits so well:
- The PARALLEL nature suits transformation, much of which is "EMBARRASSINGLY PARALLEL" — filtering, projecting fields, and many common warehouse transformations can all be done in parallel
- Robust workflow schedulers make it easy to schedule, orchestrate, and debug. Schedulers RETRY jobs to mitigate transient issues; a job that fails repeatedly is MARKED AS FAILED, which helps developers EASILY SEE WHICH JOB STOPPED WORKING. Airflow ships built-in source, sink, and query operators for MySQL, PostgreSQL, Snowflake, Spark, Flink, and dozens of others
- Easy to troubleshoot — "Failed files can be easily inspected to see what went wrong, and ETL batch jobs can be FIXED AND RERUN." If an input file lacks a field the job needs, engineers can easily spot it and update either the transformation or the job that produced the input
The organizational shift:
"Data pipelines USED TO BE MANAGED BY A SINGLE DATA ENGINEERING TEAM, as it was considered UNFAIR to ask product teams to write and manage complex batch pipelines. Recently, improvements in batch processing models and METADATA MANAGEMENT have made it MUCH EASIER FOR ENGINEERS ACROSS AN ORGANIZATION TO CONTRIBUTE TO AND MANAGE THEIR OWN PIPELINES." — data mesh, data contracts, data fabric provide standards and tools to help teams SAFELY PUBLISH THEIR DATA for consumption by anybody in the organization.
"Many batch ETL jobs now run ON THE SAME SYSTEMS as the analytical queries that read their output — SparkSQL, Trino, or DuckDB. Such an architecture FURTHER BLURS THE LINE BETWEEN APPLICATION ENGINEERING, DATA ENGINEERING, ANALYTICS ENGINEERING, AND BUSINESS ANALYSIS."
4.2 Analytics — the lakehouse
Analysts write SQL that executes atop a query engine reading from and writing to a DFS or object store. Table metadata is managed with TABLE FORMATS such as Apache Iceberg and CATALOGS such as Unity. THIS ARCHITECTURE IS KNOWN AS A DATA LAKEHOUSE.
| Query style | Characteristics |
|---|---|
| Pre-aggregation | Data rolled up into OLAP cubes or data marts (Ch 4 §7.7). Pre-aggregated data is queried in the warehouse OR PUSHED TO PURPOSE-BUILT REAL-TIME OLAP SYSTEMS such as Druid or Pinot. Runs at a scheduled interval, managed by workflow schedulers |
| Ad hoc | Response times are IMPORTANT here. Analysts run queries ITERATIVELY as they get responses and learn more about the data. Fast query execution REDUCES WAITING TIMES |
BI integration: SQL support enables Tableau, Power BI, Looker, Apache Superset — "Tableau offers SparkSQL and Presto connectors; Superset supports Trino, Hive, Spark SQL, Presto."
4.3 Machine learning
| Use | What goes in / comes out |
|---|---|
| Feature engineering | Raw data filtered and transformed into data models can train on. Predictive models often need NUMERIC data, so engineers must transform TEXT OR DISCRETE VALUES into the required format |
| Model training | Training data is the INPUT; the WEIGHTS of the trained model are the OUTPUT |
| Batch inference | Predictions in bulk when datasets are large and real-time results aren't required — including EVALUATING THE MODEL'S PREDICTIONS ON A TEST DATASET |
(Tooling: Spark MLlib, Flink FlinkML — feature engineering tools, statistical functions, classifiers.)
Graph processing: recommendation engines and ranking systems use it heavily.
Many graph algorithms are expressed by TRAVERSING ONE EDGE AT A TIME, joining one vertex with an adjacent vertex TO PROPAGATE SOME INFORMATION, and REPEATING UNTIL A CERTAIN CONDITION IS MET — until there are no more edges to follow, or until a metric CONVERGES.
The BULK SYNCHRONOUS PARALLEL (BSP) model has become popular — implemented by Apache Giraph, Spark's GraphX, Flink's Gelly. Also known as the PREGEL MODEL, after Google's Pregel paper.
LLM data preparation — batch's newest big job:
“OpenAI uses Ray as part of its ChatGPT training process.” These frameworks ship built-in integrations for PyTorch, TensorFlow and XGBoost, and built-in support for feature engineering, model training, batch inference, and fine-tuning.
(And notebooks — Jupyter, Hex — where cells of Markdown/Python/SQL execute sequentially, many using batch processing via DataFrame APIs or SQL.)
4.4 Serving derived data — and the anti-pattern
Batch jobs build precomputed datasets — product recommendations, user-facing reports, ML features — typically served from a production database, key-value store, or search engine. THE PRECOMPUTED DATA NEEDS TO MAKE ITS WAY FROM THE BATCH PROCESSOR'S STORAGE BACK INTO THE DATABASE SERVING LIVE TRAFFIC.
❌ THE ANTI-PATTERN: writing directly to the production database from inside the job.
“You might be tempted to use the client library for your favorite database directly within a batch job and write to the database server One record at a time. This Will work (assuming your firewall rules allow it), But it is a bad idea:”
- “Making a network request For every single record is Orders of magnitude slower than the normal throughput of a batch task. Even if the client library supports batching, performance is likely to be poor.”
- “Batch frameworks run Many tasks in parallel. If all tasks concurrently write to the same output database At the rate expected of a batch process, that database can easily be overwhelmed, and its query performance is likely to suffer. This can in turn cause Operational problems in other parts of the system.”
- “Normally batch jobs provide a clean All-or-nothing guarantee... However, Writing to an external system from inside a job produces externally visible side effects that cannot be hidden. You have to worry about results from Partially completed jobs being visible, and If a task fails and restarts, it may duplicate output.”
✔ Solution A — push to a stream (Kafka):
- “Streaming systems are optimized for sequential writes, better suited to the bulk write workload of a batch job.”
- “They act as a buffer between the batch job and production databases. Downstream systems can throttle their read rate to ensure they can continue to comfortably serve production traffic.”
- “The output of a single batch job can be consumed by multiple downstream systems.”
- “They can serve as a security boundary between batch and production networks — deployed in a DMZ network that sits between them.”
Still unsolved: the all-or-nothing guarantee. “Upon completion, batch jobs must send a notification to downstream systems that the job is done and the data can now be served. Consumers must be able to keep the data they receive invisible to queries — like an uncommitted transaction with read-committed isolation — until they are notified that the job is complete.”
✔ Solution B — build the database in the job and bulk-load it:
Build a BRAND-NEW DATABASE INSIDE the batch job and BULK-LOAD those files directly. Tools: TiDB's Lightning, Apache Pinot's Hadoop import jobs, RocksDB's SST bulk-import API.
VERY FAST, and makes it easier for systems to ATOMICALLY SWITCH BETWEEN DATASET VERSIONS. On the other hand, IT CAN BE CHALLENGING TO INCREMENTALLY UPDATE datasets from jobs that build brand-new databases.
It's common to take a HYBRID approach when both bootstrapping and incremental loads are needed — Venice supports hybrid stores allowing batch row-based updates AND full dataset swaps.