Learn Labs
11. Batch Processing

11.1 Batch Processing with Unix Tools

The NGINX access log line (one line, wrapped for readability):

216.58.210.78 - - [27/Jun/2025:17:55:11 +0000] "GET /css/typography.css HTTP/1.1"
200 3377 "https://martin.kleppmann.com/" "Mozilla/5.0 (Macintosh; …) Chrome/137.0.0.0 …"

Format: $remote_addr - $remote_user [$time_local] "$request" $status $body_bytes_sent "$http_referer" "$http_user_agent"

Though log parsing might seem contrived, it's A CRITICAL PART OF THE OPERATIONS OF MANY MODERN TECHNOLOGY COMPANIES — used for everything from AD PIPELINES to PAYMENT PROCESSING. Indeed, IT WAS A DRIVING FORCE BEHIND THE RAPID ADOPTION OF MAPREDUCE AND THE "BIG DATA" MOVEMENT.

1.1 The five-command pipeline

cat /var/log/nginx/access.log |   # read the log
  awk '{print $7}'            |   # 7th field = the requested URL
  sort                        |   # so identical URLs become ADJACENT
  uniq -c                     |   # collapse adjacent duplicates, -c = count them
  sort -r -n                  |   # sort by the leading NUMBER, REVERSED
  head -n 5                       # top 5
4189 /favicon.ico
3631 /2016/02/08/how-to-do-distributed-locking.html
2124 /2020/11/18/distributed-systems-and-elliptic-curves.html
1369 /
 915 /css/typography.css

It will process GIGABYTES of log files IN A MATTER OF SECONDS, and you can easily modify the analysis. Omit CSS files: change awk to $7 !~ /\.css$/ {print $7}. Count top client IPs: {print $1}.

Many data analyses can be done in a few minutes using a combination of awk, sed, grep, sort, uniq, and xargs, and THEY PERFORM SURPRISINGLY WELL.

1.2 The crucial contrast: sorting vs in-memory aggregation

The Python equivalent keeps an in-memory hash table url → count. The Unix pipeline has NO hash table — it relies on SORTING a list in which multiple occurrences are simply repeated.

The working set is the amount of memory to which the job needs random access. Which approach you can afford follows from it.

Hash tableSorting
The working set depends only on the number of distinct URLs. “If there are a million log entries for a single URL, the space required is still just one URL plus a counter.”Works when the working set exceeds available memory.
Fits in roughly 1 GB for most small-to-midsize sites — fine on a laptop.The same principle as log-structured storage (Ch 4): chunks sorted in memory → written out as segment files → multiple sorted segments merged into a larger sorted file. “Mergesort has sequential access patterns that perform well on disks.”

GNU Coreutils sort AUTOMATICALLY handles larger-than-memory datasets by SPILLING TO DISK and AUTOMATICALLY PARALLELIZES sorting across multiple CPU cores. The simple chain of Unix commands EASILY SCALES TO LARGE DATASETS without running out of memory. THE BOTTLENECK IS LIKELY TO BE THE RATE AT WHICH THE INPUT FILE CAN BE READ FROM DISK.

A limitation of Unix tools is that THEY RUN ON A SINGLE MACHINE — and that's where distributed batch processing frameworks come in.


On this page