11.1 Batch Processing with Unix Tools
The NGINX access log line (one line, wrapped for readability):
216.58.210.78 - - [27/Jun/2025:17:55:11 +0000] "GET /css/typography.css HTTP/1.1"
200 3377 "https://martin.kleppmann.com/" "Mozilla/5.0 (Macintosh; …) Chrome/137.0.0.0 …"Format: $remote_addr - $remote_user [$time_local] "$request" $status $body_bytes_sent "$http_referer" "$http_user_agent"
Though log parsing might seem contrived, it's A CRITICAL PART OF THE OPERATIONS OF MANY MODERN TECHNOLOGY COMPANIES — used for everything from AD PIPELINES to PAYMENT PROCESSING. Indeed, IT WAS A DRIVING FORCE BEHIND THE RAPID ADOPTION OF MAPREDUCE AND THE "BIG DATA" MOVEMENT.
1.1 The five-command pipeline
cat /var/log/nginx/access.log | # read the log
awk '{print $7}' | # 7th field = the requested URL
sort | # so identical URLs become ADJACENT
uniq -c | # collapse adjacent duplicates, -c = count them
sort -r -n | # sort by the leading NUMBER, REVERSED
head -n 5 # top 54189 /favicon.ico
3631 /2016/02/08/how-to-do-distributed-locking.html
2124 /2020/11/18/distributed-systems-and-elliptic-curves.html
1369 /
915 /css/typography.cssIt will process GIGABYTES of log files IN A MATTER OF SECONDS, and you can easily modify the analysis. Omit CSS files: change awk to
$7 !~ /\.css$/ {print $7}. Count top client IPs:{print $1}.Many data analyses can be done in a few minutes using a combination of
awk,sed,grep,sort,uniq, andxargs, and THEY PERFORM SURPRISINGLY WELL.
1.2 The crucial contrast: sorting vs in-memory aggregation
The Python equivalent keeps an in-memory hash table url → count. The Unix pipeline has NO hash table — it relies on SORTING a list in which multiple occurrences are simply repeated.
The working set is the amount of memory to which the job needs random access. Which approach you can afford follows from it.
| Hash table | Sorting |
|---|---|
| The working set depends only on the number of distinct URLs. “If there are a million log entries for a single URL, the space required is still just one URL plus a counter.” | Works when the working set exceeds available memory. |
| Fits in roughly 1 GB for most small-to-midsize sites — fine on a laptop. | The same principle as log-structured storage (Ch 4): chunks sorted in memory → written out as segment files → multiple sorted segments merged into a larger sorted file. “Mergesort has sequential access patterns that perform well on disks.” |
GNU Coreutils
sortAUTOMATICALLY handles larger-than-memory datasets by SPILLING TO DISK and AUTOMATICALLY PARALLELIZES sorting across multiple CPU cores. The simple chain of Unix commands EASILY SCALES TO LARGE DATASETS without running out of memory. THE BOTTLENECK IS LIKELY TO BE THE RATE AT WHICH THE INPUT FILE CAN BE READ FROM DISK.A limitation of Unix tools is that THEY RUN ON A SINGLE MACHINE — and that's where distributed batch processing frameworks come in.