11.9 Self-test
Self-test38 questions
Give the four benefits of treating batch inputs as immutable with no side effects. Which one do read/write databases fundamentally lack, and why?
State the two costs of batch processing. Which one motivates stream processing?
Walk through the five-command Unix pipeline and say what each stage contributes. Why is the first
sortthere?What is a job's "working set"? Why does it depend on distinct keys rather than record count?
When does sorting beat an in-memory hash table, and what property of mergesort makes it work on disk?
Map the three components of a single machine onto the three components of a distributed batch framework.
Why are DFS blocks 128 MB rather than 4 KB? Give two reasons.
What plays the role of the VFS in a distributed filesystem? Give a concrete example of why that matters commercially.
Contrast shared-nothing DFS with NAS/SAN on hardware, cost, and failure handling.
Give five ways object stores differ from filesystems that would break naive code.
Why does HDFS allow computation on the node holding the data, and why is that less compelling now?
Name the three orchestrator components and what each is responsible for. Where is cluster state stored in YARN and in Kubernetes?
Work through the 160-core, two-job scheduling example. Define gang scheduling, starvation, and preemption, and give the downside of each.
Why is optimal scheduling intractable, and what do real schedulers do instead?
In batch processing, what does "workflow" mean, and how does it differ from Ch 5's usage?
Contrast pipe-style coupling with file-style coupling between jobs. Which is more typical, and why?
Why doesn't Spark's own scheduler manage workflows? What tools do?
What are spot instances, and why is batch processing unusually well suited to them? What's the catch?
Why are task failures easier to handle in batch than online? At what granularity is work retried?
Compare MapReduce, Spark, and Flink on how they handle intermediate data and its loss.
Give the four steps of MapReduce and map each onto the Unix pipeline. Which step do you not write?
Why does avoiding mutable state enable both parallelism and retries?
Give the six advantages of dataflow engines over MapReduce. Which one directly fixes MapReduce's inability to pipeline?
Why is "shuffle" a misleading name? Describe the full data path from mapper output to reducer output.
What determines the number of map tasks? What determines the number of reduce tasks?
Explain a sort-merge join. What is secondary sort, and what does it buy the reducer?
Why does a reducer in a sort-merge join need only one user record in memory and no network requests?
Give two machine-level reasons SQL improves batch jobs, not just human-level ones.
Name three workloads that are hard for cloud data warehouses and three ways warehouses and batch frameworks have converged.
Give two ways distributed DataFrames surprise people coming from Pandas.
List four real-world domains where batch dominates. Which one is the most surprising?
Give three reasons batch fits ETL particularly well.
What is a data lakehouse? Contrast pre-aggregation queries with ad hoc queries.
Name three ML uses of batch processing and describe the BSP/Pregel model.
List the three reasons you should not write directly to a production database from a batch job. Which one is about correctness rather than performance?
Give the four benefits of pushing batch output through a stream. What problem does streaming not solve, and what's the fix?
When is bulk-loading a purpose-built database file better than streaming, and what does it make hard?
- Design question
you must produce daily personalized recommendations for 50 million users from 2 TB/day of clickstream data, serve them at p99 < 20 ms, and be able to fix a bad model version within one hour of discovering it. Design the pipeline: storage layer, processing engine, workflow orchestration, fault-tolerance strategy, and the mechanism for getting output into the serving path. For each choice, name the specific trade-off from this chapter you're relying on, and state what happens when a single task fails, when the whole job produces bad output, and when the serving database is at capacity.