1.0 The mental model for the whole book
The hard part is never one block. It's choosing between blocks with different characteristics, and gluing blocks together when no single tool does the job.
Definition worth memorizing: an application is data-intensive if data management is the primary engineering challenge — storing large volumes, managing change, ensuring consistency under failure and concurrency, staying available. Contrast with compute-intensive, where the challenge is parallelizing a single big computation.
Standard building blocks every non-trivial app assembles:
| Block | Job |
|---|---|
| Databases | store data so it can be found again later |
| Caches | remember the result of an expensive operation |
| Search indexes | search by keyword / filter in arbitrary ways |
| Stream processing | react to events as they occur |
| Batch processing | periodically crunch accumulated data |
The hard part is never one block. It's choosing between blocks with different characteristics, and gluing blocks together when no single tool does the job.
1. Trade-Offs in Data Systems Architecture
Chapter 1 of Designing Data-Intensive Applications — 12 sections.
1.1 Operational vs Analytical Systems
Key observation: analysts and scientists both read data that users and backend services generated, and they do not modify it (they may create derived datasets).