Learn Labs
1. Trade-Offs in Data Systems Architecture

1.0 The mental model for the whole book

The hard part is never one block. It's choosing between blocks with different characteristics, and gluing blocks together when no single tool does the job.

HTTP / WebSocketwritesETL / CDCusers / devicesfrontendbrowser, mobile — one user's dataapplication codeusually statelessdatabasesystem of recordcachederivedsearch indexderivedqueuederiveddata lakefileswarehousetablesBI / ML— data infrastructure — everything but the database is derived dataAnalytical side is a read-only copy: losing it costs a rebuild, not the business.
Figure 1.0.1The mental model for the whole book

Definition worth memorizing: an application is data-intensive if data management is the primary engineering challenge — storing large volumes, managing change, ensuring consistency under failure and concurrency, staying available. Contrast with compute-intensive, where the challenge is parallelizing a single big computation.

Standard building blocks every non-trivial app assembles:

BlockJob
Databasesstore data so it can be found again later
Cachesremember the result of an expensive operation
Search indexessearch by keyword / filter in arbitrary ways
Stream processingreact to events as they occur
Batch processingperiodically crunch accumulated data

The hard part is never one block. It's choosing between blocks with different characteristics, and gluing blocks together when no single tool does the job.