Learn Labs
4. Storage and Retrieval

4. Storage and Retrieval

Chapter 4 of Designing Data-Intensive Applications — 15 sections.

"A computer does not primarily compute in the sense of doing arithmetic. […] They primarily are filing systems." — Richard Feynman

Ch 3 was the user's view (what format you give the database, what interface you query it through). Ch 4 is the database's view: how it stores what you give it, and how it finds it again.

Why you should care even though you'll never write a storage engine: you must select the right one from many available, and to configure it to perform well on your workload you need a rough idea of what it's doing under the hood.

The chapter's map:

storage enginesOLTPsmall reads/writes by keyOLAPscan + aggregate many rowslog-structuredLSM, SSTables — immutableupdate-in-placeB-trees — fixed pagescolumn-orientedParquet, ORCmaterialized viewsand data cubesPlus the specialised kinds: multidimensional (R-tree), full-text (inverted index), vector (HNSW / IVF).
Figure 4.0.1The chapter's map