4.12 Worked examples
Leaf pages = 500³ = 125,000,000. Capacity = 125,000,000 × 4 KiB = 500 GB of leaf pages; the book's figure of ~250 TB assumes 500⁴ leaf entries' worth of addressable data (4 levels…
① B-tree capacity. Page 4 KiB, branching factor 500, depth 4. Leaf pages = 500³ = 125,000,000. Capacity = 125,000,000 × 4 KiB = 500 GB of leaf pages; the book's figure of ~250 TB assumes 500⁴ leaf entries' worth of addressable data (4 levels of references). The point to retain: depth 3–4 covers essentially any real database, so a lookup is 3–4 page reads.
② Bloom filter sizing. You have 10 million keys per SSTable and want a 0.1% false-positive rate. Rule: 10 bits/key → 1%; each +5 bits/key divides FPP by 10. So 0.1% needs 15 bits/key = 150,000,000 bits = ~18.8 MB per SSTable. For 0.01%: 20 bits/key = ~25 MB.
③ Write amplification. An LSM with leveled compaction, 7 levels, fan-out 10.
A key is written: once to the WAL, once on memtable flush, and once per level it is promoted through — roughly 1 + 1 + 7 ≈ 10× write amplification (real-world leveled RocksDB is typically 10–30×). Size-tiered is lower (~4–10×) but has higher space amplification. A B-tree writing a full 8 KiB page for a 100-byte row change has ~80× amplification for that write — which is why full_page_writes after a checkpoint is so expensive.
④ Columnar I/O saving. Fact table: 1 billion rows × 100 columns × 8 bytes = 800 GB. Query touches 3 columns → row store reads 800 GB; column store reads 1e9 × 3 × 8 = 24 GB, and with 4:1 compression ~6 GB. That's the >100× difference that makes interactive analytics possible.
⑤ Bitmap compression. product_sk with 100,000 distinct values over 1 billion rows.
Raw: 1e9 × 4 bytes = 4 GB. Bitmaps: 100,000 bitmaps × 1e9 bits = 12.5 TB uncompressed — worse! But each bitmap is ~99.999% zeros, so run-length encoding collapses it. The lesson: bitmap indexes only pay off when combined with RLE/roaring, and the win grows as cardinality falls. For a column with 5 distinct values sorted as the primary key, RLE takes it to a few kilobytes over a billion rows.
⑥ HNSW memory. 5 million documents, 1536-dim float32 embeddings, M=16. Vectors: 5e6 × 1536 × 4 B = 30.7 GB. Graph: 5e6 × 16 × 8 B ≈ 0.64 GB per layer, ~1.3 GB total. ~32 GB RAM — before you've served a single query. With product quantization to 8× compression, ~4 GB. This is why "just add vector search" is a capacity-planning conversation.