4. Storage and Retrieval
4.10 Production failure catalog for this chapter
| Symptom | Underlying mechanism |
|---|---|
| Write latency jumps 1 ms → 30 s, no code change | LSM compaction can't keep up; engine applied backpressure |
| Reads slow down over days, then recover after a restart | L0 file accumulation / read amplification |
| Query times out scanning a "small" table | Tombstones not yet compacted away |
| Deleted data still readable weeks later | Tombstone hasn't propagated through all compaction levels |
| Disk full during routine maintenance | Size-tiered compaction temp space |
| Benchmark 5× better than production | Benchmarked an empty LSM — no compaction yet |
| Range queries far slower than point queries | Bloom filters don't help range queries |
| Insert throughput collapses after switching to UUIDv4 PK | Random writes + constant page splits in a clustered B-tree |
| Table is 4× its logical size after a mass delete | B-tree fragmentation; needs vacuum/rebuild |
| Periodic latency spikes every few minutes | Checkpoint flushing buffered dirty pages |
| Everything fine until data exceeded RAM, then 100× slower | Working set no longer fits the buffer pool / page cache |
| Warehouse query costs $4,000 | No sort key / no partition pruning → full scan |
| Query planning takes longer than the query | Small-file problem — a footer read per file |
SELECT * is 50× slower than naming 3 columns | Defeats columnar storage |
| Search cluster falls over with thousands of shards | Oversharding — every shard is a full Lucene index |
| Semantic search quality quietly degrades | Recall collapse from low ef_search/nprobe, or mixed embedding model versions |
| Vector search returns nothing when filtered | Filtered ANN — pre/post-filtering interacting badly with the graph |