Learn Labs
4. Storage and Retrieval

4.10 Production failure catalog for this chapter

Production failure catalog
0 rows
SymptomUnderlying mechanism
Write latency jumps 1 ms → 30 s, no code changeLSM compaction can't keep up; engine applied backpressure
Reads slow down over days, then recover after a restartL0 file accumulation / read amplification
Query times out scanning a "small" tableTombstones not yet compacted away
Deleted data still readable weeks laterTombstone hasn't propagated through all compaction levels
Disk full during routine maintenanceSize-tiered compaction temp space
Benchmark 5× better than productionBenchmarked an empty LSM — no compaction yet
Range queries far slower than point queriesBloom filters don't help range queries
Insert throughput collapses after switching to UUIDv4 PKRandom writes + constant page splits in a clustered B-tree
Table is 4× its logical size after a mass deleteB-tree fragmentation; needs vacuum/rebuild
Periodic latency spikes every few minutesCheckpoint flushing buffered dirty pages
Everything fine until data exceeded RAM, then 100× slowerWorking set no longer fits the buffer pool / page cache
Warehouse query costs $4,000No sort key / no partition pruning → full scan
Query planning takes longer than the querySmall-file problem — a footer read per file
SELECT * is 50× slower than naming 3 columnsDefeats columnar storage
Search cluster falls over with thousands of shardsOversharding — every shard is a full Lucene index
Semantic search quality quietly degradesRecall collapse from low ef_search/nprobe, or mixed embedding model versions
Vector search returns nothing when filteredFiltered ANN — pre/post-filtering interacting badly with the graph