4.6 Keeping everything in memory
Everything so far is an answer to the limitations of disks.
Everything so far is an answer to the limitations of disks. We tolerate the awkwardness because disks are durable and have lower cost per gigabyte than RAM.
As RAM gets cheaper, the cost-per-GB argument erodes. Many datasets are simply not that big. Hence in-memory databases.
| System | Durability approach |
|---|---|
| Memcached | Caching only — acceptable to lose data on restart |
| VoltDB, SingleStore, Oracle TimesTen | In-memory relational; vendors claim big gains from removing on-disk data structure overheads |
| RAMCloud | Open source in-memory KV store with durability, log-structured for both memory and disk |
| Redis, Couchbase | Weak durability — asynchronous disk writes |
Durability is achieved by special hardware (battery-powered RAM) or, more commonly, writing a change log to disk, writing periodic snapshots, or replicating the in-memory state to other machines.
These are still "in-memory" databases because the disk is merely an append-only log for durability, and reads are served entirely from memory. Writing to disk also brings operational advantages: files can be backed up, inspected, and analyzed by external utilities.
⚠️ The counterintuitive point most people get wrong:
The performance advantage of in-memory databases is NOT because they avoid reading from disk. A disk-based engine may never read from disk either, if you have enough memory, because the OS caches recently used disk blocks anyway.
They are faster because they avoid the overheads of ENCODING in-memory data structures into a form that can be written to disk.
A second, under-appreciated use case: in-memory databases can offer data models that are difficult to implement with disk-based indexes. Redis offers a database-like interface to priority queues and sets — its implementation is comparatively simple because everything is in memory.
4.5 Secondary indexes and where values live
Both B-trees and log-structured storage can implement an index.
4.7 Data Storage for Analytics
Data warehouses are usually relational, because SQL fits analytical queries well, and many graphical tools generate SQL and support drill-down and slicing and dicing.