3.0 The layer stack
Most applications are built by layering one data model on another.
Most applications are built by layering one data model on another. For each layer the key question is: how is it represented in terms of the next-lower layer?
| Layer | What it is | Whose problem |
|---|---|---|
| 1 · Real world | Application objects, data structures, APIs — people, organizations, goods, actions, money flows, sensors | App developers |
| 2 · General-purpose data model | JSON/XML documents · tables and rows · vertices and edges — this chapter | App developers |
| 3 · Bytes | In memory, on disk, on the network — queryable, searchable, manipulable representations (Ch 4) | Database engineers |
| 4 · Physics | Electrical currents, pulses of light, magnetic fields | Hardware engineers |
Each layer hides the complexity of the layers below by providing a clean data model. That's what lets database vendors' engineers and application developers work together effectively without knowing each other's internals.
Declarative query languages — the key terminology note
SQL, Cypher, SPARQL, and Datalog are declarative: you specify the pattern of the data you want — what conditions results must meet, how they should be transformed (sorted, grouped, aggregated) — but not how to achieve it. The query optimizer decides which indexes and join algorithms to use, and in which order.
With imperative languages (Python, Java) you write the algorithm: which operations, in which order.
Why declarative wins:
- More concise and easier to write than an explicit algorithm.
- More importantly, it hides implementation details of the query engine, so the database can introduce performance improvements without any changes to your queries.
- It enables automatic parallelism — the DB can execute the query across multiple CPU cores and machines without you implementing that. In a handcoded algorithm, implementing parallel execution yourself is a lot of work.