3.10 Self-test
Self-test35 questions
Give the four layers of the data-model stack and say who owns each.
What are the two advantages of declarative query languages beyond conciseness? Which one is more important, and why?
Why did every challenger to the relational model fail while SQL absorbed their ideas?
What is the impedance mismatch? Give three criticisms of ORMs and two genuine advantages.
Explain the N+1 problem and how you avoid it.
Why does the document model suit the résumé example? Name the specific limit expressed by "one-to-few."
Give five reasons to store a region ID rather than the region name.
State the normalization trade-off in one sentence for reads and one for writes. Why is denormalization "a form of derived data"?
What exactly does X's materialized timeline store, and what does it deliberately not store? Give both reasons.
What is hydration, and why does the book say join-on-read is not an obstacle to scalability?
Which relationship types don't fit in a single document? Give the two ways to query a many-to-many bidirectionally and the risk of each.
Draw a star schema. What is in a fact table row, and what do the dimensions represent?
Why is a fact table's multi-item transaction not represented explicitly?
What is OBT, and why is aggressive denormalization safe in analytics but not OLTP?
Contrast schema-on-read with schema-on-write using the static/dynamic typing analogy. Show the migration for splitting a name field both ways.
Why is "schemaless" misleading? When is schema-on-read genuinely better?
What is the locality advantage of documents, and what are its two limits? Name three non-document systems that provide locality.
Why does a graph query mean a variable number of joins? Why is that hard in SQL?
Describe the property graph model. Why are there indexes on both tail and head vertex?
Give three things about the Lucy/Alain example that are hard in a relational schema.
What does
*0..mean, and what is its regex analogy? Give the two execution strategies an optimizer might choose.Map triple-store concepts onto property-graph concepts. What are the two possible kinds of object?
Why does RDF use URIs as predicates?
What is the one syntactic advantage SPARQL gets from RDF not distinguishing properties from edges?
How does Datalog differ in style from Cypher and SPARQL? Trace the
within_recursiverules for Idaho.Why is GraphQL deliberately restricted? Name two things it forbids and the reason.
Why does GraphQL accept duplication in its responses? Give both cases.
Define event sourcing and CQRS. What is a command, when does it become an event, and what may a projection never do?
Why must events be named in the past tense?
Give two differences between an event log and a star-schema fact table.
List four advantages of event sourcing and all three downsides. Which downside is a determinism problem?
What is the one hard requirement on the event storage system, and why is it difficult in a distributed system?
What is a DataFrame, what is
merge, and why do data scientists prefer it to SQL?Explain one-hot encoding and why a sparse matrix doesn't fit a relational database.
- Design question
you're modeling a project-management product: workspaces contain projects, projects contain tasks, tasks have ordered subtasks users drag to reorder, tasks have many assignees, and users want full-text search plus "show me everything blocking this task" (transitive dependencies of unknown depth). Choose a model (or models) for each part, justify each against a specific trade-off from this chapter, and identify the one query that would push you toward a graph database — and the one that would make you regret it.