Learn Labs
13. A Philosophy of Streaming Systems

13.10 Self-test

Self-test44 questions

—/44
  1. Why does "every piece of software, even a general-purpose database, is designed for a particular usage pattern" lead inevitably to composing multiple systems?

  2. Draw the good and bad dataflow topologies for a database + search index. What exactly goes wrong in the bad one?

  3. What principle matters more than the choice between CDC and event sourcing, and why?

  4. Compare distributed transactions and log-based derived data on mechanism and on the one guarantee that genuinely differs.

  5. Give the four limits of total ordering. Which one applies to microservices, and which to offline-capable clients?

  6. Tell the unfriending story. Why is it fundamentally a join problem? Give the three partial mitigations and the limitation of each.

  7. Why is asynchrony "what makes systems based on event logs robust"? Contrast with what distributed transactions do to a local fault.

  8. Explain the railway gauge migration and map each step onto a data-system migration. What property makes it safe?

  9. What was the lambda architecture, and what replaced it? List the three features required to unify batch and stream.

  10. Walk through what CREATE INDEX does. Name the two other operations in this book that follow the same four steps.

  11. Explain "the dataflow across an entire organization looks like one huge database." What are batch and stream processors, in that metaphor?

  12. Distinguish federation from unbundling. Which tradition does each follow, and which problem is harder?

  13. Give the two levels at which log-based integration provides loose coupling.

  14. State the argument against unbundling. What is the goal of unbundling — and what is it explicitly not?

  15. What did VisiCalc have in 1979 that data systems still lack? What three requirements make it hard to replicate?

  16. Give four derived datasets and their derivation functions. Which are cookie-cutter and which require custom code?

  17. Why are databases poor deployment environments for application code? Explain the Church-and-state joke.

  18. Why is a database a "passive" mutable variable, and what would it take to make it active?

  19. Explain the currency-conversion example. Why is the dataflow version both faster and more robust? What problem does it not remove?

  20. Define the write path and the read path. Which is eager and which is lazy? What sits at the boundary?

  21. Walk the full-text search spectrum from no-index to precompute-everything. Why is the far end impossible, and what's the practical middle?

  22. What does it mean to say the pixels on screen are a materialized view? What does extending the write path to the device require, and why is offline already solved?

  23. Explain "reads are events too." What join is being performed, and what distinguishes a one-off read from a subscription?

  24. Why might you want to log read events, and what does it cost?

  25. Trace the four layers at which duplicate suppression fails for a money transfer. At which layer does it actually break, and why can't the layer below fix it?

  26. Write the request-ID transaction. Why does the uniqueness constraint work even at weak isolation, and what does the requests table give you for free?

  27. State the end-to-end argument verbatim in substance. Apply it to duplicate suppression, integrity checking, and encryption.

  28. If low-level reliability mechanisms can't provide end-to-end correctness, why keep them?

  29. Describe enforcing unique usernames with a shared log. What general principle does it embody, and what replication model does it rule out?

  30. Walk through the four steps of the multishard money transfer. Where exactly does atomicity come from? What are the three requirements?

  31. What happens if the source processor crashes mid-request? Show why the outcome is still correct.

  32. Define timeliness and integrity precisely. State the slogan. Which one is catastrophic when violated, and why?

  33. Use the credit card statement to illustrate both. What would be a timeliness violation, and what an integrity violation?

  34. Name the four mechanisms by which dataflow systems achieve integrity without distributed transactions.

  35. Give three business situations where a "hard" constraint is deliberately violated. What is the forklift argument?

  36. What is a compensating transaction? Give examples ordered by cost of apology.

  37. State the two observations that combine into coordination-avoiding data systems. What guarantee do they keep, and what do they give up?

  38. Explain the apology calculus. Why can't you drive apologies to zero?

  39. Why does ACID consistency "make sense only if we assume the transaction is free from bugs"?

  40. What do HDFS and S3 do that most systems don't? What is the corresponding advice about backups?

  41. Why do event-based systems audit better than mutation logs? What can you check for the event log, and what for derived state?

  42. Why is end-to-end integrity checking better than per-component checking? What does it implicitly cover?

  43. What can a hardware-signed transaction log not guarantee? What do blockchains add, and why are they usually the wrong tool?

  44. Design question

    you're building a ride-hailing platform. Requirements: (a) a driver can accept at most one ride at a time; (b) surge pricing must be computed from live demand within 5 seconds; (c) the rider's app must show the car moving in real time and keep working through a tunnel; (d) daily payouts to drivers must be exactly correct, with a full audit trail; (e) the whole thing runs in 3 regions and must survive losing one. For each requirement: state whether you need timeliness or integrity or both, whether the constraint is hard or loosely-interpretable, which mechanism you'd use, where the write path/read path boundary sits, and what audit proves it's working. Identify the one place you'd accept synchronous coordination and justify why nothing else needs it.