Learn Labs
3. Data Models and Query Languages

3.3 Event Sourcing and CQRS

A conference management system is a genuinely complex domain:

The setup: in every model so far, data is queried in the same form it is written. But in complex applications it can be hard to find a single representation satisfying all the ways data needs to be queried and presented.

The move: write data in one form, then derive representations optimized for different types of reads. We saw this with systems-of-record vs derived data, and ETL. Now push it further:

If we're going to derive one representation from another anyway, we can choose representations optimized for writing and reading, respectively. How would you model data if you wanted to optimize it for ONLY writing, with efficient queries of no concern?

Answer: an event log. Encode each write as a self-contained string (perhaps JSON) including a timestamp, and append it. Events are immutable — never changed or deleted, only appended to (later events may supersede earlier ones). An event can contain arbitrary properties.

3.1 The conference example

A conference management system is a genuinely complex domain:

  • Individual attendees register and pay by card
  • Companies order seats in bulk, pay by invoice, and later assign seats to individuals
  • Seats reserved for speakers, sponsors, volunteers
  • Reservations may be canceled
  • The organizer might change the capacity by moving to a different room

With all this going on, simply calculating the number of available seats becomes a challenging query.

invalidvalid — now a factcommandsuser requestsvalidaterejectedappend-only event logthe source of truth · immutablebooking statusorganizer chartsbadge printer filesThe views are projections, or read models — any data model, denormalized freely, all rebuildable from the log.
Figure 3.3.13.1 The conference example

Definitions:

  • Event sourcing — using events as the source of truth, expressing every state change as an event
  • CQRS (command query responsibility segregation) — maintaining separate read-optimized representations derived from the write-optimized representation

Both terms originated in the DDD community, though similar ideas are old — e.g. state machine replication.

The command→event lifecycle:

  1. A request from a user is a command — it must first be validated
  2. Once executed and determined to be valid (e.g. there were enough seats), it becomes a fact, and the corresponding event is appended to the log
  3. Therefore the event log contains only valid events, and a consumer building a materialized view is NOT ALLOWED TO REJECT AN EVENT

Name your events in the PAST TENSE ("the seats were booked") — an event records that something has happened. Even if the user later cancels, the fact remains true that they formerly held a booking; the cancellation is a separate event added later.

Event sourcing vs star-schema fact table — similar (both are collections of past events), but:

Fact tableEvent log
ShapeAll rows have the same set of columnsMany event types, each with different properties
OrderUnordered collectionOrder is important — a booking made then canceled must not be processed in the wrong order

3.2 Advantages

  1. Events communicate intent. "The booking was canceled" is far easier to understand than "the active column on row 4001 of bookings was set to false, three rows were deleted from seat_assignments, and a refund row was inserted into payments." Those row modifications may still happen when a view processes the event — but driven by an event, the reason for the updates becomes much clearer.
  2. Reproducibility. A key principle: views are derived from the log in a reproducible way. You should always be able to delete the materialized views and recompute them by processing the same events in the same order with the same code. If the view-maintenance code had a bug, delete the view and recompute with the fixed code. Finding the bug is easier too, because you can rerun the view-maintenance code as often as you like and inspect its behavior.
  3. Multiple views, each optimized for particular queries. Stored in the same database as the events or a different one; any data model; denormalized for fast reads. You can even keep a view only in memory and never persist it, as long as recomputing from the log on restart is acceptable.
  4. Easy evolution. Present existing information in a new way → build a new view from the existing log. Support new features → add new event types or new properties to existing types (older events remain unmodified). Chain new behaviors off existing events — e.g. when an attendee cancels, offer their seat to the next person on the waiting list.
  5. Reduced irreversibility. If an event was written in error, write a subsequent deletion event to reverse it; downstream views incorporate it automatically and correct the data. In a database where you update and delete directly, a committed transaction is often difficult to reverse. This ties straight back to Ch 2's irreversibility is the main obstacle to evolvability.
  6. Audit log of what has occurred — valuable in regulated industries requiring auditability.
  7. Higher write throughput than databases, because of sequential access patterns. A temporary burst is absorbed by the log, and downstream view maintainers catch up at their own pace without being overwhelmed — natural backpressure.

3.3 Downsides (all three are real and commonly underestimated)

  1. External information breaks determinism. An event contains a price in one currency; a view needs it converted. Fetching the exchange rate from an external source at processing time is wrong — you'd get a different result if you recomputed the view on another date. To keep processing deterministic you must either include the exchange rate in the event itself, or have a way of querying the historical rate at the event's timestamp that always returns the same result for the same timestamp.
  2. Immutability vs GDPR. Users may request deletion of their data. If the log is per-user you can delete that user's whole log — but that doesn't work if the log contains events relating to multiple users. Options: store personal data outside the event, or encrypt it with a key you can later delete (crypto-shredding) — but both make it harder to recompute derived state when needed.
  3. Reprocessing with externally visible side effects. You probably don't want to resend confirmation emails every time you rebuild a materialized view.

3.4 Implementation

Implementable on any database. Purpose-built: EventStoreDB, MartenDB (on PostgreSQL), Axon Framework. You can also use Apache Kafka to store the event log with stream processors keeping views up to date (Ch 12).

The only important requirement: the event storage system must guarantee that ALL materialized views process the events in EXACTLY the same order as they appear in the log. As Ch 10 shows, this is not always easy to achieve in a distributed system.


On this page