1. Trade-Offs in Data Systems Architecture
1.6 Cross-cutting production failure catalog for this chapter
| Failure | Root trade-off it comes from |
|---|---|
| Analytics query tanks production latency | Ran OLAP on the OLTP system — reason #3 for warehouses |
| Dashboards show stale data, nobody notices | Derived data with no freshness monitoring |
| "The numbers don't match between two dashboards" | Two derived paths from one system of record, no single definition |
| Vendor raises prices 4× / sunsets the product | Cloud lock-in with no compatible alternative API |
| Cloud service is slow and you can't tell why | No access to OS metrics, server logs, or internals |
| Application is intermittently slow, disks look fine | Virtual block device — every I/O is a network call |
| One tenant's heavy job degrades everyone | Multitenancy without proper resource isolation |
| Retry causes duplicate side effects | Timeout gives no information about whether the request was received |
| A cluster is slower than one big machine | Distributed by default; data movement cost exceeded parallelism gain |
| Deploy breaks 6 downstream clients | Microservice API evolution without schema management |
| Cannot honour a GDPR erasure request | Immutable logs + untracked derived copies |
| Surprise $40k cloud bill | Capacity planning became financial planning and nobody owned it |
1.5 Technology deep dives
For each: what problem, why wasn't the alternative enough, how it works internally, deployment, monitoring, scaling, backup, what breaks in production.
1.7 Decision cheat sheet
Yes if any of: analysts need to join across ≥2 operational systems; analytical queries would compete with user traffic; analysts need ad-hoc SQL; or dataset is heading past a few…