1.2 Cloud vs Self-Hosting
Two separable decisions: who builds the software and who deploys it.
2.1 The build/buy spectrum
Two separable decisions: who builds the software and who deploys it.
| Bespoke, in-house | Off-the-shelf, self-hosted | SaaS / cloud service |
|---|---|---|
| You write it and you run it. | You download MySQL and install it — on your own hardware (on premises), on a cloud VM (IaaS), or as a modified fork of the open source. | The vendor implements and operates it; you access it by web UI or API. |
| ← more control · less operational burden → | ||
Rule of thumb: core competency / competitive advantage → in-house. Non-core, routine, commonplace → vendor. (Most companies don't fabricate their own CPUs.)
2.2 Honest scorecard
Cloud wins when:
- You don't already know how to deploy and operate the system. Hiring/training specialists is expensive.
- Load varies a lot. If you provision for peak and sit idle most of the time, you're wasting money. Analytical systems are the archetype: a big query needs lots of parallel compute, then the resources sit idle until the next query. Predefined daily reports can be enqueued and smoothed; interactive queries are the variable ones — the faster you want them, the burstier the load.
- The vendor's operational expertise (gained across many customers) exceeds yours.
Self-hosting wins when:
- You already have the operational skill and load is predictable → often just cheaper to buy machines.
- You need to tune the system for your particular workload. A cloud service won't customize for you.
- You need full control of hardware — e.g. very latency-sensitive high-frequency trading.
The downsides of cloud, all of which reduce to "no control":
| Risk | What you can actually do about it |
|---|---|
| Missing feature | Politely ask the vendor. You cannot implement it. |
| Service goes down | Wait. |
| You trigger a bug or perf problem | Hard to diagnose — no OS metrics, no server logs, no internals |
| Vendor shuts down / raises price / changes product | Forced migration; running an old version isn't an option → vendor lock-in (mitigated only if alternatives expose a compatible API, and most cloud services have no standard API) |
| Geopolitics | Sanctions can lock you out of a provider in another country |
| Security & compliance | You must trust the provider with your data |
2.3 Cloud native architecture — the technical consequence
"Cloud native" = designed from the ground up to build on cloud services, not just to run on cloud VMs. Demonstrated advantages: better performance on the same hardware, faster recovery from failures, faster scaling to match load, larger datasets.
| Category | Self-hosted | Cloud native |
|---|---|---|
| Operational / OLTP | MySQL, PostgreSQL, MongoDB | AWS Aurora, Azure SQL DB Hyperscale, Google Cloud Spanner |
| Analytical / OLAP | Teradata, ClickHouse, Spark | Snowflake, Google BigQuery, Azure Synapse Analytics |
Layering. Self-hosted software assumes generic resources: CPUs, RAM, a filesystem, an IP network. Cloud native services instead build higher-level services on top of lower-level cloud services:
- Object storage (S3, Azure Blob, Cloudflare R2) — more limited API than a filesystem (basic reads/writes of large files), but it hides the underlying physical machines: automatically distributes data across many machines so you never run out of disk on one, and survives machine/disk failure with no data loss.
- Snowflake is a data warehouse built on S3. Other services in turn build on Snowflake.
General abstraction rule: higher-level abstractions are more oriented to particular use cases. If your needs match, use the high-level thing. If nothing fits, compose from lower-level parts.
2.4 Separation of storage and compute — the defining cloud-native move
Traditional model: disk is durable; RAID keeps copies across several disks on the same machine, transparent to applications.
Cloud model breaks that:
- Local instance disks are treated as an ephemeral cache, not long-term storage — the disk becomes inaccessible if the instance fails, or if the instance is resized (moving it to a different physical machine).
- Virtual disks (EBS, Azure managed disks, GCP persistent disks) can detach from one instance and attach to another. But a virtual disk is not a physical disk — it's a service run by a separate set of machines emulating a block device (typically 4 KiB blocks). Two costs: (a) block-device emulation overhead that a purpose-built cloud system avoids, and (b) every I/O is a network call, making the application very sensitive to network glitches.
- So cloud native services avoid virtual disks and build on dedicated storage services optimized per workload. Object stores are designed for large files (hundreds of KB → several GB). Individual DB rows are far smaller — so cloud databases manage small values in a separate service and pack larger blocks (containing many values) into the object store.
Multitenancy. Cloud native systems typically share hardware across customers rather than one machine per customer. Benefits: better hardware utilization, easier scaling, easier management. Cost: careful engineering so one customer's activity doesn't affect another's performance or security (the noisy-neighbour problem).
2.5 Operations in the cloud era
Roles: DBAs/sysadmins → DevOps → SRE (Google's implementation of the idea). The role of operations: deliver services reliably to users (configuring infra, deploying apps) and keep production stable (monitoring and diagnosing anything that affects reliability).
Traditional self-hosted ops work is machine-level: capacity planning (watch disk space, add disks before you run out), provisioning machines, moving services between machines, OS patching.
Cloud services present an API that hides individual machines — e.g. cloud storage replaces fixed-size disks with metered billing (store without planning capacity, get charged for space used), and stays available even when individual machines fail.
DevOps/SRE emphasis:
- Automation, repeatable processes over manual one-off jobs
- Ephemeral VMs and services rather than long-running servers
- Frequent application updates
- Learning from incidents
- Preserving organizational knowledge as people come and go
The bifurcation: ops teams at infrastructure companies specialize in running a reliable service for many customers; customers of those services spend as little time on infrastructure as possible.
But cloud customers still need operations, just different operations:
- Choosing the right service for a task; integrating services; migrating between services
- Capacity planning becomes financial planning; performance optimization becomes cost optimization — metered billing removes capacity planning but you must still know what resources you use and why, or you burn money
- Quotas and resource limits (e.g. max concurrent processes) must be known and planned for before you hit them
- Integration between services is a growing challenge and there are no standards — it's manual effort
- Cannot be outsourced at all: application/library security, interactions between your own services, monitoring load, and root-causing performance degradations and outages
The need for operations is as great as ever — the cloud changed its shape, not its existence.
1.1 Operational vs Analytical Systems
Key observation: analysts and scientists both read data that users and backend services generated, and they do not modify it (they may create derived datasets).
1.3 Distributed vs Single-Node Systems
A distributed system = several machines communicating over a network.