2.5 Maintainability
Software doesn't wear out or suffer material fatigue.
Software doesn't wear out or suffer material fatigue. But requirements evolve, the environment changes (dependencies, platform), and bugs need fixing.
The majority of the cost of software is not initial development but ongoing maintenance — fixing bugs, keeping systems operational, investigating failures, adapting to new platforms, modifying for new use cases, repaying technical debt, adding features.
Legacy pain compounds: outdated technologies few engineers understand (mainframes, COBOL); institutional knowledge of how and why the system was designed lost as people leave; fixing other people's mistakes. Because systems are intertwined with the human organizations they support, maintenance is as much a people problem as a technical one.
Every system we create today will one day become a legacy system, if it is valuable enough to survive.
Three principles:
| Principle | Goal |
|---|---|
| Operability | Make it easy for the organization to keep the system running smoothly |
| Simplicity | Make it easy for new engineers to understand — well-understood, consistent patterns; no unnecessary complexity |
| Evolvability | Make it easy to change the system later, for unanticipated use cases |
5.1 Operability
"Good operations can often work around the limitations of bad (or incomplete) software, but good software cannot run reliably with bad operations."
Automation is essential at thousands of machines — manual maintenance would be unreasonably expensive. But automation is two-edged:
- There will always be edge cases (rare failure scenarios) requiring manual intervention, and the cases that can't be automated tend to be the most complex — so greater automation requires a MORE skilled operations team
- An automated system that goes wrong is often harder to troubleshoot than one where an operator does some steps manually
- Therefore more automation is not always better for operability. The sweet spot depends on your application and organization.
What data systems can do to be operable:
- Support monitoring of key metrics and observability tools for runtime behavior
- Avoid dependency on individual machines — let machines be taken down for maintenance while the system keeps running
- Good documentation and an easy-to-understand operational model: "If I do X, Y will happen"
- Good defaults, but freedom to override them
- Self-healing where appropriate, but manual control over system state when needed
- Predictable behavior, minimizing surprises
5.2 Simplicity
Complexity slows everyone down and raises maintenance cost. A project mired in it is a big ball of mud. In complex software there is greater risk of introducing bugs when making a change, because hidden assumptions, unintended consequences, and unexpected interactions are more easily overlooked.
Simplicity is subjective — there is no objective standard. The book's own counterexample: is a system that hides a complex implementation behind a simple interface simpler than one with a simple implementation that exposes more internal detail? No settled answer.
Essential vs accidental complexity (essential = inherent to the problem domain; accidental = arising only from limitations of our tooling) is a useful frame but also flawed, because the boundary shifts as tooling evolves.
Abstraction is the best tool we have. A good abstraction hides implementation detail behind a clean façade and can serve a wide range of applications. Beyond reuse efficiency, quality improvements in the abstracted component benefit every application that uses it.
Examples of abstraction:
- High-level languages hide machine code, CPU registers, and system calls
- SQL hides complex on-disk and in-memory data structures, concurrent requests from other clients, and inconsistencies after crashes
Application-level abstraction methodologies: design patterns, domain-driven design (DDD). This book is instead about general-purpose abstractions you build applications on: database transactions, indexes, and event logs. You can implement DDD on top of these foundations.
5.3 Evolvability
Requirements are in constant flux: new facts learned, unanticipated use cases, changed business priorities, new feature requests, platform replacement, legal/regulatory change, growth forcing architectural change.
Agile gives an organizational framework; TDD and refactoring are its technical tools. Evolvability is the word for agility at the data-system level (across several applications/services with different characteristics).
Evolvability is closely linked to simplicity and abstraction. Loosely coupled, simple systems are easier to modify than tightly coupled, complex ones.
The major factor making change difficult in large systems is IRREVERSIBILITY. Migrating from one database to another: if you cannot switch back when the new one has problems, the stakes are much higher. Irreversible actions must be taken very carefully. Minimizing irreversibility improves flexibility.
This is the design principle that most directly translates into daily practice: prefer dual-writes over cutovers, expand-then-contract migrations over rename-in-place, feature flags over branch-and-deploy, and shadow traffic over big-bang switches.