5. Encoding and Evolution
5.0 The two compatibilities — memorize this
| Direction | Definition | Difficulty |
|---|---|---|
| Backward compatibility | Newer code can read data written by OLDER code | Normally not hard. As the author of the newer code, you know the format written by older code, so you can explicitly handle it (worst case, keep the old code around to read old data) |
| Forward compatibility | OLDER code can read data written by NEWER code | Trickier — it requires older code to IGNORE additions made by a newer version |
For APIs, work it out from who is newer:
| Scenario | Requirement |
|---|---|
| Older client → newer service | Backward compat on the REQUEST, forward compat on the RESPONSE |
| Newer client → older service | Forward compat on the REQUEST, backward compat on the RESPONSE |
The forward-compatibility trap: unknown-field loss
- New code writes a record with a new field:
{ name: "Martin", photoURL: "…", age: 30 } - Old code reads it. It decodes into a model object that has only
{ name, photoURL }— it never heard ofage. - Old code updates and writes it back:
{ name: "Martin", photoURL: "…" }—ageis silently lost.
The desirable behaviour is for old code to keep the new field intact even though it cannot interpret it. That requires the decoder to preserve unknown fields.
This is the single most under-appreciated bug class in schema evolution. It doesn't error, it doesn't log — it just deletes data on a round-trip through an old node during a rolling upgrade.
How the data models differ on this to begin with:
- Relational: all data conforms to one schema; it can be changed via migrations (
ALTER), but exactly one schema is in force at any point in time. - Schema-on-read ("schemaless"): no enforcement, so the database can contain a mixture of older and newer data formats written at different times.