Learn Labs
5. Encoding and Evolution

5.11 Worked examples

JSON 81 B → MessagePack 66 B (−19%) → Protobuf 33 B (−59%) → Avro 32 B (−60%).

① Size comparison, same record. JSON 81 B → MessagePack 66 B (−19%) → Protobuf 33 B (−59%) → Avro 32 B (−60%). The step from JSON to a binary-JSON variant buys almost nothing; the step to a schema-driven format halves it, because field names leave the payload. At 1 billion records/day, that's ~49 GB/day saved on the wire versus JSON.

② Varint cost. favorite_number = 1337 → 2 bytes. favorite_number = 1_000_000_000_000 → 6 bytes. favorite_number = -1 in a plain int64 field → 10 bytes, because negative numbers set the high bits. This is why protobuf has sint64 (zigzag encoding), where −1 costs 1 byte. Choosing int64 for a field that holds negative deltas is a real, measurable waste.

③ Compatibility matrix for one change. You add string email = 4; to Person.

Old code readsNew code reads
Old data (no field 4)fineemail = "" (default) — can you distinguish "no email" from "empty email"? No.
New data (has field 4)skips tag 4 by length, preserves it ✔fine

④ The same change in Avro. Add email with "default": "" → both directions fine. Add it without a default → new reader + old data = failure, because there's nothing to fill in. This asymmetry is the whole Avro rule in one line.

⑤ Rolling upgrade window. 100 nodes, 2-minute health check per node, deployed 10 at a time = ~20 minutes of mixed versions. During those 20 minutes, every write from a new node may be read by an old node. Now consider a mobile client: the equivalent window is months to years. That's the real reason forward compatibility isn't optional.

⑥ Durable execution replay cost. A workflow with 50 activities that fails at activity 49 replays all 49 prior activities from history on the retry — cheap, because they're history reads, not RPCs. But if the history has 20,000 events (a loop), replay itself becomes the bottleneck, and you need continueAsNew.