5.11 Worked examples
JSON 81 B → MessagePack 66 B (−19%) → Protobuf 33 B (−59%) → Avro 32 B (−60%).
① Size comparison, same record. JSON 81 B → MessagePack 66 B (−19%) → Protobuf 33 B (−59%) → Avro 32 B (−60%). The step from JSON to a binary-JSON variant buys almost nothing; the step to a schema-driven format halves it, because field names leave the payload. At 1 billion records/day, that's ~49 GB/day saved on the wire versus JSON.
② Varint cost. favorite_number = 1337 → 2 bytes. favorite_number = 1_000_000_000_000 → 6 bytes. favorite_number = -1 in a plain int64 field → 10 bytes, because negative numbers set the high bits. This is why protobuf has sint64 (zigzag encoding), where −1 costs 1 byte. Choosing int64 for a field that holds negative deltas is a real, measurable waste.
③ Compatibility matrix for one change. You add string email = 4; to Person.
| Old code reads | New code reads | |
|---|---|---|
| Old data (no field 4) | fine | email = "" (default) — can you distinguish "no email" from "empty email"? No. |
| New data (has field 4) | skips tag 4 by length, preserves it ✔ | fine |
④ The same change in Avro. Add email with "default": "" → both directions fine. Add it without a default → new reader + old data = failure, because there's nothing to fill in. This asymmetry is the whole Avro rule in one line.
⑤ Rolling upgrade window. 100 nodes, 2-minute health check per node, deployed 10 at a time = ~20 minutes of mixed versions. During those 20 minutes, every write from a new node may be read by an old node. Now consider a mobile client: the equivalent window is months to years. That's the real reason forward compatibility isn't optional.
⑥ Durable execution replay cost. A workflow with 50 activities that fails at activity 49 replays all 49 prior activities from history on the retry — cheap, because they're history reads, not RPCs. But if the history has 20,000 events (a loop), replay itself becomes the bottleneck, and you need continueAsNew.