Learn Labs
5. Encoding and Evolution

5.4 Protocol Buffers

Binary encoding library from Google; similar to Apache Thrift (originally Facebook) — most of this section applies to Thrift too.

Encoding size

JSON · 216 B/record0.2 GB
MessagePack · 184 B/record0.2 GB
Protobuf · 112 B/record0.1 GB
Avro · 104 B/record0.1 GB
Safe

Avro is 2.1× smaller than JSON here, saving 0.1 GB at this volume. The trade is that the bytes are meaningless without the writer's schema, so you now need a schema registry and a compatibility policy — infrastructure JSON did not require.

Binary formats drop the field names because the schema already knows them. That is where the space goes — and it is also why you can no longer read the bytes without the schema.

Binary encoding library from Google; similar to Apache Thrift (originally Facebook) — most of this section applies to Thrift too.

Requires a schema, written in the interface definition language (IDL):

syntax = "proto3";
message Person {
    string user_name      = 1;
    int64  favorite_number = 2;
    repeated string interests = 3;
}

A code generation tool produces classes implementing the schema in various languages; your application calls the generated code to encode/decode.

The schema language is very simple compared to JSON Schema: it defines fields and their types, but does NOT support other restrictions on possible values.

4.1 The encoding — 33 bytes

Field tags replace field names entirely:

0a  06  M a r t i n     tag=1, type=string(2), len=6, "Martin"
▲▲
││└── wire type 2 (length-delimited)
│└─── field tag 1        ← type and tag packed into a single byte

10  b9 0a                tag=2, type=varint, value=1337

1a  0b  daydreaming      tag=3, string
1a  07  hacking          tag=3 AGAIN

A varint uses the top bit of each byte to say “more bytes follow”, with the least significant 7 bits in the first byte to simplify reconstructing the integer as bytes arrive: −64…63 in one byte, −8,192…8,191 in two, bigger numbers in more.

There is no explicit list datatype. repeated means the field holds a list, and in the binary encoding list elements are simply repeated occurrences of the same field tag within one record.

Compare with MessagePack: no userName, favoriteNumber, interests strings anywhere. Field tags are aliases for fields — a compact way of indicating which field we mean without spelling out the name.

4.2 Schema evolution rules

An encoded record is just the concatenation of its encoded fields. Each field is identified by tag number and annotated with a datatype. If a field value is not set, it is simply omitted.

ChangeSafe?Why
Change a field name✔The encoded data never refers to names.
Change a field tag✗Would make all existing encoded data invalid.
Add a field with a new tag✔Forward: old code sees an unrecognized tag and ignores it — the datatype annotation tells the parser how many bytes to skip while preserving the unknown field, which avoids the data-loss problem above. Backward: new code reading old data gets a default value.
Remove a field✔The same as adding, with the two compatibilities reversed.
Reuse a tag number✗Data written somewhere may still carry the old tag. Reserve past tag numbers in the schema so they are not forgotten.
Change a datatype~Possible for some types, but values may truncate. int32 → int64: new code reads old data fine, but old code reading new data still uses a 32-bit variable, so a 64-bit value that does not fit is truncated.

On this page