5.4 Protocol Buffers
Binary encoding library from Google; similar to Apache Thrift (originally Facebook) — most of this section applies to Thrift too.
Encoding size
Avro is 2.1× smaller than JSON here, saving 0.1 GB at this volume. The trade is that the bytes are meaningless without the writer's schema, so you now need a schema registry and a compatibility policy — infrastructure JSON did not require.
Binary encoding library from Google; similar to Apache Thrift (originally Facebook) — most of this section applies to Thrift too.
Requires a schema, written in the interface definition language (IDL):
syntax = "proto3";
message Person {
string user_name = 1;
int64 favorite_number = 2;
repeated string interests = 3;
}A code generation tool produces classes implementing the schema in various languages; your application calls the generated code to encode/decode.
The schema language is very simple compared to JSON Schema: it defines fields and their types, but does NOT support other restrictions on possible values.
4.1 The encoding — 33 bytes
Field tags replace field names entirely:
0a 06 M a r t i n tag=1, type=string(2), len=6, "Martin"
▲▲
││└── wire type 2 (length-delimited)
│└─── field tag 1 ← type and tag packed into a single byte
10 b9 0a tag=2, type=varint, value=1337
1a 0b daydreaming tag=3, string
1a 07 hacking tag=3 AGAINA varint uses the top bit of each byte to say “more bytes follow”, with the least significant 7 bits in the first byte to simplify reconstructing the integer as bytes arrive: −64…63 in one byte, −8,192…8,191 in two, bigger numbers in more.
There is no explicit list datatype. repeated means the field holds a list, and in the binary encoding list elements are simply repeated occurrences of the same field tag within one record.
Compare with MessagePack: no userName, favoriteNumber, interests strings anywhere. Field tags are aliases for fields — a compact way of indicating which field we mean without spelling out the name.
4.2 Schema evolution rules
An encoded record is just the concatenation of its encoded fields. Each field is identified by tag number and annotated with a datatype. If a field value is not set, it is simply omitted.
| Change | Safe? | Why |
|---|---|---|
| Change a field name | ✔ | The encoded data never refers to names. |
| Change a field tag | ✗ | Would make all existing encoded data invalid. |
| Add a field with a new tag | ✔ | Forward: old code sees an unrecognized tag and ignores it — the datatype annotation tells the parser how many bytes to skip while preserving the unknown field, which avoids the data-loss problem above. Backward: new code reading old data gets a default value. |
| Remove a field | ✔ | The same as adding, with the two compatibilities reversed. |
| Reuse a tag number | ✗ | Data written somewhere may still carry the old tag. Reserve past tag numbers in the schema so they are not forgotten. |
| Change a datatype | ~ | Possible for some types, but values may truncate. int32 → int64: new code reads old data fine, but old code reading new data still uses a 32-bit variable, so a 64-bit value that does not fit is truncated. |
5.3 JSON, XML, CSV, and binary variants
Widely adopted: in OpenAPI specs, in schema registries (Confluent Schema Registry, Red Hat Apicurio), and in databases (PostgreSQL's pg_jsonschema, MongoDB's $jsonSchema validator…
5.5 Avro
Started 2009 as a Hadoop subproject, because Protocol Buffers was not a good fit for Hadoop's use cases.