Learn Labs
5. Encoding and Evolution

5.3 JSON, XML, CSV, and binary variants

Widely adopted: in OpenAPI specs, in schema registries (Confluent Schema Registry, Red Hat Apicurio), and in databases (PostgreSQL's pg_jsonschema, MongoDB's $jsonSchema validator…

3.1 The honest list of flaws

FormatProblems
XMLToo verbose and unnecessarily complicated
XML & CSVCannot distinguish a number from a string that happens to consist of digits (except via an external schema)
JSONDistinguishes strings and numbers, but not integers from floating-point, and doesn't specify precision
CSVNo schema at all — the application defines what each row and column means; adding a row or column must be handled manually. Also quite vague (what if a value contains a comma or newline?). Escaping rules are formally specified but not all parsers implement them correctly
JSON & XMLGood Unicode string support, but no binary strings. Workaround: Base64, indicated by the schema — hacky, and increases data size by about a third
XML Schema & JSON SchemaPowerful and thus quite complicated to learn and implement. Since correct interpretation depends on schema info, applications not using schemas must hardcode encoding/decoding logic

3.2 The 2⁵³ problem — worth knowing exactly

Integers greater than 2⁵³ cannot be exactly represented in an IEEE 754 double-precision float, so they become inaccurate when parsed by a language that uses floating point — like JavaScript.

The canonical real-world instance: X uses a 64-bit number to identify each post. The JSON returned by the API includes post IDs TWICE — once as a JSON number and once as a decimal string — to work around incorrect parsing by JavaScript applications.

Real ID:       1234567890123456789   (fits in int64)
Parsed in JS:  1234567890123456800   ← silently wrong, no error
2^53 =            9007199254740992   ← the exact boundary

The API's fix is to send both: { "id": 1234567890123456789, "id_str": "1234567890123456789" } — and clients should use the string.

Practical rule: any ID that could exceed 2⁵³ (Snowflake IDs, bigint PKs, nanosecond timestamps) must be transported as a string in JSON.

3.3 JSON Schema

Widely adopted: in OpenAPI specs, in schema registries (Confluent Schema Registry, Red Hat Apicurio), and in databases (PostgreSQL's pg_jsonschema, MongoDB's $jsonSchema validator).

Offers standard primitives (string, number, integer, object, array, boolean, null) plus a separate validation specification overlaying constraints — e.g. a port field with minimum 1 and maximum 65,535.

Open vs closed content models:

ModeladditionalPropertiesMeaning
Open (the default)trueAny field not defined in the schema may exist, with any datatype
ClosedfalseOnly explicitly defined fields are allowed

Consequence: JSON Schemas are usually a definition of what ISN'T permitted (invalid values on defined fields) rather than what IS permitted.

Example — a map from integer-like keys to strings, which JSON can't express natively (JSON objects always use string keys):

{
  "$schema": "http://json-schema.org/draft-07/schema#",
  "type": "object",
  "patternProperties": { "^[0-9]+$": { "type": "string" } },
  "additionalProperties": false
}

JSON Schema also supports conditional if/else logic, named types, references to remote schemas, and more. All of which makes for a very powerful schema language — and unwieldy definitions. It can be challenging to resolve remote schemas, reason about conditional rules, or evolve schemas compatibly. Same concerns apply to XML Schema.

3.4 Binary variants of JSON/XML

A profusion: MessagePack, CBOR, BSON, BJSON, UBJSON, BISON, Hessian, Smile (JSON); WBXML, Fast Infoset (XML). Adopted in various niches — more compact, sometimes faster to parse — but none as widely adopted as the textual versions.

Some extend the datatype set (integers vs floats, binary strings) but keep the JSON/XML data model unchanged. Crucially:

Since they don't prescribe a schema, they must include ALL the object field names within the encoded data.

The running example record:

{ "userName": "Martin", "favoriteNumber": 1337, "interests": ["daydreaming", "hacking"] }

MessagePack encoding, byte by byte:

83                object, 3 fields   (0x80 = object | 0x03 = 3 fields)
a8 userName       string, 8 bytes    (0xa0 = string | 0x08 = length)
a6 Martin         string, 6 bytes
ae favoriteNumber string, 14 bytes
cd 05 39          uint16 = 1337
a9 interests      string, 9 bytes
92                array, 2 elements
ab daydreaming    string, 11 bytes
a7 hacking        string, 7 bytes
                  ────────────────
                  TOTAL: 66 bytes

Because the length is known up front there is no terminator and no escaping — which is what makes a binary encoding both smaller and faster to parse than JSON.

(If an object has more than 15 fields — too many for four bits — it gets a different type indicator and the count is encoded in two or four bytes.)

The verdict:

EncodingSize
JSON (whitespace removed)81 bytes
MessagePack66 bytes
Protocol Buffers33 bytes
Avro32 bytes

66 vs 81 bytes is only a little less. It's not clear whether such a small space reduction (and perhaps a parsing speedup) is worth the loss of human-readability. The schema-driven formats do twice as well — because they drop the field names entirely.


On this page