5.3 JSON, XML, CSV, and binary variants
Widely adopted: in OpenAPI specs, in schema registries (Confluent Schema Registry, Red Hat Apicurio), and in databases (PostgreSQL's pg_jsonschema, MongoDB's $jsonSchema validator…
3.1 The honest list of flaws
| Format | Problems |
|---|---|
| XML | Too verbose and unnecessarily complicated |
| XML & CSV | Cannot distinguish a number from a string that happens to consist of digits (except via an external schema) |
| JSON | Distinguishes strings and numbers, but not integers from floating-point, and doesn't specify precision |
| CSV | No schema at all — the application defines what each row and column means; adding a row or column must be handled manually. Also quite vague (what if a value contains a comma or newline?). Escaping rules are formally specified but not all parsers implement them correctly |
| JSON & XML | Good Unicode string support, but no binary strings. Workaround: Base64, indicated by the schema — hacky, and increases data size by about a third |
| XML Schema & JSON Schema | Powerful and thus quite complicated to learn and implement. Since correct interpretation depends on schema info, applications not using schemas must hardcode encoding/decoding logic |
3.2 The 2⁵³ problem — worth knowing exactly
Integers greater than 2⁵³ cannot be exactly represented in an IEEE 754 double-precision float, so they become inaccurate when parsed by a language that uses floating point — like JavaScript.
The canonical real-world instance: X uses a 64-bit number to identify each post. The JSON returned by the API includes post IDs TWICE — once as a JSON number and once as a decimal string — to work around incorrect parsing by JavaScript applications.
Real ID: 1234567890123456789 (fits in int64)
Parsed in JS: 1234567890123456800 ← silently wrong, no error
2^53 = 9007199254740992 ← the exact boundaryThe API's fix is to send both: { "id": 1234567890123456789, "id_str": "1234567890123456789" } — and clients should use the string.
Practical rule: any ID that could exceed 2⁵³ (Snowflake IDs, bigint PKs, nanosecond timestamps) must be transported as a string in JSON.
3.3 JSON Schema
Widely adopted: in OpenAPI specs, in schema registries (Confluent Schema Registry, Red Hat Apicurio), and in databases (PostgreSQL's pg_jsonschema, MongoDB's $jsonSchema validator).
Offers standard primitives (string, number, integer, object, array, boolean, null) plus a separate validation specification overlaying constraints — e.g. a port field with minimum 1 and maximum 65,535.
Open vs closed content models:
| Model | additionalProperties | Meaning |
|---|---|---|
| Open (the default) | true | Any field not defined in the schema may exist, with any datatype |
| Closed | false | Only explicitly defined fields are allowed |
Consequence: JSON Schemas are usually a definition of what ISN'T permitted (invalid values on defined fields) rather than what IS permitted.
Example — a map from integer-like keys to strings, which JSON can't express natively (JSON objects always use string keys):
{
"$schema": "http://json-schema.org/draft-07/schema#",
"type": "object",
"patternProperties": { "^[0-9]+$": { "type": "string" } },
"additionalProperties": false
}JSON Schema also supports conditional if/else logic, named types, references to remote schemas, and more. All of which makes for a very powerful schema language — and unwieldy definitions. It can be challenging to resolve remote schemas, reason about conditional rules, or evolve schemas compatibly. Same concerns apply to XML Schema.
3.4 Binary variants of JSON/XML
A profusion: MessagePack, CBOR, BSON, BJSON, UBJSON, BISON, Hessian, Smile (JSON); WBXML, Fast Infoset (XML). Adopted in various niches — more compact, sometimes faster to parse — but none as widely adopted as the textual versions.
Some extend the datatype set (integers vs floats, binary strings) but keep the JSON/XML data model unchanged. Crucially:
Since they don't prescribe a schema, they must include ALL the object field names within the encoded data.
The running example record:
{ "userName": "Martin", "favoriteNumber": 1337, "interests": ["daydreaming", "hacking"] }MessagePack encoding, byte by byte:
83 object, 3 fields (0x80 = object | 0x03 = 3 fields)
a8 userName string, 8 bytes (0xa0 = string | 0x08 = length)
a6 Martin string, 6 bytes
ae favoriteNumber string, 14 bytes
cd 05 39 uint16 = 1337
a9 interests string, 9 bytes
92 array, 2 elements
ab daydreaming string, 11 bytes
a7 hacking string, 7 bytes
────────────────
TOTAL: 66 bytesBecause the length is known up front there is no terminator and no escaping — which is what makes a binary encoding both smaller and faster to parse than JSON.
(If an object has more than 15 fields — too many for four bits — it gets a different type indicator and the count is encoded in two or four bytes.)
The verdict:
| Encoding | Size |
|---|---|
| JSON (whitespace removed) | 81 bytes |
| MessagePack | 66 bytes |
| Protocol Buffers | 33 bytes |
| Avro | 32 bytes |
66 vs 81 bytes is only a little less. It's not clear whether such a small space reduction (and perhaps a parsing speedup) is worth the loss of human-readability. The schema-driven formats do twice as well — because they drop the field names entirely.