Learn Labs
Storage Engine

Compression

Codecs & column encoding

Columnar storage doesn't just mean "read fewer columns" — it also compresses dramatically better than row storage, because every value sitting next to another value on disk is now the same type of thing: a column of country codes next to more country codes, a column of timestamps next to more timestamps. Similar values compress far better than a shuffled row of unrelated types ever could.

Two layers of compression, per column

Each column's data passes through an optional specialized encoding, then a general-purpose compressor. Both are configurable per column.

Raw columne.g. timestamps
CodecDelta / DoubleDelta / Gorilla
Compressed blockwritten to the part
CompressorLZ4 (default) / ZSTD

Specialized codecs exploit structure the compressor can't see

  • Delta — stores the difference between consecutive values. Great for slowly-increasing IDs or timestamps, where the deltas are much smaller numbers than the values themselves.
  • DoubleDelta — deltas of deltas. Even better for near-constant-interval timestamps.
  • Gorilla — designed for floating-point time-series (metrics) where consecutive values are close together.
  • T64 — transposes bits of fixed-width integers to expose more redundancy before general compression.

General compressors trade speed for ratio

  • LZ4 (default) — very fast to decompress, which matters more than raw ratio for most analytical queries that decompress a column, use it, and move on.
  • ZSTD — noticeably better compression ratio, more CPU per read. Common choice for cold/rarely-queried data or when storage cost dominates.
CREATE TABLE metrics
(
  ts     DateTime CODEC(DoubleDelta, ZSTD),
  value  Float64  CODEC(Gorilla, ZSTD),
  host   LowCardinality(String) CODEC(ZSTD)
)
ENGINE = MergeTree
ORDER BY ts;

On this page