Learn Labs
Replication & Scaling

Replication

Streaming, physical, logical, and replication slots

Every write to Postgres already produces a byte-for-byte record of what changed: the write-ahead log. Replication is what happens when you ship that log to another server and have it replay the same changes. Nothing about replication requires a second "replication engine" — it's WAL, sent somewhere else, applied in order.

Physical (streaming) replication

A replica connects to the primary and streams WAL as it's generated, applying each record to its own copy of the data files. Because it's replaying the exact same WAL the primary wrote, the replica ends up byte-for-byte identical — same tables, same indexes, same bloat, same everything.

Primary
Replica 1streams + replays WAL
Replica 2streams + replays WAL
wal_level = replica          # or 'logical' — see below
max_wal_senders = 10         # max concurrent replication connections
wal_keep_size = 1GB          # WAL retained for replicas that fall behind

A replica is read-only by default — it will reject writes with ERROR: cannot execute INSERT in a read-only transaction. This is what makes read replicas useful: point read-heavy traffic at them without any risk of them diverging from the primary.

Synchronous vs asynchronous

By default, replication is asynchronous — the primary commits a transaction and returns success to the client without waiting for any replica to receive it. This means a crashed primary can lose the last few transactions that never made it to a replica.

Synchronous replication (synchronous_standby_names) makes the primary wait for at least one replica to confirm it has received the WAL before the client's COMMIT returns. This trades latency (every commit now waits on a network round-trip) for a durability guarantee: an acknowledged commit is not lost even if the primary dies immediately after.

Replication slots

Without a slot, a lagging replica is the primary's problem to solve on its own — it will happily recycle old WAL segments once wal_keep_size is exceeded, and a replica that falls behind that point can no longer catch up; it has to be rebuilt from scratch.

A replication slot flips that: the primary tracks exactly how far each slot's consumer has confirmed receiving WAL, and refuses to remove any segment a slot still needs — no matter how far behind it falls.

SELECT pg_create_physical_replication_slot('replica_1');
SELECT slot_name, active, restart_lsn FROM pg_replication_slots;

Logical replication

Physical replication ships raw bytes and requires an identical replica — same major version, same full copy of every database in the cluster. Logical replication instead decodes WAL back into row-level changes (INSERT/UPDATE/DELETE on specific tables) and replays those as SQL, which unlocks things physical replication can't do:

  • Replicate a subset of tables, not the whole cluster.
  • Replicate into a database with extra tables, columns, or a different major Postgres version — useful for near-zero-downtime upgrades.
  • Feed changes into non-Postgres consumers (this is what tools like Debezium build on).
-- on the publisher (source)
CREATE PUBLICATION orders_pub FOR TABLE orders, order_items;

-- on the subscriber (destination)
CREATE SUBSCRIPTION orders_sub
CONNECTION 'host=primary dbname=learning user=admin password=admin123'
PUBLICATION orders_pub;

Logical replication needs wal_level = logical (a superset of replica) because it has to decode enough information from WAL to reconstruct each row's before/after values, not just apply raw page changes.

This repo's local postgres/compose.yaml runs a single instance, so there's no replica to point at — but everything above works the same whether the replica is a second container on your laptop or a separate machine in another region. The concepts, not the container count, are what matter here. Next: what to do when a single primary and its replicas aren't enough — Scaling & Pooling.

On this page