Write-Ahead Log
Checkpoints, crash recovery, archiving
Every transaction that commits has to
survive a crash the instant after it commits — a power failure one
millisecond after COMMIT returns must not be able to lose that data.
Postgres doesn't guarantee this by forcing every changed data page to
disk before acknowledging a commit; that would make every write pay
for a random-I/O flush of scattered 8KB pages. Instead it writes
everything to one append-only log first, and that log is what actually
has to hit disk before a commit is considered done.
Write-ahead, literally
The rule the WAL (Write-Ahead Log) is named after: a change is described in the log before the data page it modifies is allowed to reach disk. The log entry itself is small and sequential — cheap to fsync — while the actual data page can be flushed later, lazily, in whatever order is convenient.
The client only waits on the top path — buffer, fsync, ack. The heap page carrying the actual row data can sit dirty in shared buffers for a while after the commit already returned; it gets written out later by the background writer or the next checkpoint, whichever comes first. This is exactly what makes commits fast: one small sequential write instead of N scattered ones, per transaction.
Why this makes crash recovery possible
If Postgres crashes before a dirty heap page reaches disk, the change
still exists — durably — in the WAL. On restart, Postgres replays every
WAL record since the last checkpoint against the on-disk pages,
reconstructing exactly the state that existed at the moment of the
crash. This is why a postgres container that gets kill -9'd
mid-write comes back with all committed data intact and nothing else:
recovery isn't guesswork, it's deterministic replay of a durable log.
This same mechanism is why a ROLLBACKed transaction's changes
never corrupt anything: its WAL records get written too (Postgres
doesn't know in advance a transaction will abort), but each record
carries the transaction ID that wrote it, and the visibility rules
from MVCC simply treat that transaction's tuples as
never having happened. Replay restores physical state; MVCC decides
what's visible.
Checkpoints
Replaying WAL back to the very beginning of time on every restart would make recovery take longer the older the database gets. A checkpoint periodically flushes all currently-dirty buffers to their data files and records the WAL position at that moment. Recovery after a crash only ever needs to replay WAL from the most recent checkpoint forward, not from the dawn of the cluster.
-- force one manually (mostly useful before a physical backup)
CHECKPOINT;
-- see what's currently configured
SHOW checkpoint_timeout; -- default: 5min
SHOW max_wal_size; -- checkpoints also trigger by WAL volumeCheckpoints are a tradeoff: too infrequent and recovery after a crash takes longer (more WAL to replay); too frequent and you pay repeated write I/O flushing pages that would otherwise have been overwritten again before ever being flushed.
WAL segments and pg_wal/
WAL isn't one infinite file — it's split into fixed-size segments,
16MB each by default, stored in the pg_wal/ directory inside the
data directory. Postgres writes into the current segment sequentially
and rolls to a new one when it fills up.
docker exec -it postgres psql -U admin -d learning -c \
"SELECT pg_current_wal_lsn(), pg_walfile_name(pg_current_wal_lsn());"
docker exec -it postgres ls -la /var/lib/postgresql/data/pg_wal/Segments no longer needed for crash recovery (older than the latest checkpoint) are normally recycled — renamed and reused for future WAL, rather than deleted and recreated — unless something is configured to keep them around longer, which is exactly what archiving does.
Archiving: the backbone of point-in-time recovery
By default, an old WAL segment is just recycled once it's no longer
needed for crash recovery — its contents are gone. Turning on
archive_mode runs archive_command against every completed segment
before it's allowed to be recycled, copying it somewhere durable (S3,
another disk, a dedicated archive server) instead of letting it be
overwritten.
wal_level = replica # or logical, if using logical replication
archive_mode = on
archive_command = 'cp %p /mnt/wal-archive/%f'A continuous archive of every WAL segment, combined with one physical base backup taken at some point in the past, is what lets Postgres replay forward to any moment since that backup — not just the moments a full backup happened to be taken. That combination is the entire mechanism behind Point-in-Time Recovery, and it's also the same stream of WAL records that replication ships to standby servers in real time instead of archiving to a file.
If archive_mode is on but archive_command starts failing
silently (disk full, permissions, network), Postgres keeps every WAL
segment since the last successful archive — it can't recycle them,
because losing an unarchived segment would break the archive's
continuity. pg_wal/ grows without bound until the disk fills up,
which is itself an outage. Monitor archiving success, not just that
the server is up.
Next: what happens to the old row versions that WAL-logged updates leave behind, in VACUUM & Autovacuum.