Learn Labs
WAL, Vacuum & Recovery

Write-Ahead Log

Checkpoints, crash recovery, archiving

Every transaction that commits has to survive a crash the instant after it commits — a power failure one millisecond after COMMIT returns must not be able to lose that data. Postgres doesn't guarantee this by forcing every changed data page to disk before acknowledging a commit; that would make every write pay for a random-I/O flush of scattered 8KB pages. Instead it writes everything to one append-only log first, and that log is what actually has to hit disk before a commit is considered done.

Write-ahead, literally

The rule the WAL (Write-Ahead Log) is named after: a change is described in the log before the data page it modifies is allowed to reach disk. The log entry itself is small and sequential — cheap to fsync — while the actual data page can be flushed later, lazily, in whatever order is convenient.

Client COMMITwrite + commit
↓ appends
WAL Bufferin-memory, sequential
↓ fsync
WAL Segment on Diskpg_wal/, durable
Ack to clientcommit confirmed
Heap pageflushed later, lazily

The client only waits on the top path — buffer, fsync, ack. The heap page carrying the actual row data can sit dirty in shared buffers for a while after the commit already returned; it gets written out later by the background writer or the next checkpoint, whichever comes first. This is exactly what makes commits fast: one small sequential write instead of N scattered ones, per transaction.

Why this makes crash recovery possible

If Postgres crashes before a dirty heap page reaches disk, the change still exists — durably — in the WAL. On restart, Postgres replays every WAL record since the last checkpoint against the on-disk pages, reconstructing exactly the state that existed at the moment of the crash. This is why a postgres container that gets kill -9'd mid-write comes back with all committed data intact and nothing else: recovery isn't guesswork, it's deterministic replay of a durable log.

Checkpoints

Replaying WAL back to the very beginning of time on every restart would make recovery take longer the older the database gets. A checkpoint periodically flushes all currently-dirty buffers to their data files and records the WAL position at that moment. Recovery after a crash only ever needs to replay WAL from the most recent checkpoint forward, not from the dawn of the cluster.

-- force one manually (mostly useful before a physical backup)
CHECKPOINT;

-- see what's currently configured
SHOW checkpoint_timeout;      -- default: 5min
SHOW max_wal_size;            -- checkpoints also trigger by WAL volume

Checkpoints are a tradeoff: too infrequent and recovery after a crash takes longer (more WAL to replay); too frequent and you pay repeated write I/O flushing pages that would otherwise have been overwritten again before ever being flushed.

WAL segments and pg_wal/

WAL isn't one infinite file — it's split into fixed-size segments, 16MB each by default, stored in the pg_wal/ directory inside the data directory. Postgres writes into the current segment sequentially and rolls to a new one when it fills up.

docker exec -it postgres psql -U admin -d learning -c \
"SELECT pg_current_wal_lsn(), pg_walfile_name(pg_current_wal_lsn());"

docker exec -it postgres ls -la /var/lib/postgresql/data/pg_wal/

Segments no longer needed for crash recovery (older than the latest checkpoint) are normally recycled — renamed and reused for future WAL, rather than deleted and recreated — unless something is configured to keep them around longer, which is exactly what archiving does.

Archiving: the backbone of point-in-time recovery

By default, an old WAL segment is just recycled once it's no longer needed for crash recovery — its contents are gone. Turning on archive_mode runs archive_command against every completed segment before it's allowed to be recycled, copying it somewhere durable (S3, another disk, a dedicated archive server) instead of letting it be overwritten.

wal_level = replica          # or logical, if using logical replication
archive_mode = on
archive_command = 'cp %p /mnt/wal-archive/%f'

A continuous archive of every WAL segment, combined with one physical base backup taken at some point in the past, is what lets Postgres replay forward to any moment since that backup — not just the moments a full backup happened to be taken. That combination is the entire mechanism behind Point-in-Time Recovery, and it's also the same stream of WAL records that replication ships to standby servers in real time instead of archiving to a file.

Next: what happens to the old row versions that WAL-logged updates leave behind, in VACUUM & Autovacuum.

On this page