MVCC
xmin/xmax, snapshots, and visibility rules
Postgres never blocks a reader to let a writer finish, and never blocks a writer to let a reader finish. That guarantee — high concurrency without constant locking — comes from MVCC (Multi-Version Concurrency Control), and it's the single idea that explains the most about how Postgres actually behaves under load.
UPDATE doesn't update
Recall from Heap Storage that every tuple carries
an xmin (creating transaction) and xmax (deleting transaction) in its
header. UPDATE uses both: it marks the old tuple's xmax with the current
transaction ID and inserts a brand-new tuple with a fresh xmin — the old
row isn't touched or overwritten, it's just marked dead.
DELETE is the same idea with no new tuple: it just sets xmax. The old
tuple version keeps occupying its page until
VACUUM comes along and reclaims it — which is why a
table with heavy update traffic grows on disk even though its row count
stays flat.
This is exactly why UPDATE-heavy tables need routine vacuuming: every
update leaves a dead tuple behind. Skip vacuum for long enough and the
table can bloat to many times its logical size.
Snapshots and visibility
Every transaction gets a snapshot the moment it starts (or, under Read Committed — see Transactions & Isolation — the moment each statement starts): a record of which transaction IDs were already committed, in-progress, or not yet started at that instant.
A tuple is visible to your transaction only if:
- its
xmincommitted before your snapshot was taken, and - its
xmaxis either unset, or belongs to a transaction that hadn't committed by your snapshot
A transaction with a snapshot at txn 8 sees only v1 — the version whose xmin already committed and whose xmax hadn't (yet).
-- Session A -- Session B
BEGIN;
SELECT balance FROM accounts
WHERE id = 1; -- sees 100
BEGIN;
UPDATE accounts SET balance = 150
WHERE id = 1;
COMMIT;
SELECT balance FROM accounts
WHERE id = 1; -- still sees 100,
-- same snapshot as before
COMMIT;Session A keeps seeing the pre-update value for its whole transaction (under Repeatable Read or Serializable) because its snapshot was taken before B committed — not because anything was locked. B's write never had to wait for A's read, and A's read never had to wait for B's write.
This is the actual mechanism behind "readers don't block writers, writers don't block readers." There's no read lock being avoided — there's simply more than one physical version of the row on disk at once, and each transaction is handed the version consistent with its own snapshot.
Multiple versions, one table
At any moment a busy table can have several live versions of the "same"
logical row, each visible to a different set of in-flight transactions.
Postgres — and specifically VACUUM — is responsible for eventually
deciding a version is visible to nobody and reclaiming it.
A transaction ID is a 32-bit counter that wraps around. If a table went
unvacuumed forever, old xmin values would eventually look like they're
"in the future" again as the counter wraps — Postgres prevents this by
forcing a freeze (rewriting old xmins to a special frozen marker) well
before that can happen.
MVCC explains what happens when two transactions touch the same row at the same time. What happens when they lock the same row on purpose is next, in Transactions & Isolation.