Duck standing on a brick irrigation channel in a flat agricultural field with rows of green crops stretching to the horizon.

min_wal_size has a better description in pg_settings than it has in the documentation: “Sets the minimum size to shrink the WAL to.” That is all it does. It is a floor under how far pg_wal is allowed to shrink, and it keeps no WAL that anything could read, though the name invites reading it as a retention setting. Here is a primary whose standby was stopped 640 MB ago:

1=# SHOW min_wal_size;
2 1GB
3=# SELECT pg_walfile_name(pg_current_wal_lsn());
4 0000000100000002000000F3
5
6$ du -sh pg_wal
71.2G pg_wal
8$ ls pg_wal | grep -c ^0000
976
10$ ls pg_wal | head -1
110000000100000002000000F3

And the standby, when it came back:

1FATAL: could not receive data from WAL stream: ERROR: requested WAL segment 0000000100000002000000CB has already been removed

There are 76 files in that directory and the oldest is the segment being written right now. The other 75 are old segments renamed to names the server has not reached yet, full of stale bytes waiting to be overwritten. If you want WAL kept for a replica, that is wal_keep_size or a slot.

The default is 80MB, five 16 MB segments (initdb writes the line into postgresql.conf itself and scales it with the segment size, so --wal-segsize=64 gets 320MB). The context is sighup, the unit is megabytes, and it arrived in 9.5 alongside max_wal_size when Heikki Linnakangas replaced checkpoint_segments.

The floor under an autotuner

The max_wal_size post covered recycling: at the end of a checkpoint, segments older than the new redo pointer are either renamed into future segments or unlinked. How many get renamed is decided by the estimate figure in every log_checkpoints line, a moving average of WAL written per checkpoint cycle that jumps up at once when a cycle exceeds it and comes down a tenth of the way per checkpoint when one doesn’t. The server keeps enough spare files to cover 1 + checkpoint_completion_target cycles at that rate, plus 10%. max_wal_size is the ceiling on that number and min_wal_size is the floor; Heikki’s commit message notes that setting the two equal switches the autotuning off.

Two properties of the floor are not in the name. First, nothing ever creates files to reach it. A fresh cluster with min_wal_size = 4GB has one 16 MB file in pg_wal, and so does a new replica from pg_basebackup, which does not copy the pool. A promoted standby recycles at most ten segments onto the new timeline and unlinks the rest (mine went from 26 spare segments to 10), so the server that has just taken over production starts with 160 MB of pool whatever the setting says. The pool only grows by recycling WAL that has already been written.

Second, the shrinking is lazy. Checkpoints never revisit segments ahead of the insert point, so the pool comes down one file per segment of WAL written through it. After a 1.5 GB burst my pg_wal held 1,520 MB; five checkpoints with almost nothing written between them took the estimate from 1.5 GB to 0.9 GB and removed nothing; it then took about 1.5 GB of trickle, 16 MB per checkpoint, to reach five files. For the floor to matter at all, a server has to write a pool’s worth of WAL slowly, over enough checkpoints for the estimate to fall (it halves about every seven), and then get busy again. That is a nightly batch. A server with a steady write rate has a pool sized to that rate, and the floor does nothing unless it is higher.

Twenty-one milliseconds, 53 times

What the floor buys is not creating files. When WAL runs off the end of the pool, the process that needs the next segment (nearly always a backend) makes it: 16 MB of zeroes, an fsync, a rename, all while holding WALWriteLock, so every other commit queues behind it. I ran a 3 GB bulk write, trickled 16 MB per checkpoint through the server until the pool had drained to its floor, then started pgbench (scale 50, eight clients, two minutes, 18.6):

1min_wal_size = 80MB pool 5 segments 3,396 tps p50 2.1 ms p99.9 22.8 ms 53 segments created, 1,103 ms
2min_wal_size = 8GB pool 227 segments 3,439 tps p50 2.1 ms p99.9 11.4 ms 0 segments created

A second pair came out the same (22.6 ms against 11.2). Zero-filling and syncing each new segment took 21 ms, which does not count the rename’s own fsync calls, and sampling pg_stat_activity in a separate run found that whenever one backend was creating a segment, six of the other seven clients were waiting on LWLock:WALWrite. The median and p99 did not move, and throughput differed by 1 to 2%, which is about the noise between runs. The p99.9 doubled. That is on a disk with a 0.3 ms fsync. On a network volume throttled to 125 MB/s, the zeroes alone take 128 ms.

In 18 the place to look is pg_stat_io: the rows with object = 'wal' and context = 'init' count segment creations by backend type, with times if track_wal_io_timing is on. Before 18 the wait events WALInitWrite and WALInitSync (spelled WalInit… from 17) are about all there is. log_checkpoints will not tell you. Its “WAL file(s) added” counts only the segment the checkpointer sometimes preallocates for itself, and the checkpoint after a burst in which other processes created 45 segments reported 0.

The documentation’s other argument is that the space is “reserved,” and it is, since recycled segments are allocated disk. I filled a volume to 100% and kept writing. With the floor at 80MB the server wrote 78 MB more WAL before PANIC: could not write to file "pg_wal/xlogtemp.15129": No space left on device; at 512MB, 506 MB. Both then failed crash recovery on the same message and stayed down. (With checkpoints running it is less than that: the next checkpoint has to create a small file of its own and panics when it can’t.) At the write rate of the test above, the difference is ten seconds against a minute.

The check that runs too late

pg_settings gives the range as 2 to 2147483647, and the first number is wrong in practice. The real minimum is two segments, and it is checked only when the postmaster reads the control file:

1=# ALTER SYSTEM SET min_wal_size = '16MB';
2=# SELECT pg_reload_conf();
3=# SHOW min_wal_size;
4 16MB

The server runs on that until the next restart, or until the next time any backend crashes and the postmaster reinitializes, whichever comes first:

1LOG: all server processes terminated; reinitializing
2FATAL: "min_wal_size" must be at least twice "wal_segment_size"
3LOG: database system is shut down

The second case turns one OOM kill into an outage that lasts until someone edits a file. Bharath Rupireddy proposed enforcing the limit at SET time in 2022, Tom Lane pointed out that a check hook cannot validate one parameter against another, and there it has stayed; 19 beta 4 behaves identically. With 64 MB segments the real minimum is 128MB, which is above the compiled-in default of 80MB, and that is why initdb writes the line: comment it out on that cluster and it will not start.

Two other settings fail quietly instead. wal_recycle = off (the copy-on-write filesystem setting) makes min_wal_size inert, because nothing is recycled at all; with a 1GB floor my next checkpoint removed 56 segments and left 32 MB. And a floor above max_wal_size is accepted without comment, and the ceiling wins.

A high floor costs disk that max_wal_size can claim within one checkpoint cycle anyway, which means disk you already had to provision. The one other cost I found is pg_rewind, which copies the source’s spare segments along with the real ones (1,235 MB in my test, nearly all of it pool; 19 learns to skip WAL that both sides already have, and still copies the pool), so every rewind gets slower by the size of the pool. What a high floor buys is the p99.9 above, for the first checkpoint cycle of a burst that follows a long slow stretch. Neither side is large, and knowing that is most of what there is to know about this parameter. Set it equal to max_wal_size on any server whose disk was sized for that number, and let pg_wal grow to its high-water mark and stay there. If max_wal_size is deliberately enormous, set the floor to what you are willing to give pg_wal permanently. PGTune’s server profiles all hand out a quarter of max_wal_size; nothing derives that ratio, and nothing is wrong with it either. The only value that can hurt you is one under two segments.