max_wal_size is not a maximum, and it is not the size of anything you can du. It is a checkpoint trigger denominated in bytes: once the WAL written since the last checkpoint started reaches a fixed fraction of this value (the fraction is below, and it is not 1), the next one starts, whether or not checkpoint_timeout has come around. The checkpoint tour covers the two triggers and why you want the clock to win; I have also written about this parameter twice before. What those posts don’t have is the arithmetic, the measurement, or the two ways the name lies to you, so that is this one.
The default is 1GB, the context is sighup, the unit is megabytes, and the range runs from 2 to 2147483647 (about two petabytes; nobody has tested it). It arrived in 9.5 when Heikki Linnakangas replaced checkpoint_segments, whose default of 3 had been triggering a checkpoint every 48 MB of WAL since 7.1 in 2001. The release notes gave a conversion ((3 * checkpoint_segments) * 16MB) and noted that the new default was already far higher than the old one, which was true and also a very low bar.
The number the trigger actually uses
The checkpointer does not fire at max_wal_size. It fires at max_wal_size / (1 + checkpoint_completion_target), rounded down to whole segments, because the parameter is defined as the most WAL that should exist between one checkpoint’s redo pointer and the end of the next one, and the server keeps writing WAL while a spread checkpoint runs. With the defaults on 14 and later that is 64 / 1.9 = 33 segments, 528 MB, and you can see the exact figure in every log_checkpoints line:
1 checkpoint complete: wrote 36828 buffers (56.2%), ... 0 WAL file(s) added, 0 removed, 33 recycled; write=22.929 s, sync=0.150 s, total=23.108 s; ... distance=540676 kB, estimate=540681 kB; ...
That distance is the WAL written between the last two checkpoint starts, estimate is the moving average the server uses to decide how many segments to keep around, and both sit on 33 × 16 MB. The trigger has moved twice without the parameter changing: before 11 the server kept WAL for two checkpoint cycles and divided by 2 + checkpoint_completion_target instead, so 1GB meant a checkpoint every 400 MB, and 11 through 13 (with the old 0.5 target) meant 672 MB. Same setting, three different servers.
What 1GB costs
Here is the same pgbench workload, four clients against a scale-50 database in 512 MB of shared_buffers, three minutes each, checkpoint_timeout at 15min, on 18.6:
1 max_wal_size = 1GB 3,700 tps WAL written: 3815 MB checkpoints: 7 (requested) wal_fpi: 447,963
2 max_wal_size = 8GB 4,131 tps WAL written: 1080 MB checkpoints: 0 wal_fpi: 91,335
Three and a half times the WAL for the same transactions, and 10% fewer of them in this run; and the ratio understates it, because the 8GB run’s three minutes still began with a checkpoint’s worth of re-imaging, which a fifteen-minute cycle pays once. The mechanism is full_page_writes: the first change to any page after a checkpoint writes the whole 8 kB page into the WAL, and at 1GB this workload was starting a checkpoint every 22 to 28 seconds, so every hot page was re-imaged that often. pg_stat_wal.wal_fpi is where to look for this on your own system; at 1GB it was 10% of all WAL records and, with wal_compression off, about 90% of the bytes. The checkpointer also wrote 263,737 buffers in those three minutes and none at all in the other run (backends and the background writer took up part of that slack, not all of it). On a real server every one of those extra gigabytes would also be archived, shipped to every replica, and read by every logical decoder, and the log said checkpoints are occurring too frequently (23 seconds apart) with a hint naming this parameter, seven times, which is the one place PostgreSQL tells you what to do.
The two ways the name lies
First, it is not the size of pg_wal. At the end of a checkpoint the server removes segments older than the new redo pointer, but it recycles the first several into preallocated future segments, keeping at least min_wal_size worth and at most max_wal_size worth measured from that checkpoint’s redo pointer, based on that estimate. So under a steady write load pg_wal sits near max_wal_size, which is presumably where the name came from, and the documentation is right to call it “a soft limit.” It is soft in one direction only: an inactive replication slot, a failing archive_command, or wal_keep_size all keep segments that have already been written, and nothing here stops them (the slot case is max_slot_wal_keep_size’s job; the other two are yours). It is also, less obviously, soft in the other direction. Lowering it frees nothing on the spot. Segments already preallocated ahead of the insert point are never revisited by a checkpoint; on my server, dropping the setting from 4GB to 1GB and checkpointing three times left pg_wal at 2.9 GB, all 184 of those segments sitting in front of the current position, waiting to be written through. It came back down a checkpoint at a time, as each cycle unlinked old segments instead of recycling them, and reached 1 GB after about 1.9 GB of new WAL had gone by.
Second, “maximum” suggests a cost you pay by going over, when the real cost is paid by going under, and the only cost of going high is the one Lukas Fittl raised on the earlier post: after a crash, recovery replays from the last completed checkpoint’s redo pointer, and if the crash lands near the end of the next checkpoint that is the whole trigger distance plus everything written while the checkpoint ran: a full max_wal_size, which is what the parameter was defined to bound. I measured it. Killing the postmaster with 1,080 MB of WAL since the last checkpoint cost 6.0 seconds of redo; with 1,627 MB, 12.5 seconds; and 682 MB of the full-page-image-heavy WAL from the 1GB regime took 1.5 seconds, because restoring a page image is a copy while an ordinary record has to read the page first. Call it 10 to 12 seconds per gigabyte of ordinary WAL on this box, single-threaded, with the data in cache. At 8GB the worst case is a minute and a half here; at 16GB, three. Your disks and your rows per record will move that number, but it scales with the WAL you let accumulate, and it is the whole trade.
A minute or two of recovery on a system that crashes rarely, against 3.5× the WAL volume on one that checkpoints every 25 seconds, is not a close call. Set checkpoint_timeout to 15min and max_wal_size to 8GB, run for a week, and grep the log for checkpoint starting: wal. If it appears outside a bulk load, double max_wal_size and look again; if it appears only during the nightly import, that is what the parameter is for, leave it. If you have a hard recovery-time budget, divide it by your measured seconds per gigabyte and that is your ceiling; measure it, on your hardware, with your WAL, because mine is not yours. The disk it takes is the cheapest thing in the building.