Vintage illustration of a lumber yard with a green delivery truck labeled "Quack Wall Delivery" loaded with bricks, surrounded by wooden structures, stacked lumber, and other period trucks in a forest setting.

Errata: the max_connections post told you to size that parameter for “the replication connections.” On 12 and later, don’t. WAL senders have their own seats, sized by this parameter, and have not drawn from max_connections since PostgreSQL 12. The rest of that post stands; that clause is wrong for every version you should be running.

max_wal_senders is the number of processes allowed to read your WAL on someone else’s behalf, and the reason it bites is that “someone else” is a longer list than the standbys you were counting. Both connections of every default pg_basebackup are on it. Every logical replication subscription is on it, plus one more per table that subscription is currently copying. pg_receivewal and everything built on it (Barman’s streaming mode, for one) are on it. And every client that just dropped off the network without closing its socket is still on it, for up to a wal_sender_timeout after it went quiet.

The default is 10, the context is postmaster, and the range is 0 to 262143. It was 0 until PostgreSQL 10, when Magnus Hagander and Dang Minh Huong raised it and max_replication_slots to 10, moved wal_level to replica and hot_standby to on, and streaming backup and replication started working out of the box. wal_level still has to be replica or logical, and the check is in the postmaster, so a wal_level = minimal you set for a bulk load, with this parameter left at its default, does this instead of starting:

1FATAL: WAL streaming ("max_wal_senders" > 0) requires "wal_level" to be "replica" or "logical"

Zero, which the documentation describes as “replication is disabled,” disables less than that. With max_wal_senders = 0 and wal_level = logical on 18.6 I created a logical slot with pg_create_logical_replication_slot() and read changes from it with pg_logical_slot_get_changes(), because those run in an ordinary backend. What zero disables is WAL sender processes: physical standbys, logical subscriptions, and pg_basebackup, which fails with the same message as a full pool, currently 0.

Their own pool, for better and worse

Before 12, a WAL sender took a seat from max_connections, which meant a busy application could lock a standby out of reconnecting, and superuser_reserved_connections plus this parameter had to fit under max_connections or the server wouldn’t start. Alexander Kukushkin’s fix for 12 gave them a free list of their own, and the arithmetic changed on both sides. Here is 18.6 with max_connections = 5 and every seat taken by an idle session:

1$ psql -c "SELECT 1"
2psql: error: connection to server on socket "/tmp/.s.PGSQL.5432" failed: FATAL: sorry, too many clients already
3$ pg_receivewal -D rw1 &
4$ pg_receivewal -D rw2 &
5$ pg_basebackup -D bb -X stream
6pg_basebackup: error: connection to server on socket "/tmp/.s.PGSQL.5432" failed: FATAL: number of requested standby connections exceeds "max_wal_senders" (currently 2)

Both pg_receivewal clients connected without complaint while psql could not, because they never asked for a regular seat. None of the rules for regular seats apply to a replication connection: not max_connections, not the reserved seats, not a role’s or a database’s CONNECTION LIMIT. What can turn it away is its own limit, pg_hba.conf, authentication, and the REPLICATION attribute. That is the better half. The worse half is that nothing borrows in the other direction either: a full max_wal_senders is full while ninety application seats sit idle beside it, and the message that tells you so arrives at the standby or the backup job, not at the application, and often at 3 a.m.

pg_basebackup is the usual way to find this. Since 10 it streams WAL on a second connection while the backup runs (-X stream, the default), so it needs two seats at once. With one seat free it connects, starts the backup, fails to open the second connection, and removes everything it wrote on its way out; -X fetch needs one seat and succeeded on the same server a moment later. A logical subscription is the other way: one seat for the apply worker and up to max_sync_workers_per_subscription more while it initializes, all of them WAL senders on the publisher, which max_logical_replication_workers has already had to explain once.

The seats the dead are sitting in

The documentation says the parameter “should be set slightly higher than the maximum number of expected clients,” and the reason is worth a transcript. I started two pg_receivewal clients against a two-seat server with wal_sender_timeout at 15s, then sent both SIGSTOP: the processes stay alive, the sockets stay open, and nothing on the server side has any way to know the client isn’t coming back.

1$ kill -STOP <both pg_receivewal pids> # 14:35:46
2$ pg_basebackup -D bb -X fetch --no-slot
3pg_basebackup: error: ... FATAL: number of requested standby connections exceeds "max_wal_senders" (currently 2)
4$ psql -c "SELECT pid, state, now() - reply_time AS since_last_reply FROM pg_stat_replication"
5 pid | state | since_last_reply
6-----+-----------+------------------
7 726 | streaming | 00:00:06
8 727 | streaming | 00:00:06
914:35:59 walsender LOG: terminating walsender process due to replication timeout
1014:35:59 walsender LOG: terminating walsender process due to replication timeout
11$ pg_basebackup -D bb -X fetch --no-slot # exit status 0

Thirteen seconds in this run, on a server configured to notice quickly; the clock runs from the last reply the WAL sender received, not from the moment the client froze, so the wait is at most the timeout and usually a little less. At the default wal_sender_timeout of a minute, a standby that lost its network and came back has to wait for its own ghost to be evicted before it can reconnect, and if the pool was sized exactly, so does everyone else. (A client whose process dies closes its socket and is gone at once. The ghost is the one that went silent: a pulled cable, a frozen VM, a NAT that timed out.) A pg_stat_replication row in streaming with a stale reply_time is that ghost; it looks healthy right up until it is killed.

Two more things the pool has to cover. A cascading standby serves its own downstream from its own max_wal_senders, so it needs seats for them, not just for itself. And since 12 a hot standby must have this parameter at least as high as its primary, for the same reason as max_connections and max_prepared_transactions: the value is in pg_control and in the WAL, and a hot standby that replays a larger value from its primary pauses recovery (on 14 and later; 12 and 13 shut down instead) until you restart it with a value that fits, which max_prepared_transactions demonstrates. Copy the primary’s configuration to the standby and this never comes up.

A seat costs what a connection costs: on 18.6, a thousand of either adds 51 MB of shared memory, and going from the default to 60 cost 3.8 MB when I measured it; the 262,143 ceiling is MAX_BACKENDS because a WAL sender is, for shared-memory purposes, a backend (the four backend pools together have to fit under that number, so nobody gets all of it). There is no performance reason to keep it small, and there is a restart between you and any correction. Set it to 50 on every node, in the same restart as max_replication_slots, whose post already asks you to keep this one at least as large as it plus your physical standbys. Then do the real count: standbys, two per base backup you’d run at once, one per subscription plus its synchronization workers, one per streaming archiver, and a wal_sender_timeout’s worth of ghosts on top. If that number is anywhere near 50, you already knew it wasn’t a small installation.