max_worker_processes has been 8 since PostgreSQL 9.4 introduced it in 2014, and in 2014 nothing in the server used a background worker. The only callers in the 9.4 tree were two example modules. Parallel query showed up in 9.6, logical replication in 10, extensions moved in along the way, and all of them draw on the same eight slots. Seven, on a stock install, because one is spoken for before you connect.
The parameter is the length of an array of worker slots in shared memory. The array is allocated when the postmaster starts, so the context is postmaster and changing it means a restart. The range is 0 to 262143. I went through how the worker pools nest in Workers of the World, Unite!, and the posts on max_parallel_workers and max_logical_replication_workers cover the two big consumers. This one is about the pool they share.
Who takes a slot
The logical replication launcher takes a slot at startup on any server where max_logical_replication_workers is above zero, which by default is all of them. It keeps that slot on a standby, where the launcher doesn’t run until promotion: I ran the same greedy parallel query on an 18.6 primary and on its standby, both at the default, and both launched seven workers.
After the launcher come parallel workers (for queries, index builds and VACUUM), the apply, table synchronization and parallel apply workers behind subscriptions, and anything in shared_preload_libraries that runs a process of its own. pg_cron holds a slot for its scheduler, plus one per running job if you set cron.use_background_workers. pg_prewarm holds one for the autoprewarm leader. With both of those preloaded, my seven became five.
Autovacuum workers are not in this pool; they have their own (autovacuum_worker_slots in 18, autovacuum_max_workers before that). Neither are WAL senders, the I/O workers that arrived in 18, the slot sync worker, or the checkpointer and its fixed-count colleagues.
PostgreSQL 19 moves more in. Subscriptions get a sequence synchronization worker, autovacuum can borrow parallel workers (autovacuum_max_parallel_workers, default 0), and there is a new consumer that does not degrade politely. REPACK (CONCURRENTLY) needs one slot for its decoding worker, and on beta 4 with none free:
1 => REPACK (CONCURRENTLY) r;
2 ERROR: out of background worker slots
3 HINT: You might need to increase "max_worker_processes".
That is the honest way to fail. The others are quieter. A parallel query that can’t get a slot runs with fewer workers and says nothing. A subscription writes a WARNING every five seconds and waits. A pg_cron job in background-worker mode waits ten seconds for a slot and then gets a row in cron.job_run_details with status failed and the message could not start background process; more details may be available in the server log. I produced that by running one long parallel report on a server with four slots. The scheduled job did nothing wrong. A query it had never heard of was holding the slots.
Then there is the failure at startup. A worker registered from shared_preload_libraries that doesn’t fit is dropped, and the server starts without it. This is 18.6 with pg_cron preloaded and max_worker_processes = 1:
1 LOG: too many background workers
2 DETAIL: Up to 1 background worker can be registered with the current settings.
3 HINT: Consider increasing the configuration parameter "max_worker_processes".
4 LOG: starting PostgreSQL 18.6 (Ubuntu 18.6-1.pgdg24.04+2) on x86_64-pc-linux-gnu ...
5 ...
6 LOG: database system is ready to accept connections
The complaint is printed before the startup banner and does not name the worker it refused. After that, CREATE EXTENSION pg_cron succeeded, cron.schedule() returned a job id, and nothing ran. You need more preloaded workers than slots to get here, so at the default it takes some effort. A configuration template that sets the parameter to 2 for small instances will do it as soon as someone preloads a second extension that runs a worker.
You also can’t see the pool. PostgreSQL has no view of the slot array. pg_stat_activity is the nearest thing: a background worker is any row whose backend_type is parallel worker, starts with logical replication, or is a name no core process uses, such as pg_cron launcher. It misses workers that never set themselves up for database access (the autoprewarm leader was in ps, holding a slot, with no row in the view), and it misses slots that are registered with no process running, like the launcher’s on a standby. The count you get is a floor.
The standby rule
A standby’s max_worker_processes has to be at least the primary’s. The documentation says that otherwise “queries will not be allowed in the standby server.” That is true of a standby you try to start with a lower value, which refuses to start at all. It is not true of a standby that is already running when the primary’s value goes up. Before 14 that standby shut down. Since 14 it does this, shown here on 18.6 after I raised the primary from 8 to 16 and restarted only the primary:
1 WARNING: hot standby is not possible because of insufficient parameter settings
2 DETAIL: max_worker_processes = 8 is a lower setting than on the primary server, where its value was 16.
3 LOG: recovery has paused
4 DETAIL: If recovery is unpaused, the server will shut down.
5 HINT: You can then restart the server after making the necessary configuration changes.
The standby stayed up and kept answering queries. It answered them with the data as of the primary’s shutdown, and went on doing so: a table I created on the primary afterward did not exist on the standby, pg_last_wal_receive_lsn() was well past pg_last_wal_replay_lsn(), which sat at the record announcing the change, and pg_get_wal_replay_pause_state() said paused. A health check that connects and runs SELECT 1 passes. Calling pg_wal_replay_resume() shut the server down with a FATAL, and it would not start again until its own setting was 16.
The mechanism is the same one behind max_connections, max_wal_senders and max_prepared_transactions. The primary keeps the value in pg_control, writes a WAL record when it starts up with a different one, and the standby compares on replay. So the standby finds out at the moment you restart the primary, which is the moment you were hoping for an uneventful maintenance window.
The comparison is against every value in the WAL the standby still has to replay, which is not the same as the value the primary has now. I stopped a standby, took the primary from 24 to 64 and back to 24 with a restart each time, and started the standby at 24. It paused, reporting that the primary’s value “was 64.” A replica that is behind, or one being rebuilt from an older base backup and archived WAL, has to be configured for the highest value the primary had anywhere in that stretch.
Raise it on the standbys first and the primary last. To lower it, do the primary first, let the standbys replay that restart, and then lower them; I ran both directions and neither one complained. If Patroni runs the cluster, it takes this parameter only from the shared configuration in the DCS and ignores a value set locally, for exactly this reason.
Cheap to raise, expensive to raise later
On 18.6, with shared_buffers at 2GB and max_locks_per_transaction at its default, a slot costs a little over 80kB of shared memory. That is within a couple of percent of what a max_connections slot costs, because to the shared memory code it is the same thing: worker slots are counted into the same MaxBackends total that sizes the process array and the lock tables. Going from 8 to 64 added 4.3MB. A thousand slots is 83MB. An empty slot runs no process and uses no CPU.
So there is no reason to size this one closely, and one good reason not to: it is the limit that the parallel and logical replication settings all sit under, and raising it takes a restart of every node. Use the sum from the earlier posts, which is max_parallel_workers plus max_logical_replication_workers, plus one for the launcher and one for each worker your preloaded extensions run, plus four to eight spare, and then round up to a number you won’t have to revisit. On a 16-core server that lands on 64. Put the same value on every node, standbys first, in a restart you were already taking for something else.