Flock of ducks wearing small signs waddling across a residential street lined with single-story houses and a parked vintage car.

max_sync_workers_per_subscription is how many tables a subscription copies at once. It is not how fast the initial copy goes, which is what people raise it hoping for. Each table synchronization worker copies exactly one table, start to finish, so the parameter buys parallelism across tables and nothing within one. Your largest table takes as long as it takes with this set to 2, 8, or 200, and the initial copy of a subscription cannot finish before that table does.

The default is 2. The context is sighup, the range is 0 to 262143, and none of that has changed since the parameter arrived with logical replication in PostgreSQL 10. The workers come out of the pool sized by max_logical_replication_workers, which is where the arithmetic and the “out of logical replication worker slots” warning live; this post assumes you have read that one. The one thing to add about the pool here is that this limit is per subscription, so three subscriptions initializing at the same time want three times this many free slots.

Reload really does mean reload. The apply worker re-reads the setting in its main loop, so raising it takes effect on a copy that is already in progress: on 18.6 I started a six-table subscription at 1, reloaded with 3 four seconds in, and by the next second’s sample there were three synchronization workers running. You can raise it mid-migration without touching the subscription.

What one worker costs, and on which server

A synchronization worker is a full logical replication client, and on the subscriber it costs a worker slot from the pool and one replication origin (from max_active_replication_origins on 18 and later; from max_replication_slots before that). On the publisher it costs more than that. Here is the publisher during the copy of one 1.1 GB table on 18.6:

1postgres=# SELECT application_name, state, sent_lsn FROM pg_stat_replication;
2 application_name | state | sent_lsn
3-----------------------------------------+-----------+------------
4 pg_16587_sync_16489_7690312813895012441 | startup |
5 sub | streaming | 3/EE5EA350
6
7postgres=# SELECT slot_name, temporary, active FROM pg_replication_slots;
8 slot_name | temporary | active
9-----------------------------------------+-----------+--------
10 pg_16587_sync_16489_7690312813895012441 | f | f
11 sub | f | t
12
13postgres=# SELECT application_name, backend_xmin, left(query, 40) AS query
14 FROM pg_stat_activity WHERE backend_type = 'walsender';
15 application_name | backend_xmin | query
16-----------------------------------------+--------------+------------------------------------------
17 pg_16587_sync_16489_7690312813895012441 | 1171656 | COPY public.t_big (id, payload) TO STDOU
18 sub | | START_REPLICATION SLOT "sub" LOGICAL 0/0

Each worker is a WAL sender (so count it against max_wal_senders), a permanent replication slot named pg_<subscription oid>_sync_<table oid>_<system identifier>, all three taken from the subscriber, so resolving that table OID on the publisher gets you the wrong table (count the slot against max_replication_slots, whose post covers what happens when a publisher runs out of them), and a REPEATABLE READ transaction that stays open for the whole COPY. That backend_xmin is the important column. For as long as the copy runs, VACUUM on the publisher cannot remove any row in that database that died after the snapshot was taken, in any table, not just the one being copied. Eight workers on a busy publisher are eight of those. (The slot showing active = f while its walsender is visibly busy is normal; the walsender creates the slot, does the COPY outside it, and only acquires it for the catch-up phase. Do not drop it.)

The copy is one transaction on the subscriber too, so the table is empty right up until it isn’t, and the subscriber’s indexes are maintained row by row as it fills.

The table you didn’t pick and the pause you didn’t expect

You don’t choose the order. The apply worker walks pg_subscription_rel in physical order and starts a worker for each not-ready table it finds while it has a free slot, so the order is, to begin with, whatever order CREATE SUBSCRIPTION inserted the rows, which is whatever order the publisher listed them (state changes are heap updates, so it drifts from there). It is not by size, and it is not by name.

A table whose copy fails is retried from the beginning, every time. The worker exits, and the apply worker starts another one as soon as wal_retrieve_retry_interval has passed since that table’s previous start (for any copy that ran longer than five seconds, that is no wait at all), and the new worker drops the old slot on the publisher, creates a new one, and starts the COPY over. I put one conflicting row into the subscriber’s copy of that 1.1 GB table before subscribing:

113:40:55.849 [7110] logical replication tablesync worker LOG: logical replication table synchronization worker for subscription "sub", table "t_big" has started
213:41:27.009 [7110] logical replication tablesync worker ERROR: duplicate key value violates unique constraint "t_big_pkey"
313:41:27.009 [7110] logical replication tablesync worker DETAIL: Key (id)=(6000000) already exists.
413:41:27.009 [7110] logical replication tablesync worker CONTEXT: COPY t_big, line 6000000
513:41:27.022 [4208] postmaster LOG: background worker "logical replication tablesync worker" (PID 7110) exited with exit code 1
613:41:27.025 [7118] logical replication tablesync worker LOG: logical replication table synchronization worker for subscription "sub", table "t_big" has started

Thirty seconds of copy, an error on the last row, and three milliseconds later a new worker starts the thirty seconds again. It did that for as long as I let it, with pg_stat_subscription_stats.sync_error_count ticking up once per attempt, and every attempt occupied one of the two sync worker slots for the duration. With the default, one bad table halves your parallelism for the rest of the migration, and the log and that counter are the only places it shows. On 15 and later, disable_on_error = true on the subscription turns that loop into one error and a disabled subscription, which is the version of this you want.

The other thing a worker does to the leader happens when it succeeds. Once the copy finishes, the worker has to catch its table up from the copy’s snapshot to wherever the leader apply worker is now, and the leader waits for it: it sits in the LogicalSyncStateChange wait event, applying nothing for any table, until the sync worker reports done. The sync worker’s WAL sender decodes everything the publisher wrote during the copy to find that one table’s changes, and sends it every published table’s changes along the way, to be discarded on arrival. I ran two clients of 2 kB updates against an already-replicating table during a copy that took 34 seconds; when the copy finished, the sync worker’s WAL sender was 681 MB behind the publisher, the leader’s own slot fell 40 MB behind while it waited, and the pause lasted about two seconds. That is the cheap case. The pause scales with the WAL the publisher wrote during the copy, and a copy that took all afternoon on a busy publisher earns a proportionally longer stall at the end of it, once per table, one table at a time.

The two ends of the range

Zero is a trap. With max_sync_workers_per_subscription = 0, CREATE SUBSCRIPTION succeeds, the apply worker starts, every table sits at srsubstate = 'i' forever, and nothing is logged, because nothing is failing; the apply worker simply never asks for a worker. Rows inserted on the publisher meanwhile are discarded by the leader, since it only applies changes for tables that are ready. I could not find a legitimate use. Reload with a positive value and the copies begin.

PostgreSQL 19 adds sequence synchronization, done by one worker per subscription for all its sequences, and that worker draws from this same per-subscription cap. The documentation says “one additional worker is also needed for sequence synchronization”, which I read as sizing advice, not as an exemption: on 19 beta 4 with the cap at 1, the sequence worker took its turn in the single slot between two table copies and finished in ten milliseconds. It is not a reason to raise the number.

Leave it at 2 on a subscriber that is going to sit there replicating. For a migration, set it to the number of tables you can afford to copy at once, which is bounded by cores and disks on both servers (each copy is a COPY ... TO STDOUT on a publisher core, holding an xmin there, feeding a COPY FROM with index maintenance on a subscriber core), and reserve that many WAL senders and slots on the publisher before you start. On my two-core test box, with both servers on it, six workers finished the same six tables in 14.5 seconds and two finished them in 13.3, which is what running out of cores looks like in a transcript. Then put it back. The next ALTER SUBSCRIPTION ... REFRESH PUBLICATION someone runs on a Tuesday afternoon will start that many copies against production, and it will not ask first.