Three ducks stand in front of a large chalkboard covered with a grid of handwritten numbers.

multixact_offset_buffers defaults to 16 buffers, and through PostgreSQL 18, 16 buffers hold the offsets of 32,768 multixacts. multixact_member_buffers defaults to 32, which holds 52,352 members. Whether those numbers are generous or absurd depends on one thing: how many multixacts your system creates while its longest row-locking transaction is open.

These are two of the seven SLRU cache sizes that PostgreSQL 17 promoted from hard-wired sizes to parameters. The commit_timestamp_buffers post has the tour of what an SLRU is, and autovacuum_multixact_freeze_max_age has the tour of what a multixact is. The short version of the second: when more than one transaction has an interest in the same row (two foreign-key checks holding FOR KEY SHARE on one parent, say), the row’s xmax stops being a transaction ID and becomes a multixact ID, and the list of transactions behind that ID lives on disk in pg_multixact. The offsets directory maps each multixact ID to a position; the members directory holds the transaction IDs and lock modes found at that position. These two parameters size the caches in front of those two directories.

Both are counts of 8kB buffers, with a minimum of 16 and a maximum of 131072 (1GB). Both must be multiples of 16. ALTER SYSTEM refuses anything else:

1ERROR: invalid value for parameter "multixact_offset_buffers": 24
2DETAIL: "multixact_offset_buffers" must be a multiple of 16.

Put the same value in postgresql.conf by hand and the server will not start. The context is postmaster, so every change costs a restart. Unlike commit_timestamp_buffers, transaction_buffers, and subtransaction_buffers, these two do not scale with shared_buffers; you get 16 and 32 on a laptop and on a machine with 512GB of RAM. On PostgreSQL 14 through 16 there is nothing to set at all. The sizes are 8 and 16, and they live in a header file.

Usually, nobody reads an old multixact

Through 18, an offsets entry is four bytes, so a page holds 2,048 of them. A member is a transaction ID plus a flag byte, and a page holds 1,636. Resolving a multixact means reading its offsets page and then one or more member pages.

Most of the time, old multixacts never get resolved. A multixact that only records row locks, and that was created before any still-open transaction first updated, deleted, or locked a row, cannot have a live member. PostgreSQL can tell that from the ID and a flag on the tuple, and skips the lookup without touching either cache. A system whose transactions are all short therefore only touches the newest page or two of each file, and the defaults are plenty. I ran 40 clients fighting over FOR KEY SHARE locks on 2,000 rows for 40 seconds on 18.6. They created about 25,000 multixacts, and blks_read for both caches in pg_stat_slru stayed at zero.

Then someone runs a batch job

Any open transaction that has updated, deleted, or locked a single row switches that shortcut off for the whole server, for every multixact created after it did so. (One that has only read, or only inserted into a table with no foreign keys, does not; I tested both.) One idle session is survivable, because the multixacts that get looked up are still the recent ones. What breaks things is a long transaction holding row locks on a great many rows that other transactions also want to lock. The ordinary way to get one is a batch that inserts a few hundred thousand child rows in a single transaction. Each insert takes FOR KEY SHARE on its parent row, and the batch holds every one of those locks until it commits.

For as long as the batch is open, every other transaction that inserts a child of one of those parents does two things. It creates a new multixact (itself plus the batch), and it reads the multixact that was on the parent row before, which can no longer be skipped because the batch is still running. Each read is of the last multixact put on that particular parent, so the reads scatter back across roughly as many multixacts as there are busy parents under the batch’s locks. With a few thousand parents, 32,768 is enough. With twenty thousand, it already is not.

Here is that situation on 18.6: one session has inserted a child row for each of 200,000 parents and has not committed, while eight clients insert children of random parents (synchronous_commit = off, counters reset after a 400,000-insert warm-up). The next 800,000 inserts at the defaults:

1 name | blks_zeroed | blks_hit | blks_read | blks_written
2------------------+-------------+----------+-----------+--------------
3 multixact_member | 978 | 853356 | 720399 | 978
4 multixact_offset | 390 | 866810 | 707404 | 390

That is nearly one physical read from each file per insert. (With 3,000 parents under the batch, the same 800,000 inserts did 37 offset reads; with 20,000 parents, 275,174.) The same 200,000-parent run with multixact_offset_buffers = 512 and multixact_member_buffers = 2048:

1 multixact_member | 978 | 1573390 | 0 | 0
2 multixact_offset | 390 | 1573753 | 405 | 216

Throughput went from about 35,500 inserts per second to about 42,000, a fifth more. That was on a two-core machine with every SLRU file sitting in the operating system’s page cache, so each miss cost a system call and not a disk read. It is the cheapest a miss will ever be.

Raising the setting buys locks as well as pages. Since 17, each SLRU cache is divided into banks of 16 buffers, each bank has its own lock, and a page belongs to the bank given by its page number modulo the number of banks (which is where the multiple-of-16 rule comes from). The multixact code takes that lock in exclusive mode even when it only wants to read. At the default, the offsets cache is a single bank, so every multixact lookup and every multixact creation on the server queues on one exclusive lock. At 512 buffers there are 32 of them. When Andrey Borodin started the thread that became these parameters, in May 2020, that lock was the complaint: it “turns MultiXact Offset subsystem to single threaded.” The fix shipped in September 2024.

The place to look is pg_stat_slru, in the rows named multixact_offset and multixact_member (spelled MultiXactOffset and MultiXactMember through 16). A healthy system shows blks_read at or near zero and not moving. blks_zeroed is useful too: it counts new pages, so on the offsets row, multiplying it by 2,048 gives the number of multixacts created since the last reset. In pg_stat_activity, the wait events to recognize are SlruRead, MultiXactOffsetSLRU, and MultiXactMemberSLRU. (I have diagnosed an outage that consisted of a wall of sessions waiting on the second of those.)

PostgreSQL 19 cuts one of them in half

PostgreSQL 19 widens multixact member offsets to 64 bits, which retires a wraparound. It also makes every offsets entry eight bytes instead of four, so a page holds 1,024 of them, and the default for multixact_offset_buffers is still 16. The default cache that covered 32,768 multixacts on 18 covers 16,384 on 19.

The same test on 19 beta 4 filled 781 offsets pages where 18.6 filled 390. With multixact_offset_buffers = 512 it did 62,109 offset reads against 18.6’s 405, and it took 1,024 buffers to get back to 404. Whatever you have set for multixact_offset_buffers on 18, double it when you upgrade. Member pages are unchanged, so multixact_member_buffers can stay where it is. PostgreSQL 19 also adds pg_get_multixact_stats(), which reports the number of multixacts and members in existence without the arithmetic.

If blks_read on those two rows is flat, leave both parameters alone; the defaults are correct for a system with short transactions, and no amount of buffer makes them more correct. If it is climbing, find the long transaction first (pg_stat_activity ordered by xact_start), and break the batch into smaller transactions if that is in your power. If it isn’t, the number of multixacts created during the longest transaction, divided by 2,048 (1,024 on 19), is a safe upper bound for the offsets cache. Set the member cache to four times that on 17 and 18, where member pages fill at least two and a half times as fast as offsets pages (faster still when rows collect more than two lockers), and to twice that on 19. Lacking a measurement, multixact_offset_buffers = 512 and multixact_member_buffers = 2048 cover a million multixacts on 17 and 18 for about 20MB of shared memory; on 19, the same coverage is 1024 and 2048, for about 25MB. It is a restart either way, so do not creep up on the number by doubling.