Parallel autovacuum: leader crashes when no DSM segment can be created - Mailing list pgsql-hackers

From Bharath Rupireddy
Subject Parallel autovacuum: leader crashes when no DSM segment can be created
Date
Msg-id CALj2ACVjBN03Fq4XWgKH1zOfkapK6-P3ufNLHUHnZcW1zPm+Mg@mail.gmail.com
Whole thread
Responses Re: Parallel autovacuum: leader crashes when no DSM segment can be created
List pgsql-hackers
Hi,

AI review found a bug in parallel autovacuum (1ff3180ca01). I checked
it myself and it reproduces on HEAD and PG19. Patch and reproducer
attached.

When no DSM segment can be created, InitializeParallelDSM() does not
fail. It sets up the parallel context in the leader's own memory with
zero workers and leaves pcxt->seg NULL. Parallel query checks that
pointer before using it, see ExecInitParallelPlan(), and VACUUM
(PARALLEL) never looks at it and vacuums all the indexes in the
leader. Parallel autovacuum registers a DSM detach callback on it, so
the autovacuum worker segfaults and the database goes through crash
recovery. It also leaves pv_shared_cost_params pointing into the
leader's private memory with nothing left to reset it.

It is not easy to hit. Right after the fallback, vacuum creates the
DSA area for its dead items, and with no segments left that fails with
a plain "too many dynamic shared memory segments" error before the
callback is reached. Another backend has to free a segment in the
window between the two.

The fix is to set up the cost parameters and the callback only when
the context has workers.

The attached script fills the segments with parallel queries and keeps
autovacuum busy on tables needing parallel index vacuuming. Most
attempts fail with that error, then one lands in the window and
crashes the worker, usually within a few seconds.

[1]
2026-09-26 21:31:55.790 UTC [7908] LOG:  autovacuum worker (PID 8904)
was terminated by signal 11: Segmentation fault
2026-09-26 21:31:55.790 UTC [7908] DETAIL:  Failed process was
running: autovacuum: VACUUM ANALYZE public.t3

#0  slist_push_head (head=0x38, node=0x10468ff8) at
../../../../src/include/lib/ilist.h:1008
#1  on_dsm_detach (seg=0x0, function=0x775bce
<parallel_vacuum_dsm_detach>, arg=0) at dsm.c:1148
#2  parallel_vacuum_init (rel=0x7f39f5bf0cf0, indrels=0x104be9a0,
nindexes=3, nrequested_workers=2, vac_work_mem=65536, elevel=13,
bstrategy=0x104ec410) at vacuumparallel.c:470
#3  dead_items_alloc (vacrel=0x104be050, nworkers=2) at vacuumlazy.c:3471
#4  heap_vacuum_rel (rel=0x7f39f5bf0cf0, params=0x7fffcadde380,
bstrategy=0x104ec410) at vacuumlazy.c:863
#5  table_relation_vacuum (rel=0x7f39f5bf0cf0, params=0x7fffcadde380,
bstrategy=0x104ec410) at ../../../src/include/access/tableam.h:1802
#6  vacuum_rel (relid=16399, relation=0x10501360, params=...,
bstrategy=0x104ec410, isTopLevel=true) at vacuum.c:2353
#7  vacuum (relations=0x10501480, params=0x104fd5e0,
bstrategy=0x104ec410, vac_context=0x10501260, isTopLevel=true) at
vacuum.c:635
#8  autovacuum_do_vac_analyze (tab=0x104fd5d8, bstrategy=0x104ec410)
at autovacuum.c:3351
#9  do_autovacuum () at autovacuum.c:2535
#10 AutoVacWorkerMain (startup_data=0x0, startup_data_len=0) at
autovacuum.c:1634
#11 postmaster_child_launch (child_type=B_AUTOVAC_WORKER,
child_slot=10, startup_data=0x0, startup_data_len=0, client_sock=0x0)
at launch_backend.c:268
#12 StartChildProcess (type=B_AUTOVAC_WORKER) at postmaster.c:4043
#13 StartAutovacuumWorker () at postmaster.c:4107
#14 process_pm_pmsignal () at postmaster.c:3864
#15 ServerLoop () at postmaster.c:1720
#16 PostmasterMain (argc=3, argv=0x1043c510) at postmaster.c:1414
#17 main (argc=3, argv=0x1043c510) at main.c:227

--
Bharath Rupireddy
Amazon Web Services: https://aws.amazon.com

Attachment

pgsql-hackers by date:

Previous
From: Bharath Rupireddy
Date:
Subject: Parallel autovacuum: DROP DATABASE WITH (FORCE) fails on the parallel workers
Next
From: Bharath Rupireddy
Date:
Subject: Parallel vacuum: wrong error context when the leader vacuums an index