Re: Parallel autovacuum: DROP DATABASE WITH (FORCE) fails on the parallel workers - Mailing list pgsql-hackers

From Masahiko Sawada
Subject Re: Parallel autovacuum: DROP DATABASE WITH (FORCE) fails on the parallel workers
Date
Msg-id CAD21AoDUTngHM7PRzNEr=EhGU47rWDQOcgFnxPWiBYW2VEQ75g@mail.gmail.com
Whole thread
In response to Parallel autovacuum: DROP DATABASE WITH (FORCE) fails on the parallel workers  (Bharath Rupireddy <bharath.rupireddyforpostgres@gmail.com>)
List pgsql-hackers
Hi,


On Sun, Sep 27, 2026 at 5:00 PM Bharath Rupireddy
<bharath.rupireddyforpostgres@gmail.com> wrote:
>
> Hi,
>
> AI review found a bug in parallel autovacuum (1ff3180ca01). I checked
> it myself and it reproduces on HEAD and PG19. Patch with a test
> attached.
>
> A role with pg_signal_backend that is not a superuser gets "permission
> denied to terminate process" [1] from DROP DATABASE WITH (FORCE) while
> a parallel autovacuum on that database is in its index phase, which on
> a large table is where the vacuum spends its time. The same role can
> terminate the autovacuum worker itself, and DROP DATABASE without
> FORCE succeeds with the same vacuum running.
>
> Both the autovacuum worker and its parallel workers run as the
> bootstrap superuser, so that is not what separates them.
> TerminateOtherDBBackends() looks at the role a process published in
> its PGPROC, and only the workers published one. The autovacuum worker
> gets its user from InitializeSessionUserIdStandalone(), which sets
> AuthenticatedUserId directly and never calls SetAuthenticatedUserId(),
> so MyProc->roleId stays InvalidOid and superuser_arg() on it is false.
> Its parallel workers take the same user through ParallelWorkerMain(),
> which does call SetAuthenticatedUserId(), so they publish the
> bootstrap superuser and the check refuses them.
>
> A parallel worker of a VACUUM command is not affected, since its
> leader is a user session and the worker publishes the same role as its
> leader. Autovacuum is the only leader that publishes no role while its
> workers publish one.
>
> The fix treats a process whose lock group leader is an autovacuum
> worker the way the autovacuum worker itself is treated. The patch adds
> a test to the test_autovacuum module that fails with this error
> without the fix.

Thank you for the report and the patch!

IIUC the issue stems from the fact that the leader and its workers
advertise different roleIds (InvalidOid and BOOTSTRAP_SUPERUSERID). I
think the same issue can be reproduced in other cases. For instance,
suppose that a bgworker connecting to the database via
BackgroundWorkerInitializeConnection(dbname, NULL, 0) runs a parallel
query, the leader's roleId is InvalidOid whereas the parallel query
workers have BOOTSTRAP_SUPERUSERID. I think we should fix it as well
and backpatch the fix to 14.

I have some review comments on the proposed patch:

+               if (leader != NULL && leader != proc &&
+                   leader->backendType == B_AUTOVAC_WORKER)
+                   roleId = InvalidOid;

I think we should check a lock group member with its leader's roleId
instead of unconditionally using InvalidOid. That would
straightforwardly fix the inconsistency between the leader and the
workers.

    if (leader != NULL && leader != proc &&
        leader->databaseId == databaseId)
        roleId = leader->roleId;

To fix this issue not only in autovacuum cases, the backendType check
should be removed.

Regards,

--
Masahiko Sawada
Amazon Web Services: https://aws.amazon.com



pgsql-hackers by date:

Previous
From: Vik Fearing
Date:
Subject: Re: ON EMPTY clause for aggregate and window functions
Next
From: "Matheus Alcantara"
Date:
Subject: Re: [PATCH v1] Fix for Bug#19724 - ALTER TYPE ... ALTER ATTRIBUTE triggers internal error for base type of domain with check