Re: [PATCH] Fix vacuum_delay_point happening inside lock - Mailing list pgsql-hackers

From Andres Freund
Subject Re: [PATCH] Fix vacuum_delay_point happening inside lock
Date
Msg-id s73kokddmelmc6xgetdaqvyu4agyzg6oyzzoch2wrgbc2wnz2e@fkchb773yzbo
Whole thread
In response to Re: [PATCH] Fix vacuum_delay_point happening inside lock  ("Kevin Rocker" <me@kevinrocker.com>)
List pgsql-hackers
Hi,

On 2026-09-28 17:07:52 +0200, Kevin Rocker wrote:
> > The reason this is more than a cancellation-latency issue on v19 is the
> > buffer content-lock rewrite in fcb9c977aa5 tracks only one lock per
> > buffer per backend (the single data.lockmode field). A second SHARE
> > acquire on an already-share-locked buffer is now:
> >
> >   - a hard Assert/crash in cassert builds (what unicorn shows), and
> >   - in a non-assert build, an asymmetric leak: BufferLockAttempt() adds
> >     a second BM_LOCK_VAL_SHARED to the shared state, data.lockmode
> >     records only one, and release subtracts one -- so that pg_class
> >     buffer is left permanently one shared-locker too high and can never
> >     again be locked exclusive. Any later exclusive waiter (VACUUM) on
> >     that buffer blocks for the life of the cluster.
> >
> > Pre-v19 the double SHARE was harmless (the held-lwlocks array could
> > represent it), which is presumably why the call site survived so long.

I think this was completely broken before 19 too. Acquiring a lock while
holding the same lock just happened to be undiagnosed. Note that if you ever
did this with an exclusive lock being involved, you'd just have ended up with
an uninterruptible endless wait.

It surely was never safe to call ProcessConfigFile(), or sane to sleep, while
holding an lwlock. And calling vacuum_delay_point() with interrupts held, made
it not actually properly work, due to not doing the CFI().


> This is a v19 regression from fcb9c977aa5, so I think it needs an open
> item. I don't have wiki edit access yet, so could someone from the RMT
> (Cc'd) add it, with Andres as owner?

I do not believe that fcb9c977aa5 is the culprit here.

Greetings,

Andres Freund



pgsql-hackers by date:

Previous
From: "Tristan Partin"
Date:
Subject: Re: Add counted_by attribute
Next
From: Álvaro Herrera
Date:
Subject: Re: REPACK (CONCURRENTLY) can silently lose updates when the toast table is rewritten