Re: Logical slot creation/synchronization on a standby may deadlock with recovery conflict resolution - Mailing list pgsql-hackers

From Xuneng Zhou
Subject Re: Logical slot creation/synchronization on a standby may deadlock with recovery conflict resolution
Date
Msg-id CABPTF7VarWiGB0Oa0c7YvPqkm79=3oHzgwWnfnUz2kw=TP=diQ@mail.gmail.com
Whole thread
In response to Re: Logical slot creation/synchronization on a standby may deadlock with recovery conflict resolution  (Xuneng Zhou <xunengzhou@gmail.com>)
List pgsql-hackers
On Sat, Sep 26, 2026 at 10:55 AM Xuneng Zhou <xunengzhou@gmail.com> wrote:
>
> On Thu, Sep 24, 2026 at 7:34 PM Xuneng Zhou <xunengzhou@gmail.com> wrote:
> >
> > On Thu, Sep 24, 2026 at 7:00 PM Xuneng Zhou <xunengzhou@gmail.com> wrote:
> > >
> > > On Thu, Sep 24, 2026 at 5:28 PM Bertrand Drouvot
> > > <bertranddrouvot.pg@gmail.com> wrote:
> > > >
> > > > Hi,
> > > >
> > > > On Thu, Sep 24, 2026 at 03:49:35PM +0800, Xuneng Zhou wrote:
> > > > > Hi hackers,
> > > > >
> > > > > I don't see a clear solution to this potential issue, because the interface
> > > > > is a function, which means that the held snapshots cannot be popped cleanly
> > > > > since they belong to the surrounding executor.
> > > >
> > > > Thanks for the report and reproducers!
> > > >
> > > > Thinking out loud, I wonder if we could add a transient PGPROC state while a
> > > > backend depends on recovery replay. After deadlock_timeout, ResolveRecoveryConflictWithVirtualXIDs()
> > > > could check whether a VXID in its waitlist has that state set and, if so, use the
> > > > existing recovery conflict cancellation path.
> > > >
> > > > Does that make sense to you and others? If so, I can have a look at preparing a
> > > > patch.
> > >
> > > The overall direction looks promising to me, and I haven't come up
> > > with a simpler fix. As a side benefit, it could also break the
> > > potential deadlock where 'WAIT' command waits for replay while
> > > recovery waits for the same backend's VXID. It would be good to hear
> > > more echo before heading to implementation.
> >
> > Here's the reproducer for the mentioned VXID issue. I think we need to
> > test the fix for it as well, since the underlying issue remains the
> > same. The reproducers could fit in existing test files like 031, but
> > for clarity, they are in standalone files. Also CCed Alexander for
> > this.
>
> After more investigation, both slot functions seem also vulnerable to
> VXID deadlock issue like the WAIT command. My original thought for the
> fix of the issue is to let ResolveRecoveryConflictWithSnapshot make
> the blocking decision based on the actual snapshot conflict rather
> than simply checking whether the VXID has gone away. However, that is
> more complex and needs more consideration than what you proposed. It
> could be a follow-up optimization, not necessarily the bug fix.
> Another problem that both functions suffered is the heavyweight
> deadlock issue[1].
>
> I also asked Astra to do a broader inspection for the same categorical
> issue in the tree, and it did find more, which I'll share later.
>
> [1] https://www.postgresql.org/message-id/CABPTF7U0gW5%2B-4oL7-qdML-yerZxUb7ku4QXp7JxCYo0qyJ_Tw%40mail.gmail.com

Hold on a bit. The deeper I dig, the more interesting it gets. I'll
share something very different soon..


--
Regards,
Xuneng Zhou
HighGo Software Co., Ltd.



pgsql-hackers by date:

Previous
From: Kirill Reshke
Date:
Subject: Re: ON CONFLICT DO SELECT returns rows hidden by a view
Next
From: Radim Marek
Date:
Subject: Re: REPACK (CONCURRENTLY) can't complete after ~105M concurrent updates/deletes