Re: Bug in logical decoding with DDL and subtransactions - Mailing list pgsql-hackers
| From | Zhijie Hou |
|---|---|
| Subject | Re: Bug in logical decoding with DDL and subtransactions |
| Date | |
| Msg-id | CAFvd2n_+UCM4g-QwpC0rxUv5EUnRyVy3LDJwafvbSSkz7nQyVA@mail.gmail.com Whole thread |
| In response to | Re: Bug in logical decoding with DDL and subtransactions (Tom Lane <tgl@sss.pgh.pa.us>) |
| List | pgsql-hackers |
Hi, On Wed, Sep 30, 2026 at 5:17 AM Tom Lane <tgl@sss.pgh.pa.us> wrote: > > Alexander Lakhin <exclusion@gmail.com> writes: > > 29.09.2026 21:31, Tom Lane wrote: > >> Also, of late the test_decoding/sql/ddl.sql test has been failing > >> often enough in the buildfarm to be quite annoying. So I'd like to > >> see this fixed sooner not later. (It's not very clear to me why > >> we are suddenly able to see this old bug in the regression tests. > >> The part of ddl.sql that's crashing hasn't changed in years, but > >> BF member "prion" has failed multiple times in the past two weeks. > >> Do we have a theory about that that's better than hand-wavy > >> "some change in timing"?) > > > Besides prion, skink managed to trigger that assert too: [1]. It didn't run > > tests from 2026-09-08 to 2026-09-25 [2], and never failed this test before, > > so I guess the change which affected the test was on Sep 15: a4b26b8f7 > > (the very first prion's failure includes this commit [3]). > > I doubt that theory, because a4b26b8f7 was "Revert UPDATE/DELETE FOR > PORTION OF", so to suppose that it caused this failure you'd have to > explain why we didn't see it before any of that went in. I experimented with this while reading the patch, and I think Lakhin's instinct was actually right: the onset commit is a4b26b8f7. The reason being that this commit changes the layout of the pg_attribute page. For the assert to fire, a catalog page needs three things to satisfied: (1) the aborted subtransaction's dead tuple must be prunable to LP_UNUSED, (2) the page must be full enough to trigger on-access pruning during the second ALTER, and (3) a later catalog insert in the same top-level transaction must reuse that line pointer with a different cmin. Since a4b26b8f7 removed the 36-line FOR PORTION OF section from ddl.sql (a WITHOUT OVERLAPS pkey table plus DML, pg_attribute page change), the pg_attribute page that receives tr_sub_ddl's attribute row now fills past the on-access-prune threshold during the failing section's own catalog churn, so the prune fires and the TID is reused for a new row. As for why it didn't fail before FOR PORTION OF went in: IIUC, it's also due to the pg_attribute layout, the old layout coincidentally did not satisfy the condition to trigger the page prune in the test. But the pg_attribute layout keeps changing, and other related commits (like the new system view pg_stat_kind_info, the new system catalog row pg_subscription.subconflictlogrelid, ...) keep shifting it further, which could be why we're seeing the failure on HEAD now. As for why it only fails intermittently: Michael's test shows that one factor in reproducing it is the standby snapshot WAL record, which helps advance the replication slot's catalog_xmin past the page's pd_prune_xid (pg_attribute's page in this case), making the TID where the old aborted row lived prunable so it gets reused by the second ALTER TABLE. While the logging of the standby snapshot by the bgwriter is timing-dependent. Best Regards, Zhijie Hou
pgsql-hackers by date: