Re: BUG #19686: Rolling back SET TABLESPACE - Mailing list pgsql-hackers
| From | Alexandre Felipe |
|---|---|
| Subject | Re: BUG #19686: Rolling back SET TABLESPACE |
| Date | |
| Msg-id | CAE8JnxPbj94K51gDaNJPJWsN_Pu3pquOyzibU66aHkuyebwWxg@mail.gmail.com Whole thread |
| In response to | Re: BUG #19686: Rolling back SET TABLESPACE (Manu <manuelreyesbravo@gmail.com>) |
| Responses |
Re: BUG #19686: Rolling back SET TABLESPACE
|
| List | pgsql-hackers |
Thank you Manu,
was v4 supposed to be attached?
was v4 supposed to be attached?
On Fri, Oct 2, 2026 at 3:36 PM Manu <manuelreyesbravo@gmail.com> wrote:
Hi Alexandre, Shihao,
Thanks to both. v4 attached; it changes the approach because of
Alexandre's last point, so let me take that one first.
> moving to a different tablespace should not take up more space in the
> original tablespace.
Agreed, and v4 does not. The indexes only need to follow the heap's
rewrite if the table is modified again in the same transaction, which
is also the only way the corruption can happen. So v4 leaves SET
TABLESPACE itself alone and gives an index a new relfilenumber (copied
within its own tablespace, so the indexes stay where the documentation
says) when the executor opens it to modify a table whose storage was
replaced in the current transaction, and before an in-place TRUNCATE of
such a table -- the case Shihao found. The ALTER alone in its
transaction copies nothing.
Measured with a 162 MB table and 188 MB of indexes in a nearly full
source tablespace: the move succeeds, the source needs no extra space
(v3 needed the size of the indexes there), the target receives the
162 MB heap, and the ALTER takes 0.13 s instead of 0.20 s. Modifying
the table in the same transaction after the move still needs room for
the index copies; with the source full that fails with ENOSPC at the
INSERT, cleanly (relfilenodes unchanged, nothing left behind beyond the
usual zero-length files until the next checkpoint, amcheck passes, same
after a restart).
The hook is one call in ExecOpenIndices and one in the in-place branch
of TRUNCATE. It keys on rd_firstRelfilelocatorSubid, which says whether
a relation's storage differs from what it was at the start of the top
transaction and is kept accurate for RelationNeedsWAL().
> I don't see the codebase talking about past bugs,
> the test just describes the correct behaviour.
Agreed. The comments and the test describe the behavior only; the bug
number is in the commit message.
> If we are recreating the indices, a better way to observe the effect
> is by inserting enough data to grow one page inside the transaction
> and checking pg_relation_size before and after the rollback.
Done, and it is a better test: on master the index goes from 16384 to
65536 bytes and stays there after the rollback; with the fix it is back
to 16384. The test also checks that the move alone leaves the index
file untouched, the same across a subtransaction, and the TRUNCATE
case.
> Do we need to keep that ticket number here?
Renamed to tbspace_rollback.
Shihao: thanks for the broader tests. Patches for REL_15 and REL_14
(old RelFileNode names; in 14 the test lives in tablespace.source) are
attached; each passes the regression suite on its branch.
Regards,
Manu
pgsql-hackers by date: