Re: REPACK (CONCURRENTLY) might keep dropped-column data - Mailing list pgsql-hackers

From Radim Marek
Subject Re: REPACK (CONCURRENTLY) might keep dropped-column data
Date
Msg-id CAJgoLk+aoiyMA5batXhKn9pWg_TQHB9er6U11jN1Wa63CtAgPg@mail.gmail.com
Whole thread
In response to REPACK (CONCURRENTLY) might keep dropped-column data  (Radim Marek <radim@boringsql.com>)
Responses Re: REPACK (CONCURRENTLY) might keep dropped-column data
List pgsql-hackers
Aha, so on my way to office I started thinking and got more silly ideas, and now can confirm this is more widespread than logical subscriber use case. 

There's a case for BEFORE UPDATE trigger that might RETURN OLD. I.e. the cases 

-- 1
BEGIN RETURN OLD; END

-- 2
BEGIN NEW := OLD; RETURN NEW; END

-- 3
BEGIN OLD.a := NEW.a; RETURN OLD; END

are all affected by the same problem

CREATE TRIGGER t BEFORE UPDATE ON demo
  FOR EACH ROW EXECUTE FUNCTION trg_return_old();

Confirmed by a single run 

  before                                         2288 kB
  REPACK                                           72 kB
  CONCURRENTLY, no updates                        128 kB
  CONCURRENTLY, rows updated (trigger)           2360 kB

How to replicate:
1. create table demo(id int primary key, a text, b text), fill it with random data in b
2. add a BEFORE UPDATE trigger that does RETURN OLD
3. drop column b
4. keep updating rows while running REPACK (CONCURRENTLY) demo

Radim


On Wed, 30 Sept 2026 at 08:27, Radim Marek <radim@boringsql.com> wrote:
Hello, 

last night I found my small issue as one of REPACK (CONCURRENTLY) testing. It's similar to an old problem with pg_squeeze reported to Antonin some time ago.

I played with the idea how it might cope under the logical replication (on subscriber) and given the previous experience with dropped columns, I managed to hit scenario where it leaves old data behind.

While the initial copy removes the dropped column data as expected, any changes that don't follow regular UPDATE path seems to retain the old value.

The table sizes shows the problem nicely

  before                                  2424 kB
  REPACK                                   224 kB
  CONCURRENTLY, no replicated updates      288 kB
  CONCURRENTLY, rows updated              2616 kB

The local apply worker seems to build the data from the original tuple, without setting dropped value to NULL. 

How to replicate:
1. On publisher create table demo(id int primary key, a text)
2. On subscriber set the table demo(id int primary key, a text, b text)
3. Populate values on publishers, set random data in b on subscriber
4. Drop the column 'b' on subscriber
5. Keep updating data on publisher
6. run REPACK (CONCURRENTLY) on subscriber table

Hope this helps. I will try to look more into the source of the problem once I have time (if needed).

Radim

pgsql-hackers by date:

Previous
From: Xuneng Zhou
Date:
Subject: Re: test: avoid redundant standby catchup in 049_wait_for_lsn
Next
From: Andrey Borodin
Date:
Subject: Re: Protocol Compression (fourth attempt)