REPACK (CONCURRENTLY) might keep dropped-column data - Mailing list pgsql-hackers

From Radim Marek
Subject REPACK (CONCURRENTLY) might keep dropped-column data
Date
Msg-id CAJgoLkK2UBzB1J9buCsSUbjf7bOqz-o_0=CeiTruU789BhBw-Q@mail.gmail.com
Whole thread
Responses Re: REPACK (CONCURRENTLY) might keep dropped-column data
List pgsql-hackers
Hello, 

last night I found my small issue as one of REPACK (CONCURRENTLY) testing. It's similar to an old problem with pg_squeeze reported to Antonin some time ago.

I played with the idea how it might cope under the logical replication (on subscriber) and given the previous experience with dropped columns, I managed to hit scenario where it leaves old data behind.

While the initial copy removes the dropped column data as expected, any changes that don't follow regular UPDATE path seems to retain the old value.

The table sizes shows the problem nicely

  before                                  2424 kB
  REPACK                                   224 kB
  CONCURRENTLY, no replicated updates      288 kB
  CONCURRENTLY, rows updated              2616 kB

The local apply worker seems to build the data from the original tuple, without setting dropped value to NULL. 

How to replicate:
1. On publisher create table demo(id int primary key, a text)
2. On subscriber set the table demo(id int primary key, a text, b text)
3. Populate values on publishers, set random data in b on subscriber
4. Drop the column 'b' on subscriber
5. Keep updating data on publisher
6. run REPACK (CONCURRENTLY) on subscriber table

Hope this helps. I will try to look more into the source of the problem once I have time (if needed).

Radim

pgsql-hackers by date:

Previous
From: Shubhra Jain
Date:
Subject: Re: docs: Include database collation check on SQL from alter_collation.sgml
Next
From: Narayanan Venkateswaran
Date:
Subject: Re: Proposal: Conflict log history table for Logical Replication